Skeleton detection and fall detection method based on improved spatio-temporal adaptive graph convolution
Through the improved space-time adaptive graph convolution network combined with YOLOv5 and Deepsort algorithm, the problem of the existing fall detection technology degradation in lighting, occlusion and viewing angle changes is solved, and high accuracy and robust fall detection in complex environments is achieved.
Patent Information
- Application Number
- PCT/CN2024/099750
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-06-18
- Publication Date
- 2025-06-12
AI Technical Summary
The existing fall detection technology reduces the detection accuracy when light changes, occlusion or viewing angle changes, and traditional methods rely on specific sensors or high-demand environmental conditions, making it difficult to adapt to complex multi-person scenes and harsh environments.
The improved space-time adaptive graph convolution network is adopted, combined with YOLOv5 object detection and Deepsort object tracking algorithm, and the accuracy and robustness of fall detection are improved by automatically learning and extracting features and adapting to different fall modes and environmental conditions.
High accuracy and robust fall detection under different fall modes and environmental conditions is achieved, reducing dependence on specific sensors and environmental conditions, and is suitable for complex multiplayer scenarios and harsh environments.
Smart Images

Figure CN2024099750_12062025_PF_FP_ABST
Abstract
Description
Skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 6, 2023, with application number 2023116620681 and invention name “Skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application belongs to the field of artificial intelligence technology, and in particular relates to a skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution. Background Art
[0003] Existing technologies use computer vision-based approaches, which use cameras or depth sensors to capture human motion and posture, then analyze key points or skeleton data for fall detection. However, computer vision methods are sensitive to changes in lighting, occlusion, or perspective, which can lead to reduced detection accuracy and present significant drawbacks in practical applications.
[0004] Existing technologies use machine learning-based approaches, which leverage machine learning algorithms to train models to identify characteristic patterns in falling behavior. However, traditional machine learning methods often require manual feature extraction and have limited ability to model complex spatiotemporal relationships.
[0005] Existing technologies use sensor-based fall detection algorithms, which primarily rely on sensors such as triaxial accelerometers, gyroscopes, and pressure sensors to detect human motion and locate the body's position to identify falls. Yodpijit et al. used a combination of accelerometers and gyroscopes to detect falls and employed an artificial neural network (ANN)-based algorithm to distinguish falls from everyday behavior. Chen KH et al. used smartphones equipped with motion sensors to detect falls. The smartphone's motion sensors and transmission module calculate changes in human motion and send the data to a server for analysis and fall detection. However, wearable devices are typically expensive and can easily cause discomfort to the wearer, making them less universally applicable. Therefore, detecting falls in public places without sensor-equipped devices is unrealistic. Methods based on ambient devices, like wearable devices, also rely on sensors. However, they primarily rely on devices such as radar or ground sensors to collect current or audio information to determine whether a fall has occurred. Because ambient devices are highly sensitive to interference such as external noise, these methods are less effective in complex multi-person scenarios.
[0006] Existing technologies use traditional computer vision-based fall detection algorithms. With the rapid development of computer vision technology, more and more researchers are focusing on behavior recognition. Human behavior recognition is widely used in security monitoring, human-computer interaction, and virtual reality. Due to the limitations of sensor-based approaches, computer vision-based fall detection has received increasing attention in recent years. Computer vision-based approaches do not require sensor devices; they only require cameras and can identify falls through video feeds, achieving higher accuracy than sensor-based approaches. Wang et al. proposed a new foreground segmentation model to detect pedestrians, detecting falls based on changes in pedestrian silhouettes. Zerrouki et al. used a hidden Markov model to identify and classify falls based on pedestrian silhouette morphology features. Yu et al. combined a hidden Markov model with an accelerometer. The sensor data was used to train the hidden Markov model separately, and a direction calibration algorithm was used to compensate for errors. Deng Zhifeng et al. combined geometric features and constructed a pedestrian circumscribed matrix to identify falls in different directions. Fan et al. classified falls by calculating dynamic and static features separately. Miao et al. used an ellipse fitting method to wrap the human body contour with an ellipse. They then combined geometric features and position information with support vector machines to identify and classify different falls. However, this traditional method focuses on feature extraction and classification and is susceptible to noise, illumination changes, and occlusion.
[0007] The above fall detection methods have several drawbacks: Sensor-based fall detection methods rely on specific sensors, which limits system deployment and usage. Sensor devices must be installed and configured in specific environments. Furthermore, the detection range and coverage of sensor-based methods are often limited by the sensor's sensing range. For example, a fixed-position accelerometer can only detect falls within a fixed range and cannot cover the entire environment. Sensor-based methods also require real-time acquisition, processing, analysis, and judgment of sensor data. This places demands on the system's real-time and stability, and relies on the quality and accuracy of the sensor data. Fall detection algorithms based on traditional computer vision methods have high environmental requirements, requiring a clear camera field of view and good lighting conditions to accurately capture and analyze human posture and motion information. Algorithm performance may degrade in complex or harsh environments. Traditional computer vision methods often rely on manually designed feature extraction methods, requiring the design and selection of features tailored to the fall detection task. However, this manual feature extraction is often influenced by human experience and subjective factors, failing to fully utilize the data's potential.
[0008] Chinese patent document (CN112966628A) discloses a perspective-adaptive multi-target fall detection method based on a graph convolutional neural network, comprising the following steps: using a target detection algorithm to detect the human target in each frame of the target video source, using a posture estimation algorithm to extract the key skeleton point data of the human target in each frame, and when the number of frames in which the same human target is detected continuously exceeds a preset detection threshold, inputting the extracted key skeleton point data into a trained perspective-adaptive sub-network to obtain perspective adjustment parameters; performing perspective adjustment on the key skeleton point data according to the perspective adjustment parameters, and then calculating motion data based on the perspective-adjusted key skeleton point data. The perspective-adjusted key skeleton point data and motion data are then input into a trained graph convolutional fall recognition main network for fall detection, and a detection result label is output. This method fails to fully utilize the potential information of the data.
[0009] Therefore, the present application provides a skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution, which is adaptable to different fall patterns and environmental conditions, and the method has higher accuracy and robustness.
[0010] Summary of the Invention
[0011] The technical problem to be solved by this application is to provide a skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution, which is adaptable to different fall patterns and environmental conditions, and the method has higher accuracy and robustness.
[0012] In order to solve the above technical problems, the technical solution adopted in this application is: the skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution specifically includes the following steps:
[0013] S1: collect image data and obtain each frame of image data;
[0014] S2: Use the pre-trained yolov5 target person detection model to detect whether the target person appears in each frame of image data. If the target person appears, go to step S3; if not, end;
[0015] S3: For each detected target person, use the Deepsort target tracking algorithm to track the target, obtain the tracking result, calculate the similarity to obtain the target association result, and update the trajectory information of each target person;
[0016] S4: Perform posture recognition on each target person based on trajectory information, and use a spatiotemporal adaptive graph convolutional network to extract the feature vector of the posture. Use a classifier to classify and identify human behavior to determine whether the target person has fallen.
[0017] The above technical solution uses deep learning methods to automatically learn and extract features, adapting to different fall patterns and environmental conditions. It can also be trained and optimized end-to-end from a large amount of data, resulting in a deep learning-based fall detection algorithm with higher accuracy and robustness. By applying an adaptive graph convolutional network to the ST-GCN fall detection model, the model's flexibility, adaptability, and representation capabilities are enhanced, enabling more accurate human action recognition and spatiotemporal feature modeling. This improves the modeling capabilities of complex spatiotemporal features, enhancing the model's flexibility and adaptability. The adaptive graph convolutional network dynamically adjusts the weights in the convolution operation based on the characteristics of the input data and contextual information. Through adaptive weight adjustment, the model can weight features according to different situations, better capturing and emphasizing key spatiotemporal features. The adaptive graph convolutional network can also adapt to different environments and scenarios, demonstrating robustness to factors such as lighting changes, background noise, and occlusion. This enables the fall detection model to maintain good performance in a variety of practical application scenarios.
[0018] Preferably, the specific steps of training the yolov5 target person detection model in step S2 are:
[0019] S21: Using a video frame extraction method, the image data collected in step S1 is subjected to frame extraction processing to generate a fall data set;
[0020] S22: performing enhancement processing on the fall dataset to obtain an enhanced dataset, and dividing it into a training set and a test set;
[0021] S23: Build the YOLOV5 algorithm model, input data, and train it to obtain the algorithm model weights, thereby obtaining the YOLOV5 target person detection model. To improve the accuracy of the model's detection, a large number of videos of real-life people falling, as well as simulated videos of people falling, are collected and processed using frame extraction to generate a fall dataset. Common data augmentation techniques, such as Mosaic data augmentation and adaptive anchor box calculation, are used to enhance the dataset and improve the model's generalization capabilities.
[0022] Preferably, the specific steps of step S3 are:
[0023] S31: Use Kalman filtering to predict the position of the next frame of the target person's image and then perform matching;
[0024] S32: Using the Hungarian algorithm to perform data association, that is, inputting the video into the detection network to obtain the location information of the target person, and then transmitting it to the tracking network of the Deepsort target tracking algorithm for data association and matching the target person in the previous and next frames of the target person, so as to obtain the tracking target;
[0025] S33: Then use Kalman filtering to update the trajectory to determine the tracking result and determine the ID.
[0026] Preferably, the similarity calculation in step S32 utilizes the motion information and appearance information of the target person, and the Mahalanobis distance is used to determine the correlation between the predicted target person and the detected target person for the motion information. The formula of the Mahalanobis distance is:
[0027] Among them, d j is the position of the detection box j; y i is the predicted position of tracker i; d j -y i It means that the feature vector of the jth data point is subtracted from the feature vector of the i-th data point to obtain a difference vector, which represents the difference or distance between the two data points in each feature dimension; (d j -y i ) T Indicates that the difference vector is transposed, that is, from a row vector to a column vector, to facilitate matrix operations; S i is the covariance matrix of the detected and predicted positions; the Mahalanobis distance takes the uncertainty of the state measurement into account by calculating the standard deviation between the detected position and the average tracking position;
[0028] When the target is occluded for a long time or the viewing angle is jittery, appearance information is introduced and the cosine distance is used to solve the problem of identity switching caused by occlusion. The expression of cosine distance is:
[0029] Among them, r j is the detection box d j The eigenvectors of ; is the set of eigenvectors of the nearest N frames (N is set to 100) corresponding to tracker i; R i It is the appearance feature vector library;
[0030] Then use linear weighted summation, the formula is: c i,j =λd (1) (i, j)+(1-λ)d (2) (i, j);
[0031] Among them, λ is the weight parameter. This formula combines the Mahalanobis distance and the cosine distance and balances the contribution of the two by adjusting the weight coefficient λ. The value range is between [0, 1]. Function c i,j It is used to represent the comprehensive distance or similarity between the i-th sample and the j-th sample. By adjusting the weight parameter λ, the importance of the two distance metrics can be balanced when combining the distance metrics. i,j Exists in (1)(i, j) and d (2) (i, j) is considered to be related to the target. The “closest distance” to tracker i refers to the N frames with the highest similarity to the target features of the current frame. It is necessary to calculate the cosine similarity between the target features of the current frame and the target features of each previous frame, and select the N frames with the highest similarity; that is, use the trained model to extract features from the jth detection frame and the i-th tracking frame to obtain the apparent feature vector r j and the apparent eigenvector r i , the kth apparent feature vector of the i-th tracking box Stored in the appearance feature library R i middle, is the cosine similarity of the apparent features between the jth detection frame and the ith tracking frame, and the minimum cosine distance d between this value and the apparent features (2) The sum of (i, j) is 1; where d (2) The smaller (i, j) is, the higher the similarity of the apparent features between the tracking box and the detection box is, and the greater the correlation matching degree between the two.
[0032] Preferably, the specific steps of step S4 are:
[0033] S41 Alphapose pose recognition: Use the Alphapose model's top-down approach to perform skeleton detection, obtain continuous skeleton frames, and then use the SPPE (single person pose estimation) algorithm to estimate the pose of the detected target person to obtain the target person's skeleton image;
[0034] S42 Fall Behavior Recognition: A spatiotemporal adaptive graph convolutional network is used to extract features in both spatial and temporal dimensions to obtain the feature vector of the target person.
[0035] S43 Human behavior classification and recognition: Use the trained classifier to classify and recognize the feature vector of the target person. The classifier will determine whether the input data belongs to the fall category based on the feature vector.
[0036] Preferably, in step S41, a symmetric space transformation network (SSTN), a pose-guided proposals generator (PGPG), and a parametric pose non-maximum suppression (PPNMS) are added to the Alphapose model to implement skeleton detection; the specific steps are:
[0037] S411 Data Preprocessing: Preprocess the cropped image segments of the target person, including image scaling, normalization, and channel order conversion, to meet the input requirements of the Alphapose model;
[0038] S412 Multi-Person Pose Estimation: Use Alphapose to perform multi-person pose estimation on cropped pedestrian image segments. The Alphapose model detects the positions of key human body points (such as head, shoulders, elbows, knees, etc.) and estimates the pose information of the target person.
[0039] S413 Result Visualization: Combines the multi-person pose estimation results with the original image or video frame, plotting the connecting lines of key points and pose angle information to obtain a skeleton diagram of the target person for subsequent analysis and presentation. This method, combining Alphapose with graph convolution, effectively avoids dependence on the video environment. This model was compared with multiple models on public datasets, demonstrating high detection accuracy and low scene dependency.
[0040] Preferably, the spatiotemporal adaptive graph convolutional network (STA-GCN) in step S42 is composed of spatial adaptive graph convolution SA-GC and temporal adaptive graph convolution TA-GC, and both the spatial adaptive graph convolution SA-GC and the temporal adaptive graph convolution TA-GC have a topological adaptive encoder TAE with embedded graph convolution.
[0041] Preferably, the specific steps of extracting features from the pedestrian skeleton graph using the spatiotemporal adaptive graph convolutional network through the graph convolution layer in step S42 are:
[0042] S421: A feature extraction operator will be introduced on the spatiotemporal graph. First, for the spatiotemporal graph Graph convolution in , where is the set of all nodes across T frames in the skeleton sequence, ε st is the space-time edge set; v nt For the space-time diagram A node in nt represents a node in the skeleton sequence, where n is the index of the node and t is the index of the time step; therefore, v nt is a node at time step t; v nt It may represent a position or feature in the skeleton sequence at a specific moment, i.e., time step t; n is the index of the node, which is used to uniquely identify different nodes in the spatiotemporal graph; in v ntIn the above equation, n represents the node number, which is an integer usually from 1 to N, where N is the total number of nodes; t usually represents the time step or time frame; it is used to represent each moment or time point in the data set; T represents the total number of time steps or time frames;
[0043] S422: Decompose a spatiotemporal graph into T spatial graphs across time and N temporal graphs across nodes; the spatial graph is represented as where ε s is a set of spatial edges, represented by the spatial adjacency matrix A s ∈R T×N×N ; Where R represents the real number domain of the matrix, R T×N×N represents a three-dimensional matrix whose elements belong to R, i.e., elements in the real field; when all spatial graphs have the same spatial correlation, the spatial adjacency matrix is reduced to A∈R N×N form; similarly, the time diagram is represented as Representation time diagram The temporal edge set in the ; temporal edge set is used to define the temporal relationship or temporal connection between nodes; specifically, ε t Contains the time edges connecting the nodes, which are used to represent the association or dependency between nodes at different time steps; the time adjacency matrix is A t ∈R N×T×T ;After the spatiotemporal graph decomposition, spatial graph convolution S-GC and temporal graph convolution T-GC were developed;
[0044] Spatial graph convolution S-GC is expressed as:
[0045] in, It's A s The element v pt Represents a node in the spatiotemporal graph, where p is the index of the node and t is the index of the time step. In spatial graph convolution S-GC, it is used to represent nodes in the spatial graph; W represents the weight matrix or weight parameter, which is usually used for linear transformation or weighted aggregation of features; that is, w(v pt ) represents the node v pt The weights applied to perform a weighted aggregation of the input features to generate the node v nt Output features on ;
[0046] Temporal graph convolution T-GC is expressed as:
[0047] in, It's A t The element v nqIt also represents a node in the spatiotemporal graph, where n is the index of the node and q is the index of the time step. In the temporal graph convolution T-GC, it is used to represent the node in the time graph; W represents the weight matrix (or weight parameter), which is usually used for linear transformation or weighted aggregation of features; that is, w(v nq ) represents the node v nq The weights applied to perform a weighted aggregation of the input features to generate the node v nt Output features on ;
[0048] Then use the spatial graph convolution S-GC and the temporal graph convolution T-GC to extract the features on the spatiotemporal graph and obtain the feature vector, where the extracted features are expressed as:
[0049] in, It's A t Elements of Represents the spatial graph convolution used to connect nodes v in S-GC pq and node v nq The weights or connection coefficients between them are the spatial adjacency matrix A in the space-time graph. s The elements of are used to measure the spatial relationship or connection strength between nodes; specifically, Represents node v pq and node v nq The spatial connection strength between them is used to determine how to propagate feature information in the spatial graph; N is used to represent the index range of the node, indicating that the node v pq and node v nq The node index in the graph; q represents the index of the time step or time frame, which is used to represent different time steps or time frames in the temporal graph convolution T-GC; p is the index of the node of the spatial graph; v pq Represents the nodes in the spatial graph convolution S-GC, where p is the index of the node and q is the index of the time step; W represents the weight matrix or weight parameter, which is used for linear transformation or weighted aggregation of features, i.e. w1(v pq ) and w2(v nq ) are used to perform weighted aggregation on the input features, w1(v pq )w1(v pq ) represents the node v pq The first weight or connection coefficient applied, w2(v nq ) represents the node v nq The second weight or connection coefficient applied; these weight matrices are typically learnable parameters in neural network models. They are used to adjust and control how features are propagated and aggregated between different nodes to capture complex feature relationships in spatiotemporal graphs; during the training process of the model, these weight matrices are adjusted according to the loss function to optimize the performance of the model.
[0050] This method uses the YOLO v5 network model to obtain human target detection frame information, confidence information and classification information; the target tracking part of the model adopts the YOLO v5 combined with Deepsort method, in which each tracking instance needs to be matched with the target detection result obtained by YOLO v5. It can be divided into three states according to different situations. It is set to a temporary state during initialization. If no detection result is matched, it will be deleted; if it is matched continuously for a certain number of times, it will be set to a tracked state; if the number of unmatched times in the tracked state exceeds the set maximum number, it will be deleted; in the posture estimation part, the target detection frame of each frame obtained by the target detection and tracking part is passed in. In the posture estimation stage, the human skeleton key points of each frame image are identified and detected in the model through Alphapose, and then the obtained skeleton key point data is Normalized data processing is performed; after posture estimation, the information of the key points of the human skeleton is obtained, and the obtained skeleton points are passed into the improved spatiotemporal adaptive graph convolutional neural network (STA-GCN) in the form of graph information for graph convolution. The tensor after graph convolution is passed to the Linear layer for classification to realize human fall detection; spatiotemporal adaptive graph convolution is introduced to model the spatiotemporal relationship of pedestrian targets, and the pedestrian's motion trajectory information and key point position information are used to realize adaptive learning and extraction of features; spatiotemporal adaptive graph convolution can better capture the pedestrian's motion characteristics and posture changes, enhancing the feature representation ability.
[0051] Preferably, the specific steps of step S43 are:
[0052] S431: First, preprocess the feature vector to obtain a preprocessed feature vector;
[0053] S432: Select a classifier model and train the classifier using the labeled training data;
[0054] S433: Evaluate the trained classifier using unlabeled test data;
[0055] S434: Using the trained classifier to classify and identify the preprocessed feature vector, and judging whether the input data belongs to the fall category according to the feature vector.
[0056] Preferably, the preprocessing method in step S431 includes feature normalization and dimensionality reduction; the classifier model in step S432 selects support vector machine SVM; and the performance of the classifier is evaluated by calculating accuracy, recall rate and F1 score in step S433.
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] (1) High Accuracy: YOLOv5 has high accuracy as a target detector and can quickly and accurately detect pedestrian targets in images or videos. Deepsort can achieve continuous tracking of pedestrian targets, thereby reducing false detections and missed detections and improving the accuracy of fall detection.
[0059] (2) Real-time performance: YOLOv5 and Deepsort are both target detection and tracking algorithms designed for real-time applications. Therefore, the entire system can run efficiently in real-time scenarios and meet the needs of real-time fall detection.
[0060] (3) Continuous posture estimation: Combined with Deepsort’s continuous tracking of pedestrian targets, fall detection can continuously estimate the pedestrian’s posture. It can not only accurately determine whether the pedestrian has fallen, but also obtain information about the pedestrian’s posture changes throughout the entire video sequence.
[0061] (4) Spatiotemporal Adaptive Graph Convolution: The introduction of spatiotemporal adaptive graph convolution can capture the motion characteristics and posture changes of pedestrian targets and enhance the representation ability of features. By learning the relationship between pedestrian targets in time and space, the feature representation of fall detection is more comprehensive and meaningful;
[0062] (5) Multi-task learning: YOLOv5, Deepsort, and pose estimation modules are integrated to form a multi-task learning framework. The modules collaborate with each other and share features, which improves the performance of the entire system.
[0063] (6) Universality: The fall detection method based on YOLOv5 and Deepsort can be applied to various types of cameras and scenes, and has strong versatility and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0065] FIG1 is a flow chart of the skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution of the present application;
[0066] FIG2 is a flowchart of the overall pedestrian detection and tracking process in the skeleton detection and fall detection method based on the improved spatiotemporal adaptive graph convolution of the present application;
[0067] FIG3 is a flowchart of the cascade matching in the Deepsort algorithm in the skeleton detection and fall detection method based on the improved spatiotemporal adaptive graph convolution of the present application;
[0068] FIG4 is a diagram of the spatiotemporal adaptive graph convolution network structure in the skeleton detection and fall detection method based on the improved spatiotemporal adaptive graph convolution of the present application;
[0069] FIG5 is a schematic diagram of skeleton detection in the skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution of the present application. DETAILED DESCRIPTION
[0070] The embodiments of the present application are described in detail below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application and are not intended to limit the scope of protection of the present application.
[0071] The common knowledge in this technical solution specifically includes:
[0072] (1) YOLOv5 target detection algorithm:
[0073] YOLOv5 technical features: YOLOv5 is a single-stage object detection algorithm that adds some new improvements to YOLOv4, significantly improving its speed and accuracy. Compared to YOLOv4, the YOLOv5 object detection algorithm has the following improvements:
[0074] Input: During the model training phase, some improvement ideas were proposed, mainly including mosaic data enhancement and adaptive anchor box calculation.
[0075] Mosaic data augmentation randomly selects four different training images and stitches them together into a single large image, creating a "mosaic" image. During the stitching process, the size and position of each image are randomly adjusted and flipped. This allows the model to be exposed to more background and target samples during training, increasing the richness and diversity of the dataset. Mosaic data augmentation can improve the model's generalization and robustness, mitigating overfitting.
[0076] Adaptive anchor box calculation is a method that dynamically adjusts anchor boxes based on the distribution and size of objects in an image. Traditional object detection algorithms typically use fixed anchor boxes to predict the object's bounding box. Adaptive anchor box calculation, however, dynamically generates anchor boxes based on the distribution of objects in the training set, making them more adaptable to the object's size and aspect ratio. This improves the accuracy of object detection algorithms for objects of varying sizes and shapes. Adaptive anchor box calculation is often combined with clustering algorithms or statistical analysis methods to determine the optimal anchor box size and aspect ratio.
[0077] These techniques are widely used in object detection algorithms to increase data diversity and improve model adaptability and accuracy. Their use can make the model better adapted to various scenarios and targets, and achieve better performance in practical applications.
[0078] Baseline network: Integrates some new ideas from other detection algorithms, mainly including: Focus structure and CSP structure.
[0079] The Focus architecture is a special structure in YOLOv5 that replaces traditional convolution operations. Traditional convolution operations typically significantly reduce the scale of the input feature map, resulting in information loss and increased computational effort. The Focus architecture, however, employs a concept similar to grouped convolution, dividing the input feature map into smaller sub-maps and performing convolution operations within them. This reduces information loss and improves computational efficiency. The Focus architecture helps maintain the richness of feature maps while reducing computational effort.
[0080] The CSP architecture is a specialized network structure designed to improve feature map transmission and extraction. The CSP architecture splits the input feature map into two branches, one of which performs convolution operations while the other directly transmits the original feature map. This branched structure promotes feature propagation and fusion, improving the model's ability to perceive objects of varying scales. In YOLOv5, the CSP architecture is applied at different network levels to enhance feature extraction and representation.
[0081] The application of YOLOv5's Focus and CSP architectures enables the model to better handle objects of different scales, improving the accuracy and robustness of object detection. By integrating new ideas from other detection algorithms, YOLOv5's baseline network achieves excellent performance in object detection tasks.
[0082] Neck network: The target detection network often inserts some layers between the BackBone and the final Head output layer. Yolov5 adds the FPN+PAN structure, which is often used in target detection tasks to extract and fuse feature information of different scales to better detect multi-scale targets.
[0083] The Feature Pyramid Network (FPN) is a network structure designed to process multi-scale features. It builds feature pyramids at different levels of the backbone network, generating a series of feature maps at different scales. These feature maps contain rich semantic and spatial information, capturing the details and context of objects at different scales. The FPN architecture fuses these feature maps using upsampling and downsampling operations to generate a series of feature pyramids with different resolutions, meeting the detection requirements of objects of different scales.
[0084] The PAN (Path Aggregation Network) architecture is a further improvement on the FPN architecture. It introduces lateral connections and top-down pathways to better fuse and aggregate feature information. Lateral connections are used to combine high-resolution shallow features with low-resolution deep features to enrich and enhance feature representations. The top-down pathway combines lower-resolution feature maps with high-resolution feature maps through upsampling operations to obtain richer contextual information.
[0085] Head output layer: The anchor box mechanism of the output layer is the same as YOLOv4. The main improvements are the loss function GIOU_Loss during training and the prediction box screening DIOU_nms.
[0086] The head output layer of YOLOv5 uses a set of predefined anchor boxes to predict the bounding box of the object. Each anchor box is typically associated with multiple grid cells, and the object's location and size are predicted through regression. The size and aspect ratio of these anchor boxes are typically designed and adjusted based on the distribution of objects in the training set to accommodate objects of varying scales and shapes.
[0087] During training, YOLOv5 improves the loss function calculation method and adopts GIOU_Loss (Generalized Intersection over Union Loss). GIOU_Loss is an improved version of the IoU loss function, which measures the degree of overlap between the predicted bounding box and the ground-truth bounding box. It considers the area of the intersection and union between bounding boxes and optimizes the model by calculating a generalized IoU value. Compared to the traditional IoU loss function, GIOU_Loss can better handle bounding box overlap when optimizing object detection tasks, improving detection accuracy and robustness.
[0088] YOLOv5 also improves the prediction box screening method by introducing DIOU_nms (Distance-IoU Non-Maximum Suppression). DIOU_nms considers the distance and IoU values between bounding boxes during non-maximum suppression (NMS) to select the most representative bounding boxes. By comprehensively considering the position and shape of the bounding boxes, it suppresses redundant predictions while retaining the most relevant target boxes. This screening method can further improve the accuracy and recall of object detection.
[0089] (2) Deepsort target tracking algorithm: The predecessor of Deepsort is the SORT algorithm. The SORT algorithm is composed of a target detector and a tracker. The core of its tracker is the Kalman filter algorithm and the Hungarian algorithm. The Kalman filter algorithm is used to predict the state of the detection frame in the next frame, and the state is matched with the detection result of the next frame using the Hungarian algorithm to achieve tracking. Once the object is blocked or not detected for other reasons, the state information predicted by the Kalman filter will not be matched with the detection result, and the tracking segment will end early. Deepsort introduces the re-identification algorithm in deep learning to extract the appearance features (low-dimensional vector representation) of the detected object (in the detection frame object). After each (each frame) detection + tracking, the object appearance features are extracted and saved. In each subsequent step, the similarity calculation between the appearance features of the detected object in the current frame and the previously stored appearance features is performed to avoid missed detection and loss of identity ID. It can be said that Deepsort not only uses the speed and direction trend of the object to track the target, but also uses the appearance features of the object to consolidate the judgment of whether it is the same object.
[0090] Deepsort main features:
[0091] 1. Multi-target tracking: Deepsort can track multiple targets simultaneously. By associating and managing the trajectories of each target, it can accurately track multiple targets in the video.
[0092] 2. Combining deep learning and traditional methods: Deepsort combines deep learning with traditional Kalman filters. Deep learning is used for target detection and feature extraction, providing rich target feature representations; Kalman filters are used for target state estimation and prediction, providing target position and motion information.
[0093] 3. Powerful feature representation: The feature representation extracted by the deep learning network has high discriminability and can accurately measure the similarity between objects. This helps to accurately associate objects and handle problems such as appearance changes and occlusion.
[0094] 4. Robustness: Deepsort is robust enough to handle common challenges such as the appearance, disappearance, and occlusion of targets. Using Kalman filtering for state prediction and trajectory management can compensate for errors or instabilities in target detection to a certain extent.
[0095] 5. Real-time performance: Deepsort performs well in real-time target tracking, with fast speed and low computing resource requirements. This makes it suitable for real-time target tracking on devices with limited computing resources, such as mobile devices and embedded systems.
[0096] Example: A skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution, as shown in FIG1 , specifically includes the following steps:
[0097] S1: Collect image data and obtain each frame of image data through video or rtsp stream.
[0098] S2: Use the pre-trained yolov5 target person detection model to detect whether the target person appears in each frame of image data. If the target person appears, go to step S3; if not, end.
[0099] The specific steps of training the yolov5 target person detection model in step S2 are:
[0100] S21: Using a video frame extraction method, the image data collected in step S1 is subjected to frame extraction processing to generate a fall data set;
[0101] S22: performing enhancement processing on the fall dataset to obtain an enhanced dataset, and dividing it into a training set and a test set;
[0102] S23: Build the YOLOV5 algorithm model, input data and train it, obtain the algorithm model weight, and obtain the YOLOV5 target person detection model.
[0103] S3: For each detected target person, use the Deepsort target tracking algorithm to track the target, obtain the tracking result, calculate the similarity to obtain the target association result, and update the trajectory information of each target person.
[0104] As shown in FIG2 , the specific steps of step S3 are:
[0105] S31: Use Kalman filtering to predict the position of the next frame of the target person's image, and then perform matching; including cascade matching and IOU matching.
[0106] As shown in Figure 3, the Kalman filter is first used to predict the position of the next frame image of the target person, and cascade matching is performed for the tracked target person, and IOU matching is performed for the untracked trajectory; after cascade matching, unmatched tracking instances, unmatched detection instances and matched tracking instances are obtained, and IOU matching is performed again for the unmatched detection instances; after IOU matching, unmatched tracking instances, unmatched detection instances and matched tracking instances are obtained, wherein, for the matched tracking instances, step S33 is performed, for the unmatched tracking instances, if there is no tracked target person's trajectory in the unmatched tracking instances, it is deleted; if there is a tracked target person's trajectory in the unmatched tracking instances, it is judged that the number of unmatched times exceeds the set maximum number, if it is greater, it is deleted, if it is less, it is turned to the prediction of the target person using Kalman filtering.
[0107] S32: Use the Hungarian algorithm to perform data association, that is, input the video into the detection network to obtain the location information of the target person, and then transmit it to the tracking network of the Deepsort target tracking algorithm for data association and match the target persons in the previous and next frames of the target person to obtain the tracking target.
[0108] The similarity calculation in step S32 utilizes the motion information and appearance information of the target person. The Mahalanobis distance is used to determine the correlation between the predicted target person and the detected target person for the motion information. The formula of the Mahalanobis distance is:
[0109] Among them, dd j is the position of the detection box j; y i is the predicted position of tracker i; d j -y i It means that the feature vector of the jth data point is subtracted from the feature vector of the i-th data point to obtain a difference vector, which represents the difference or distance between the two data points in each feature dimension; (d j -y i ) T Indicates that the difference vector is transposed, that is, from a row vector to a column vector, to facilitate matrix operations; S i is the covariance matrix of the detected and predicted positions; the Mahalanobis distance takes the uncertainty of the state measurement into account by calculating the standard deviation between the detected position and the average tracking position;
[0110] When the target is occluded for a long time or the viewing angle is jittery, appearance information is introduced and the cosine distance is used to solve the problem of identity switching caused by occlusion. The expression of cosine distance is:
[0111] Among them, r j is the detection box dj The eigenvectors of ; is the set of the nearest N frames of eigenvectors corresponding to tracker i; in this embodiment, N is set to 100; R i is the appearance feature vector library; in order to make full use of motion information and appearance information, a linear weighted summation method is used, and the formula is: i,j =λd (1) (i, j)+(1-λ)d (2) (i, j);
[0112] Among them, λ is the weight parameter. This formula combines the Mahalanobis distance and the cosine distance and balances the contribution of the two by adjusting the weight coefficient λ. The value range is between [0, 1]. Function c i,j It is used to represent the comprehensive distance or similarity between the i-th sample and the j-th sample. By adjusting the weight parameter λ, the importance of the two distance metrics can be balanced when combining the distance metrics. i,j Exists in (1) (i, j) and d (2) (i, j) is considered to be related to the target; The “closest distance” to tracker i refers to the N frames with the highest similarity to the target features of the current frame. It is necessary to calculate the cosine similarity between the target features of the current frame and the target features of each previous frame, and select the N frames with the highest similarity; that is, use the trained model to extract features from the j-th detection frame and the i-th tracking frame to obtain the apparent feature vector r j and the apparent eigenvector r i , the kth apparent feature vector of the i-th tracking box Stored in the appearance feature library R i middle, is the cosine similarity of the apparent features between the jth detection frame and the i-th tracking frame, and the minimum cosine distance dd between the apparent features (2) The sum of (i, j) is 1; where d (2) The smaller (i, j) is, the higher the similarity of the apparent features between the tracking box and the detection box is, and the greater the correlation matching degree between the two.
[0113] S33: A Kalman filter is then used to update the trajectory, confirm the tracking result, and determine the ID. Finally, the Hungarian algorithm is used for data association to improve tracking performance and reduce identity switching issues. The video is input to the detection network to obtain the target person's location information, which is then transferred to the tracking network for data association and matching of people in previous and subsequent frames to obtain the tracking result. Deepsort calculates similarity using the target's motion information. Based on the target association results, the trajectory information of each existing pedestrian target is updated. The Kalman filter can be used to predict and update the target state, estimating the position and velocity of the pedestrian target based on the covariance matrix between the prediction and observation. For new pedestrian targets that have not been associated, a new target trajectory is created and assigned a unique identifier. The trajectory of these new targets will be updated and tracked in subsequent frames. Through the above steps, Deepsort can track pedestrian targets in real time and maintain their trajectory information. In fall detection, the pedestrian tracking results can be combined with subsequent fall event detection. For example, by analyzing information such as the pedestrian's posture and motion trajectory, a fall event can be determined.
[0114] S4: Perform posture recognition on each target person based on trajectory information, and use a spatiotemporal adaptive graph convolutional network to extract the feature vector of the posture. Use a classifier to classify and identify human behavior to determine whether the target person has fallen.
[0115] The specific steps of step S4 are:
[0116] S41 Alphapose Pose Recognition: Skeleton detection is performed using the Alphapose model in a top-down manner to obtain continuous skeleton frames. The SPPE (single person pose estimation) algorithm is then used to estimate the pose of the detected target person, resulting in a skeleton map of the target person. The effect of combining target detection and pose estimation is shown in Figure 5. It can be seen that the target person's detection box is obtained after target detection, and the human skeleton map is obtained after pose estimation. Figure 5 uses target detection followed by pose estimation, which is a common method for human behavior recognition. Its pose estimation method can provide more human information, such as joint positions and pose angles. As can be seen from the pose estimation map in Figure 5, this information can be used to more comprehensively analyze the human state, thereby improving the accuracy of fall events. In addition, the pose estimation method can capture the dynamic changes of the human body, including motion trajectory and posture transitions. This is very important for distinguishing falls from other normal behaviors, as falls are often accompanied by unusual movements. By tracking the temporal movement trajectory of key points of the human body, the relationship and continuous changes between human parts can be obtained. During a fall, the trajectory of key points may show abnormal patterns, which helps to identify fall events.
[0117] In the step S41, a symmetric space transformation network (SSTN), a pose-guided proposal generator (PGPG) and a parametric pose non-maximum suppression (PPNMS) are added to the Alphapose model to realize skeleton detection; the existing skeleton model has two main problems: positioning error and redundant detection results. To address these problems, the Alphapose model adds three modules: a symmetric space transformation network (SSTN), a pose-guided proposal generator (PGPG) and a parametric pose non-maximum suppression (PPNMS). The role of SSTN is to extract the region of interest of the human body detection frame to achieve the purpose of automatically adjusting the detection frame. By adding this module, the target detection result in the above figure is more accurate. PGPG performs posture-guided data augmentation on existing data, increasing training samples and achieving data enhancement. It is used for target detection and SPPE training. PPNMS is a parameterized posture non-maximum suppression method that eliminates redundant detection boxes by defining posture distance and calculating posture similarity. When the similarity falls below a certain threshold, the redundant boxes are deleted. By combining these three modules, Alphapose achieves more accurate skeleton detection. The specific steps are:
[0118] S411 Data Preprocessing: Preprocess the cropped image segments of the target person, including image scaling, normalization, and channel order conversion, to meet the input requirements of the Alphapose model;
[0119] S412 Multi-Person Pose Estimation: Use Alphapose to perform multi-person pose estimation on cropped pedestrian image segments. The Alphapose model detects the positions of key human body points (such as head, shoulders, elbows, knees, etc.) and estimates the pose information of the target person.
[0120] S413 Result Visualization: Combine the results of multi-person pose estimation with the original image or video frame, draw the connection lines of key points and pose angle information, and obtain the skeleton diagram of the target person for subsequent analysis and display.
[0121] S42 Fall behavior recognition: A spatiotemporal adaptive graph convolutional network is used to extract features in both spatial and temporal dimensions to obtain the feature vector of the target person.
[0122] In real videos, action behaviors do not only include single-frame situations, but also continuous time. Therefore, after extracting the skeleton using the Alphapose algorithm, continuous skeleton frames are obtained. The human skeleton frame sequence is composed of joint point coordinates. The early work of extracting skeleton features was mainly handed over to the spatiotemporal convolutional neural network, which connects the coordinate vectors of the joints of the entire frame and then forms a feature vector. Although the convolutional neural network can extract spatiotemporal features, this method does not take into account the natural connection of skeleton edges between skeleton joints and the temporal edges connecting the same joint between frames. In 2018, researchers proposed the ST-GCN model, which uses a graph convolutional neural network (GCN) to process skeleton graphs. However, the feature extraction operation on the spatiotemporal graph still has two shortcomings: (1) In the spatial feature learning stage, the fixed spatial topology is shared between all postures, which may not be optimal for actions with large posture changes. Using a fixed spatial topology may mistakenly enhance irrelevant connections or weaken key connections, and cannot accurately represent spatial dependencies. This fact indicates that the spatial topology structure should adapt to each posture in the skeleton sequence. (2) In the stage of learning temporal features, existing methods apply temporal convolution with a fixed small kernel to extract short-range temporal features. This results in a weak ability to model temporal long-range joint dependencies, which is crucial for action recognition. To learn robust feature representations in both spatial and temporal dimensions, this proposal proposes an improved ST-GCN network, the Spatio-Temporal Adaptive Graph Convolutional Network (STA-GCN). The STA-GCN in step S42 consists of a spatially adaptive graph convolution (SA-GC) and a temporally adaptive graph convolution (TA-GC). Both the SA-GC and TA-GC have a topology-adaptive encoder (TAE) embedded in a graph convolution. The SA-GC module extracts spatial features by modeling spatially adaptive joint dependencies. The TA-GC module aims to learn temporal features by capturing direct long-range joint dependencies in the temporal dimension. This model, combined with the SA-GC and TA-GC modules, can learn discriminative features in both spatial and temporal dimensions. The TAE component is used to learn spatially and temporally adaptive topologies. Existing GCN methods use a fixed spatio-temporal topology. This fixed spatial topology forces each skeleton frame to adopt the same spatial topology, while the fixed temporal topology forces all trajectories to use the same temporal topology. This fixed topology is insufficient to represent the joint dependencies for each pose or trajectory. Therefore, we propose TAE to address this problem by learning spatially adaptive topology and temporally adaptive topology. The spatially adaptive topology can generate pose-specific dependencies for each frame in the skeleton sequence to learn discriminative spatial features. The temporally adaptive topology can model the direct long-range dependencies between any two nodes in the trajectory graph to extract robust temporal features.
[0123] As shown in FIG4 , the specific steps of extracting features from the pedestrian skeleton graph using the spatiotemporal adaptive graph convolutional network through the graph convolution layer in step S42 are as follows:
[0124] S421: Most existing studies regard the human skeleton sequence as a spatiotemporal graph and extract features through spatial graph convolution and temporal convolution. However, due to the use of a fixed small temporal convolution kernel, the ability of temporal convolution to learn temporal features is not strong. Therefore, a more robust feature extraction operator is introduced on the spatiotemporal graph. First, for the spatiotemporal graph, Graph convolution in , where is the set of all nodes across frames in the skeleton sequence, ε st is the space-time edge set; v nt For the space-time diagram A node in nt represents a node in the skeleton sequence, where n is the index of the node and t is the index of the time step; therefore, v nt is a node at time step t; v nt It may represent a position or feature in the skeleton sequence at a specific moment (time step t); n is the index of the node, which is used to uniquely identify different nodes in the spatiotemporal graph; in v nt In it, n represents the node number, which is an integer, usually from 1 to N, where N is the total number of nodes; t usually represents the time step or time frame; it is used to represent each moment or time point in the data set; T represents the total number of time steps or time frames.
[0125] S422: Decompose a spatiotemporal graph into T spatial graphs across time and N temporal graphs across nodes; the spatial graph is represented as where ε s is a set of spatial edges, represented by the spatial adjacency matrix A s ∈R T×N×N ; Where R represents the real number domain of the matrix, R T×N×N represents a three-dimensional matrix whose elements belong to R, i.e., elements in the real field; when all spatial graphs have the same spatial correlation, the spatial adjacency matrix is reduced to A∈R N×N form; similarly, the time diagram is represented as ε t Representation time diagram The temporal edge set in the ; temporal edge set is used to define the temporal relationship or temporal connection between nodes; specifically, ε t Contains the time edges connecting the nodes, which are used to represent the association or dependency between nodes at different time steps; the time adjacency matrix is A t ∈R N×T×T;After the spatiotemporal graph decomposition, spatial graph convolution S-GC and temporal graph convolution T-GC were developed;
[0126] Spatial graph convolution S-GC is expressed as:
[0127] in, It's A s The element v pt Represents a node in the spatiotemporal graph, where p is the index of the node and t is the index of the time step. In spatial graph convolution S-GC, it is used to represent nodes in the spatial graph; W represents the weight matrix or weight parameter, which is usually used for linear transformation or weighted aggregation of features; w(v pt ) represents the node v pt The weights applied to perform a weighted aggregation of the input features to generate the node v nt Output features on ;
[0128] Temporal graph convolution T-GC is expressed as:
[0129] in, It's A t The element v nq Represents a node in the spatiotemporal graph, where n is the index of the node and q is the index of the time step. In the temporal graph convolution T-GC, it is used to represent the node in the temporal graph; W represents the weight matrix (or weight parameter), which is usually used for linear transformation or weighted aggregation of features; that is, w(v nq ) represents the node v nq The weights applied to perform a weighted aggregation of the input features to generate the node v nt Output features on ;
[0130] Then use the spatial graph convolution S-GC and the temporal graph convolution T-GC to extract the features on the spatiotemporal graph and obtain the feature vector, where the extracted features are expressed as:
[0131] in, It's A t Elements of Represents the spatial graph convolution used to connect nodes v in S-GC pq and node v nq The weights or connection coefficients between them are the spatial adjacency matrix A in the space-time graph. s The elements of are used to measure the spatial relationship or connection strength between nodes; specifically, Represents node v pq and node v nqThe spatial connection strength between them is used to determine how to propagate feature information in the spatial graph; N is used to represent the index range of the node, indicating that the node v pq and node v nq The node index in the space-time graph; q represents the index of the time step or time frame, which is used to represent different time steps or time frames in the time graph convolution T-GC; specifically, q represents a time point in the time dimension of the space-time graph; v pq Represents the nodes in the spatial graph convolution S-GC, where p is the index of the node and q is the index of the time step; W represents the weight matrix or weight parameter, which is used for linear transformation or weighted aggregation of features, i.e. w1(v pq ) and w2(v nq ) are used to perform weighted aggregation on the input features, w1(v pq ) represents the node v pq The first weight or connection coefficient applied, w2(v nq ) represents the node v nq The second weight or connection coefficient applied; these weight matrices are typically learnable parameters in neural network models. They are used to adjust and control how features are propagated and aggregated between different nodes to capture complex feature relationships in spatiotemporal graphs; during the training process of the model, these weight matrices are adjusted according to the loss function to optimize the performance of the model.
[0132] S43 Human behavior classification and recognition: Use the trained classifier to classify and recognize the feature vector of the target person. The classifier will determine whether the input data belongs to the fall category based on the feature vector. The specific steps of step S43 are:
[0133] S431: First, preprocess the feature vector to obtain a preprocessed feature vector;
[0134] The preprocessing methods in step S431 include feature normalization and dimensionality reduction; these preprocessing steps help optimize the performance of the classifier;
[0135] S432: Select a classifier model and train it using the labeled training data. The classifier model in step S432 is selected as a support vector machine (SVM). The classifier is trained using the labeled training data. The training data is a set of labeled feature vectors, where the label represents the fall behavior category corresponding to each feature vector.
[0136] S433: Evaluate the trained classifier using unlabeled test data. In step S433, the performance of the classifier is evaluated by calculating the precision, recall, and F1 score.
[0137] S434: Using the trained classifier to classify and identify the preprocessed feature vector, and judging whether the input data belongs to the fall category according to the feature vector.
[0138] For ordinary technicians in this field, the specific embodiments are only illustrative descriptions of the present application. It is obvious that the specific implementation of the present application is not limited to the above-mentioned methods. As long as various non-substantial improvements are made using the method concepts and technical solutions of the present application, or the concepts and technical solutions of the present application are directly applied to other occasions without improvement, they are all within the scope of protection of the present application.
Claims
1. A skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution, characterized in that: The specific steps include: S1: Collect image data and obtain each frame of image data; S2: Use the pre-trained yolov5 target person detection model to detect whether the target person appears in each frame of image data. If the target person appears, go to step S3, if not, end; S3: For each detected target person, use the Deepsort target tracking algorithm to track the target, obtain the tracking result, calculate the similarity to obtain the target association result, and update the trajectory information of each target person; S4: Perform posture recognition on each target person in combination with trajectory information, and use a spatiotemporal adaptive graph convolutional network to extract the feature vector of the posture. Use a classifier to classify and identify human behavior to determine whether the target person has fallen.
2. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 1, characterized in that: The specific steps of training the yolov5 target person detection model in step S2 are: S21: Using a video frame extraction method, the image data collected in step S1 is subjected to frame extraction processing to generate a fall data set; S22: performing enhancement processing on the fall data set to obtain an enhanced data set, and dividing the enhanced data set into a training set and a test set; S23: Build the YOLOV5 algorithm model, input data and train it, obtain the algorithm model weight, and obtain the YOLOv5 target person detection model.
3. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 2, characterized in that: The specific steps of step S3 are: S31: using Kalman filtering to predict the position of the next frame image of the target person, and then matching; S32: using the Hungarian algorithm to perform data association, that is, inputting the video into the detection network to obtain the location information of the target person, and then transmitting it to the tracking network of the Deepsort target tracking algorithm for data association and matching the target persons in the previous and next frames of the target person, so as to obtain the tracking target; S33: Update the track again to confirm the tracking result and determine the ID.
4. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 3, characterized in that: The similarity calculation in step S32 utilizes the motion information and appearance information of the target person. The Mahalanobis distance is used to determine the correlation between the predicted target person and the detected target person for the motion information. The formula of the Mahalanobis distance is: Among them, d j is the position of the detection box j; y i is the predicted position of tracker i; d j -y i It means that the feature vector of the jth data point is subtracted from the feature vector of the ith data point to obtain a difference vector, which represents the difference or distance between the two data points in each feature dimension; (d j -y i ) T Indicates transposing the difference vector, that is, changing it from a row vector to a column vector, for the convenience of matrix operations; S i is the covariance matrix of the detected and predicted positions; when the target is occluded for a long time or the viewing angle is jittered, the appearance information is introduced, and the cosine distance is used to solve the problem of identity switching caused by occlusion; the expression of the cosine distance is: Among them, r j is the detection box d j The eigenvector of j Represents the feature vector of the jth data point; Yes j The transpose of , which contains the eigenvalue information of the j-th data point in each feature dimension; is the set of the nearest N frames of eigenvectors corresponding to tracker i; R i is the appearance feature vector library; using motion information and appearance information, a linear weighted sum is used, and the formula is: c i,j =λd (1) (i,j)+(1-λ)d (2) (i,j); Among them, λ is the weight parameter, and its value range is between [0, 1]; function c i,j It is used to represent the comprehensive distance or similarity between the i-th sample and the j-th sample. By adjusting the weight parameter λ, the importance of the two distance metrics can be balanced when combining the distance metrics. i,j Exists in (1) (i, j) and d (2) (i, j) is considered to be related.
5. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 3, characterized in that: The specific steps of step S4 are: S41 Alphapose posture recognition: Use the Alphapose model's top-down method to perform skeleton detection to obtain continuous skeleton frames, and then use the SPPE algorithm to estimate the posture of the detected target person to obtain the target person's skeleton map; S42 Falling behavior recognition: A spatiotemporal adaptive graph convolutional network is used to extract features in both spatial and temporal dimensions to obtain the feature vector of the target person; S43 Human behavior classification and recognition: Use the trained classifier to classify and recognize the feature vector of the target person. The classifier will determine whether the input data belongs to the fall category based on the feature vector.
6. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 5, characterized in that: In the step S41, a symmetric space transformation network, a posture-guided sample generator and a posture non-maximum suppressor are added to the Alphapose model to realize skeleton detection; the specific steps are: S411 data preprocessing: preprocessing the cropped image segments of the target person; S412 Multi-person pose estimation: Use Alphapose to estimate the poses of multiple people on the cropped pedestrian image segments; the Alphapose model detects the positions of key points on the human body and estimates the pose information of the target person; S413 Result Visualization: Combine the results of multi-person posture estimation with the original image or video frame, draw the connection lines of key points and the information of posture angles, and obtain the skeleton diagram of the target person for subsequent analysis and display.
7. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 6, characterized in that: The spatiotemporal adaptive graph convolution network in step S42 is composed of a spatial adaptive graph convolution SA-GC and a temporal adaptive graph convolution TA-GC, and both the spatial adaptive graph convolution SA-GC and the temporal adaptive graph convolution TA-GC have a topological adaptive encoder TAE with an embedded graph convolution.
8. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 7, characterized in that: The specific steps of extracting features from the pedestrian skeleton graph by using the spatiotemporal adaptive graph convolutional network through the graph convolution layer in step S42 are: S421: A feature extraction operator is introduced on the spatiotemporal graph. First, The graph convolution in is the set of all nodes across T frames in the skeleton sequence, ε st is the space-time edge set; v nt For the space-time diagram A node in nt represents a node in the skeleton sequence, where n is the index of the node and t is the index of the time step; therefore, v nt is a node at time step t; v nt represents a position or feature in the skeleton sequence at a specific moment, i.e., step t; n is the index of the node, which is used to uniquely identify different nodes in the spatiotemporal graph; in v nt In the above, n represents the node number, which is an integer usually from 1 to N, where N is the total number of nodes; t represents the time step or time frame, which is used to represent each node in the data set. A moment or time point; T represents the total number of time steps or time frames; S422: Decompose a spatiotemporal graph into T spatial graphs across time and N temporal graphs across nodes; the spatial graph is represented as where ε s is a set of spatial edges, represented by the spatial adjacency matrix A s ∈R T×N×N ; Where R represents the real number domain of the matrix, R T×N×N represents a three-dimensional matrix whose elements belong to R, i.e., elements in the real number domain; when all spatial graphs have the same spatial correlation, the spatial adjacency matrix is reduced to A∈R N×N form; similarly, the time diagram is expressed as ε t Representation time diagram The temporal edge set in A is used to define the temporal relationship or temporal connection between nodes. The temporal adjacency matrix is A t ∈R N×T×T ; After the decomposition of the spatiotemporal graph, the spatial graph convolution S-GC and the temporal graph convolution T-GC were developed; the spatial graph convolution S-GC is expressed as: in, Yes A s The element v pt represents a node in the spatiotemporal graph, where p is the index of the node and t is the index of the time step. In the spatial graph convolution S-GC, it is used to represent the nodes in the spatial graph; W represents the weight matrix or weight parameter, which is used for linear transformation or weighted aggregation of features; that is, w(v pt ) represents the node v pt The weights applied to perform a weighted aggregation of the input features to generate the node v nt Output features on ; The temporal graph convolution T-GC is expressed as: in, Yes A t The element v nq represents a node in the spatiotemporal graph, where n is the index of the node and q is the index of the time step. In the temporal graph convolution T-GC, it is used to represent the node in the temporal graph; W represents the weight matrix or weight parameter, which is used for linear transformation or weighted aggregation of features; that is, w(v nq ) represents the node v nq The weights applied to perform a weighted aggregation of the input features to generate the node v nt Output features on ; The spatial graph convolution S-GC and the temporal graph convolution T-GC are used to extract the features on the spatiotemporal graph and obtain the feature vector, where the extracted features are expressed as: in, Represents the spatial graph convolution used to connect nodes v in S-GC pq and node v nq The weights or connection coefficients between them are the spatial adjacency matrix A in the space-time graph. s The elements of are used to measure the spatial relationship or connection strength between nodes; specifically, Represents node v pq and node v nq Between Spatial connection strength is used to determine how to propagate feature information in the spatial graph; N is used to represent the index range of the node, indicating that the node v pq and node v nq The node index in ; q represents the index of the time step or time frame, which is used to represent different time steps or time frames in the temporal graph convolution T-GC; p is the index of the node of the spatial graph; W represents the weight matrix or weight parameter, which is used for linear transformation or weighted aggregation of features, that is, w1(v pq ) and w2(v nq ) are used to perform weighted aggregation on the input features; the weight matrix is a learnable parameter in the neural network model, which is used to adjust and control the way features are propagated and aggregated between different nodes to capture the complex feature relationships in the spatiotemporal graph; during the training process of the model, the weight matrix will be adjusted according to the loss function to optimize the performance of the model.
9. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 7, characterized in that: The specific steps of step S43 are: S431: first preprocess the feature vector to obtain a preprocessed feature vector; S432: Select a classifier model and use the labeled training data to train the classifier; S433: Evaluate the trained classifier using unlabeled test data; S434: Use the trained classifier to classify and identify the preprocessed feature vector, and determine whether the input data belongs to the fall category based on the feature vector.
10. The skeleton detection and fall detection method based on improved spatiotemporal adaptive graph convolution according to claim 9, characterized in that: The preprocessing method in step S431 includes feature normalization and dimensionality reduction; the classifier model in step S432 selects support vector machine SVM; and the performance of the classifier is evaluated by calculating the accuracy, recall rate and F1 score in step S433.
Citation Information
Patent Citations
Electric power operation abnormal behavior identification method based on human skeleton key points
CN115966025A
Mine personnel target video tracking method based on YOLOv5-Deepsort algorithm and storage medium
CN116681724A
Skeleton detection and fall detection method based on improved space-time adaptive graph convolution
CN117372844A
Cited By
Low-altitude radar and high-point monitoring combined bird repelling method
CN120314908A
Spatial-temporal topology learning-based skeleton action recognition method
CN120472535A
Skeletal motion recognition method based on spatiotemporal topology learning
CN120472535B
Performance test system and equipment of fan grating sensor and storage medium
CN120740652A
Method and system for revealing animal behaviors in zoo
CN120805034A