Fall detection method and device, equipment and storage medium

By combining space-time graph convolution model and tree model for fall detection in an edge computing environment, the challenges of the human posture recognition algorithm based on deep learning in real time and accuracy are solved, and efficient and accurate fall detection is achieved.

CN120047874APending Publication Date: 2025-05-27CISDI INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510212050.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In an edge computing environment, deep learning-based human posture recognition algorithms are difficult to achieve real-time fall detection, and reducing the number of timing features to improve inference speed will affect the accuracy and reliability of the detection.

Method used

A strategy of combining spatiotemporal graph convolution model and tree model is used for fall detection. The specific steps include obtaining the data set, building a spatio-temporal graph convolution model and tree model, extracting target features through preset feature processing strategies, and combining the two models for fall detection and result-weighted fusion.

Benefits of technology

The double improvement of fall detection accuracy and real-time performance in the edge computing environment is achieved, reducing the requirements for edge device hardware conditions, and maintaining detection accuracy and efficient inference under limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047874A_ABST
    Figure CN120047874A_ABST
Patent Text Reader

Abstract

The invention provides a tumble detection method and device, equipment and a storage medium, and the method comprises the steps: obtaining a data set, the data set comprises a plurality of samples representing a target tumble process and a non-tumble process, and each sample comprises the continuously changing position of a skeleton point in a preset time period; constructing a space-time diagram convolution model, and training the space-time diagram convolution model based on the data set to obtain a first fall detection model; feature extraction is carried out on each sample according to a preset feature processing strategy, target features of the samples are determined, and the preset feature processing strategy comprises extraction of vector features of skeleton points and bending angles of the skeleton points; constructing a tree model, and training the tree model based on the target features to obtain a second tumble detection model; and performing tumble detection by combining the first tumble detection model and the second tumble detection model, and performing weighted fusion on the two detection results as a final tumble detection result. In this way, the real-time performance and accuracy of fall detection are improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fall detection, and in particular to a fall detection method, device, equipment and storage medium. Background Art

[0002] With the rapid development of intelligent technology, especially in the field of smart city monitoring, how to efficiently and accurately ensure the safety of the crowd has become an urgent problem to be solved. As a sudden and potentially life-threatening health problem, the timely detection and early warning of fall events are of great significance for reducing injuries and ensuring public safety. At present, with the rapid development of computer vision and deep learning technology, the fall detection method based on human posture recognition has gradually become a hot topic in research and application, that is, the human posture recognition algorithm based on deep learning accurately recognizes the human posture, and then realizes the detection of fall events. However, in practical applications, especially in edge computing environments, it faces the dual challenges of reasoning speed and resource consumption. Among them, the human posture recognition algorithm based on deep learning needs to process a large number of time series features, which leads to a low reasoning speed on the edge device and is difficult to meet the requirements of real-time detection. In related technologies, it is usually adopted to reduce the number of time series features, that is, to perform fall detection based on the frame data collected by the edge device for a shorter time to improve the reasoning speed.

[0003] However, reducing the number of time series features will seriously affect the results of fall detection, resulting in reduced accuracy and reliability of detection. Therefore, how to achieve both accuracy and real-time performance of fall detection is an urgent problem to be solved. Summary of the invention

[0004] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0005] In view of the above-mentioned shortcomings of the prior art, the present application discloses a fall detection method, device, equipment and storage medium to solve the above-mentioned technical problem of how to simultaneously improve the accuracy and real-time performance of fall detection.

[0006] In a first aspect, the present application provides a fall detection method, the method comprising: obtaining a data set, the data set comprising multiple samples respectively characterizing a target fall process and a non-fall process, each of the samples comprising a position of a skeleton point that continuously changes in a preset time period; constructing a spatiotemporal graph convolution model, training the spatiotemporal graph convolution model based on the data set, and obtaining a first fall detection model; performing feature extraction on each of the samples according to a preset feature processing strategy to determine the target features of the sample, the preset feature processing strategy comprising extracting vector features of the skeleton points and the bending angles of the skeleton points; constructing a tree model, training the tree model based on the target features, and obtaining a second fall detection model; performing fall detection in combination with the first fall detection model and the second fall detection model, and weighted fusion of the two detection results as the final fall detection result.

[0007] In one embodiment of the present application, feature extraction is performed on each of the samples according to a preset feature processing strategy to determine the target features of the samples, including: performing structured transformation on each of the samples, determining the coordinates of the skeleton points in each sample within a preset time period, and using them as feature items; performing feature processing on the feature items corresponding to each of the samples, obtaining vector features of the skeleton points in the sample, and using them as the target features, wherein the vector features are obtained by performing vector calculation on the coordinates of the same skeleton point within the preset time period; calculating the bending angle of the skeleton point within the preset time period according to the feature items corresponding to each of the samples, and using them as the target features of the samples, wherein at least the bending angles corresponding to the left leg joint points, the right leg joint points, the left hand joint points, the right hand joint points, the trunk center point, the head point, and the foot center point are included.

[0008] In one embodiment of the present application, the first fall detection model and the second fall detection model are constructed, including: dividing fall training samples and fall verification samples from the data set; respectively setting the hyperparameters and tuning strategies of the spatiotemporal graph convolution model and the tree model; training the spatiotemporal graph convolution model based on the fall training samples and the hyperparameters to obtain the trained spatiotemporal graph convolution model; training the tree model based on the target features of the fall training samples and the hyperparameters to obtain the trained tree model, which is a LightGBM model; tuning the trained spatiotemporal graph convolution model according to the tuning strategy and the fall verification samples to obtain the first fall detection model; tuning the trained tree model according to the tuning strategy and the target features of the fall verification samples to obtain the second fall detection model.

[0009] In one embodiment of the present application, the combination of the first fall detection model and the second fall detection model for fall detection includes: obtaining a video frame sequence of the area to be detected currently collected by the edge device; inputting the video frame sequence into a preset detection model to determine the skeleton points of at least one of the targets in the video frames corresponding to different time nodes, and the preset detection model is used to detect the skeleton points of the target; tracking the skeleton points of the same target at consecutive time nodes to obtain the skeleton points of the same target in each preset time period; and inputting the skeleton points of the same target in each preset time period into the first fall detection model and the second fall detection model, respectively.

[0010] In one embodiment of the present application, the video frame sequence is input into a preset detection model to determine the skeleton points of at least one of the targets in the video frames corresponding to different time nodes, including: inputting the video frame sequence into a target detection model to obtain target images at consecutive time nodes, the preset detection model includes a target detection model for detecting the target and a skeleton point detection model for detecting skeleton points connected in series, the target image takes the human body as the target; inputting the target image into the skeleton point detection model to obtain the position of the skeleton point in each of the target images.

[0011] In one embodiment of the present application, before combining the first fall detection model and the second fall detection model to perform fall detection, it also includes: quantizing the first fall detection model, the target detection model and the skeleton point detection model respectively; converting the model format of the first fall detection model, the target detection model and the skeleton point detection model respectively, and the model format conversion includes at least ONNX model conversion, MLIR model conversion, and bmodel model conversion, and the quantization processing is after the ONNX model conversion; deploying the second fall detection model, as well as the first fall detection model, the target detection model and the skeleton point detection model on the edge device.

[0012] In one embodiment of the present application, the data set acquisition includes: acquiring an initial data set, the initial data set including a video frame sequence of the target when it falls and when it does not fall, the video frame sequence being acquired through historical collection and customization of the edge device; dividing the initial data set according to the preset time period to obtain a plurality of initial samples, the initial sample including the video frame sequence within the preset time period; inputting each of the initial samples into the preset detection model to obtain the skeleton points of the same target within each of the preset time period, and using them as the samples to construct the data set.

[0013] In a second aspect, the present application provides a fall detection device, which includes: a data acquisition module, used to obtain a data set, the data set includes multiple samples that respectively characterize a target fall process and a non-fall process, and each of the samples includes a position of a skeleton point that continuously changes in a preset time period; a first construction module, used to construct a spatiotemporal graph convolution model, and train the spatiotemporal graph convolution model based on the data set to obtain a first fall detection model; a feature extraction module, used to extract features from each of the samples according to a preset feature processing strategy to determine the target features of the sample, and the preset feature processing strategy includes extracting vector features of skeleton points and bending angles of skeleton points; a second construction module, used to construct a tree model, and train the tree model based on the target features to obtain a second fall detection model; a fall detection module, used to perform fall detection in combination with the first fall detection model and the second fall detection model, and weightedly fuse the two detection results as the final fall detection result.

[0014] In a third aspect, the present application also provides an electronic device, comprising: a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the fall detection method as described in the above embodiment.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor of a computer, the computer executes the method described in the above embodiments.

[0016] Beneficial effects of the present application: The present application proposes a fall detection method, device, equipment and storage medium. Obtain a data set, the data set includes multiple samples that respectively characterize the target fall process and non-fall process, each sample includes the position of the skeleton point that changes continuously in a preset time period; construct a spatiotemporal graph convolution model, train the spatiotemporal graph convolution model based on the data set, and obtain a first fall detection model; extract features from each sample according to a preset feature processing strategy to determine the target features of the sample, the preset feature processing strategy includes extracting vector features of the skeleton points, and the bending angle of the skeleton points; construct a tree model, train the tree model based on the target features, and obtain a second fall detection model; combine the first fall detection model and the second fall detection model to perform fall detection, and perform weighted fusion on the two detection results as the final fall detection result. In this way, by integrating deep learning and machine learning technologies, that is, combining the spatiotemporal graph convolution and tree model for fall detection strategies, in the edge computing environment, the accuracy and real-time performance of fall detection are both improved, and the requirements for the hardware conditions of edge devices are reduced, so that detection accuracy and efficient reasoning can be met at the same time even in resource-limited situations.

[0017] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0019] Figure 1 is a flow chart of a fall detection method shown in an exemplary embodiment of the present application;

[0020] Figure 2 is a schematic diagram of a model quantization shown in an exemplary embodiment of the present application;

[0021] Figure 3 is a schematic diagram of a fall detection process shown in an exemplary embodiment of the present application;

[0022] Figure 4 is a schematic diagram of a fall detection application shown in an exemplary embodiment of the present application;

[0023] Figure 5 is a block diagram of a fall detection device shown in an exemplary embodiment of the present application;

[0024] Figure 6 It is a structural diagram of a computer system suitable for implementing the electronic device of the present application, shown as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.

[0026] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the form, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0027] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.

[0028] It should be noted that falls, as a sudden health event, may not only cause serious physical injuries, but may also endanger life due to failure to receive timely treatment. Therefore, the development of efficient and accurate fall detection technology is of great significance for preventing injuries and improving public safety. In recent years, with the rapid development of computer vision and deep learning technology, fall detection methods based on human posture recognition have gradually become a hot topic in research and application. Human posture recognition algorithms based on deep learning, such as Spatial Temporal Graph Convolutional Networks (STGCN), have been widely used in the field of fall detection. These algorithms can accurately identify fall events by capturing and analyzing the spatiotemporal characteristics of human motion.

[0029] However, in edge computing environments, the reasoning speed and resource consumption of these algorithms have become key factors that restrict their widespread application. Edge computing environments have the characteristics of limited resources and high real-time requirements. In order to meet the needs of real-time detection, edge devices can often only adopt low-frequency data collection strategies, such as collecting data once per second. However, human posture recognition algorithms based on deep learning need to process a large number of time series features to accurately detect. For example, the fall detection scenario commonly used by the original STGCN algorithm is to take 50-60 frames of data for reasoning within 1 to 2 seconds, which is significantly different from the data efficiency collected by edge devices, resulting in slow reasoning speed and difficulty in meeting the needs of real-time detection. In addition, although reducing the number of time series features can increase the reasoning speed, it often comes at the expense of detection accuracy. For example, in order to meet real-time requirements, edge devices may only collect 30 frames of data for 30 seconds for prediction. However, this approach will seriously affect the results of STGCN fall detection and reduce the accuracy and reliability of detection.

[0030] In response to the challenges of fall detection in edge computing environments, technical solutions have been proposed to improve the inference speed by optimizing the algorithm structure, reducing model parameters, or reducing resource consumption by introducing a lightweight network structure. However, these solutions often find it difficult to achieve real-time detection while ensuring detection accuracy. In addition, due to the limitations of the hardware conditions and computing power of edge devices, the effects of these technical solutions in practical applications are often unsatisfactory.

[0031] Based on this, the present application proposes a fall detection method, device, equipment and storage medium to achieve dual improvement in the accuracy and real-time performance of fall detection.

[0032] See also Figure 1 , is a flow chart of a fall detection method shown in an exemplary embodiment of the present application. Figure 1 As shown, in an exemplary embodiment, the fall detection method includes at least steps S110 to S150, which are described in detail as follows:

[0033] Step S110, obtaining a data set, the data set includes a plurality of samples respectively representing a falling process and a non-falling process of the target, and each sample includes a position of a skeleton point that changes continuously during a preset time period.

[0034] In one embodiment of the present application, each sample includes at least 17 types of human skeleton points, such as nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle; when the target moves, its skeleton points will show dynamic change characteristics in the time series, reflecting the target's motion state and posture changes, and the preset time period is a time series formed by a fixed number of continuous time nodes, therefore, the sample includes the position of the skeleton point that changes continuously in the preset time period. In addition, to ensure the real-time nature of subsequent fall detection, the preset time period is shorter than the conventional setting (i.e., when the sampling frequency of the edge device is 1 frame per second, 50 to 60 frames of data are usually required to be collected), for example, shortened to 30 seconds or less, or even 6 seconds or 5 seconds.

[0035] In step S110, a data set is obtained, including: obtaining an initial data set, the initial data set includes a video frame sequence when the target falls and does not fall, and the video frame sequence is obtained by historical acquisition and customization of the edge device; dividing the initial data set according to a preset time period to obtain multiple initial samples, and the initial sample includes a video frame sequence within a preset time period; each initial sample is input into a preset detection model to obtain the skeleton points of the same target within each preset time period, and used as a sample to construct a data set. In addition, the type of the sample is also marked, including two types of fall or non-fall.

[0036] In one embodiment of the present application, the initial data set is divided according to each preset time period. The video frame sequence can be divided according to each preset time period, and the same video frames are not included in different initial samples. Alternatively, as long as the continuous video frame sequence meets the preset time period, it is divided into one initial sample, and the same video frames can be included in different initial samples.

[0037] In one embodiment of the present application, before the initial data set is divided according to time, the initial data set needs to be preprocessed, including data cleaning, data integration, data conversion and data reduction. The main task of data cleaning is to delete irrelevant data and duplicate data in the original data set, and to handle missing values, outliers, etc. Data conversion includes data standardization, normalization and feature extraction. Among them, the processing of missing values ​​can also be handled by removing missing records and additional camera retake data. Data preprocessing is also extended with data enhancement, that is, using data enhancement technology data such as data mixing and synthetic data to expand the scale and diversity of the initial data set and its derivative data sets, so as to improve the generalization ability and performance of the subsequent established model. In addition, it is necessary to annotate the initial data set with humans as the annotated objects, that is, to annotate the targets appearing in each video frame in the video frame sequence.

[0038] In one embodiment of the present application, the setting of the preset time period is affected by the sampling frequency of the edge device, and the sampling frequency of the edge device is usually low. For example, one frame of data is collected per second. At this time, if a sequence formed by 30 frames of video frames is taken as an initial sample, the preset time period is set to 30 seconds. That is to say, the data of the skeleton points of the same target detected within 30 consecutive seconds is taken as a sample, wherein there may be multiple targets within the same preset time period, and the skeleton points of each target are taken as a separate sample.

[0039] In one embodiment of the present application, the preset detection model includes a target detection model for detecting targets and a skeleton point detection model for detecting skeleton points, which are connected in series in sequence. Obtaining skeleton points in video frames includes: inputting the video frame sequence into the target detection model to obtain target images of continuous time nodes, where the target image targets the human body; inputting the target image into the skeleton point detection model to obtain the position of the skeleton points in each target image.

[0040] In one embodiment of the present application, each target image is a human body as the target, and the target images of the continuous time nodes are actually the target images corresponding to the continuous video frames. The target images are input into the skeleton point detection model, and the skeleton point detection model can be used to identify the skeleton points of each target in the continuous video frames. In addition, before the target image is input into the skeleton point detection model, it needs to be labeled, and the labeled object is the skeleton point, and the labeled content includes the type of the skeleton point and the labeling order.

[0041] In one embodiment of the present application, in order to improve the reasoning speed of fall detection, the target detection model is constructed based on the YOLOv5s (i.e., the fifth generation of YOLO small model) model, which is a lightweight model with fewer parameters and computational complexity, suitable for resource-constrained scenarios, while maintaining high detection accuracy and real-time performance. The network structure of the YOLOv5s model mainly consists of three parts: Backbone (backbone network), Neck (neck network), and Head (head network). Among them, the backbone network adopts CSP Darkn et53 (i.e., a deep convolutional neural network model), and uses the CSP (Cross Stage Partial) module to optimize the gradient flow of the network, reduce redundant calculations, and improve the reasoning speed and model efficiency; the neck network usually includes a feature fusion module, such as PANet (Path Aggregation Network), which is used to fuse feature maps of different scales and enhance the feature expression ability of the model; the neck network usually includes some feature fusion modules, such as FPN (Feature Pyramid Network) and PANet (Path Aggregation Ne twork), etc. These modules achieve bottom-up and top-down feature fusion through upsampling, downsampling and feature map concatenation, and generate feature maps with multi-scale information. These feature maps are then passed to the head network for final target detection; the head network is responsible for predicting the target category, position and bounding box deviation based on the extracted features. It is based on Anchor (anchor point). The target detection process can include multiple groups of scale predictions to predict targets of different sizes. In addition, compared with other models in the same series, YOLOv5s has fewer convolutional layers and a relatively small number of channels in the feature map, which greatly reduces the computational overhead.

[0042] In one embodiment of the present application, the skeleton point detection model is constructed based on the HRNet (High-Resolution Network) model, which is an efficient and accurate deep neural network architecture designed for posture estimation tasks. It is particularly good at processing high-resolution input data and maintaining a fine representation of multi-scale features. It is suitable for computer vision tasks that need to retain image spatial information and extract high semantic information at the same time. Among them, the network structure of HRNet is usually composed of multiple stages, and each stage contains branches of multiple resolutions. Information is exchanged between these branches, so that the high-resolution branches can obtain the semantic information extracted by the low-resolution branches, and the low-resolution branches can also obtain the detail information extracted by the high-resolution branches, so that the mutual fusion of features of different resolutions is achieved. In this way, in human posture recognition, the high-resolution features of HRN et can locate joints very accurately. This is because the high-resolution branches can continue to retain the details in the input image, while the low-resolution branches provide contextual information, thereby improving the accuracy of skeleton point detection in the human body.

[0043] In one embodiment of the present application, before applying the target detection model, it is necessary to construct it, including: obtaining a first data set for the target detection model, wherein the first data set may be from the same source as the initial data set, or collecting new video frame sequences when the target falls and does not fall, and the first data set and the initial data set have the same processing flow before being input into the target detection model; dividing the first data set into a first training set, a first validation set, and a first test set; setting hyperparameters and tuning strategies of the YOLOv5s model; training the YOLOv5s model based on the hyperparameters of the YOL Ov5s model and the first training set to obtain a trained YO LOv5s model; tuning the trained YOLOv5s model according to the first validation set and the tuning strategy to obtain the target detection model; wherein the tuned YOLOv5s model is tested based on the first test set; if the test result does not meet the preset first test indicator, the YOLOv5s model is retrained; if the test result meets the preset first test indicator, it is used as the target detection model.

[0044] In one embodiment of the present application, before applying the skeleton point detection model, it is necessary to construct it, including: obtaining a second data set for skeleton point detection, wherein the second data set can be constructed and annotated by the target image output by the target detection model; dividing the second data set into a second training set, a second verification set and a second test set; setting the hyperparameters and tuning strategy of the HRNet model; training the HRNet model based on the hyperparameters of the HRNet model and the second training set to obtain the trained HRNet model; tuning the trained HRNet model according to the second verification set and the tuning strategy to obtain the skeleton point detection model; wherein the tuned HRNet model is tested based on the second test set; if the test result does not meet the preset second test indicator, the HRNet model is retrained; if the test result meets the preset second test indicator, it is used as the skeleton point detection model.

[0045] In one embodiment of the present application, the hyperparameters of the YOLOv5s model, the HRNet model, and the spatiotemporal graph convolution model include at least the number of training rounds (Epochs), the batch size (Batch Size), the learning rate (Learning Rate), and the optimizer (Optimizers), and the tuning strategy includes at least the cosine annealing learning rate, the sliding exponential average, etc., which can be set according to actual needs. For example, the number of training rounds of the YOLOv5s model is set to 800, the batch size is set to 8, the learning rate is set to 0.001, and the optimizer selects AdamW (an optimizer based on the gradient descent algorithm); the number of training rounds of the HRNet model is set to 200, the batch size is also set to 8, the learning rate is set to 0.001, and the optimizer selects AdamW (an optimizer based on the gradient descent algorithm).

[0046] In one embodiment of the present application, the above-mentioned annotation of any data set can be performed according to a preset standard (such as Common Objects in Context, COCO) or a custom standard.

[0047] Step S120: construct a spatiotemporal graph convolution model, train the spatiotemporal graph convolution model based on the data set, and obtain a first fall detection model.

[0048] In one embodiment of the present application, based on the dynamic change characteristics of skeleton points, the changes of skeleton points in time series are detected and analyzed, so as to more accurately judge whether the target has experienced abnormal events such as falls. Therefore, the first fall detection model is set as a spatial-temporal graph convolution (Spatial-Temporal Graph Convolutional Ne twork, STGCN) model, which is a deep learning model that combines the characteristics of graph convolutional network (GCN) and convolutional neural network (CNN) and is used to process the spatiotemporal relationship in graph structured data. It can capture the dynamic relationship between space and time in irregular graph structured data. The core idea is to map time series data to a graph, where the nodes of the graph represent entities in space (such as human skeletal joints), the edges between the nodes represent the relationship between spatial entities, and the time dimension is processed by applying a series of graph convolutions and time convolutions on the graph. The spatiotemporal graph convolution model includes graph convolutional layers, which are used to capture spatial dependencies and update the representation of nodes by combining the features of nodes with the features of their neighbors; temporal convolutional layers, which are used to capture the temporal dependencies of time series data; and spatial-temporal layers, which combine graph convolution and temporal convolution to simultaneously process spatial and temporal dependencies. In this way, the first fall detection model can simultaneously model the spatial dependencies of human skeleton points and the dynamic patterns of these skeleton points changing over time, thereby effectively identifying complex fall behaviors.

[0049] In step S120, a first fall detection model is constructed, including: dividing a fall training sample and a fall verification sample from a data set; setting hyperparameters and a tuning strategy of a spatiotemporal graph convolution model; training the spatiotemporal graph convolution model based on the fall training sample and the hyperparameters to obtain a trained spatiotemporal graph convolution model; and tuning the trained spatiotemporal graph convolution model according to the tuning strategy to obtain a first fall detection model. The first fall detection model is deployed on an edge device.

[0050] In one embodiment of the present application, the hyperparameters and optimization strategies of the spatiotemporal graph convolution model can be adjusted according to actual needs. For example, the number of training rounds is set to 400, the batch size is set to 8, the learning rate is set to 0.001, the optimizer still selects AdamW, and adopts cosine annealing learning rate strategy, sliding exponential average and other tuning strategies.

[0051] In one embodiment of the present application, the data set is also divided into fall test samples for subsequent testing of the spatiotemporal graph convolution model and the tree model. For example, the tuned spatiotemporal graph convolution model is tested based on the fall test samples; if the test result does not meet the preset third test index, the spatiotemporal graph convolution model is retrained; until the test result meets the preset third test index, it is used as the first fall detection model.

[0052] Step S130, extracting features from each sample according to a preset feature processing strategy to determine target features of the sample. The preset feature processing strategy includes extracting vector features of bone points and bending angles of bone points.

[0053] In one embodiment of the present application, each sample is structured transformed respectively, and the coordinates of the bone points in each sample within a preset time period are determined and used as feature items; feature processing is performed on the feature items corresponding to each sample to obtain vector features of the bone points in the sample and use them as target features, where the vector features are obtained by vector calculation of the coordinates of the same bone point within a preset time period; the bending angles of the bone points within the preset time period are calculated based on the feature items corresponding to each sample and used as the target features of the samples, including at least the bending angles corresponding to the left leg joint points, the right leg joint points, the left hand joint points, the right hand joint points, the trunk center points, the head points and the foot center points.

[0054] In one embodiment of the present application, the coordinates of the skeleton points at each time node include an x-axis (horizontal axis) and a y-axis (vertical axis). The structured transformation is essentially to use the x-axis and y-axis of each skeleton point as a feature item of the sample. Then, taking the number of skeleton point types as 17 and the preset time period including t time nodes (video frames) as an example, each sample should include 17*2*t feature items corresponding to the same target. In addition, other structured transformations can be implemented according to actual needs. For example, the fall type to which the sample belongs, i.e., fall or non-fall, is converted into a feature item, but it does not participate in the calculation of subsequent target features.

[0055] In one embodiment of the present application, by determining the maximum time node and the minimum time node of each sample in the preset time period, and then calculating the vector between the feature items corresponding to each sample at the maximum time node and the minimum time node, the vector features of the skeleton points in the sample are obtained, which is equivalent to compressing the number of features and extracting the motion features of each target skeleton point, wherein the vector calculation between the same feature items is aimed at, that is, the vector calculation of the coordinates between the same skeleton points. In addition, in the same sample, the bending angle of the skeleton points of the same target in the preset time period is also calculated based on the coordinates of the skeleton points at the maximum time node (the last video frame). That is to say, taking the number of feature items of a time node as 17*2*, and the preset time period including t time nodes as an example, at this time, the number of feature items corresponding to the same target in each sample is 17*2*t, the number of feature vectors determined according to the feature items in each sample is 17*2, the number of types of bending angles of the skeleton points is 6, and the number of target features corresponding to one sample is 17*2+6, totaling 40.

[0056] Step S140: construct a tree model, train the tree model based on the target features, and obtain a second fall detection model.

[0057] In one embodiment of the present application, LightGBM (Light Gradient Boosting Machine) is a machine learning model based on a gradient boosting decision tree (GBDT, Gradient Boosting Decision Tree), which has the characteristics of high efficiency, high precision and support for category features. It uses a histogram algorithm to replace the traditional decision tree algorithm, and also supports parallel training and feature parallelization. It uses leaf-based tree growth, feature selection, regularization and other methods, and performs well in processing structured data, which can significantly improve the training speed and reduce memory consumption. Therefore, the tree model uses the LightGBM model, which ensures the real-time nature of fall detection based on its ability to quickly predict, and can immediately issue a warning signal when a fall event occurs. In addition, the model's efficient algorithm design, small model size, flexible deployment options, real-time prediction capabilities, and good scalability and compatibility make it very suitable for edge devices, and efficient model operation can be achieved even on edge devices.

[0058] In step S140, a second fall detection model is constructed, including: setting hyperparameters and tuning strategies of the tree model; training the tree model based on the target features and hyperparameters of the fall training sample to obtain a trained tree model; tuning the trained tree model according to the tuning strategy and the target features of the fall verification sample to obtain a second fall detection model. The second fall detection model is deployed on the edge device.

[0059] In one embodiment of the present application, the LightGBM model adopts a boosting algorithm GBDT (Gradient Boosting Decision Tree), and its hyperparameters include at least learning rate, maximum tree depth, minimum number of samples of leaf nodes, etc. The boosting algorithm adopted is GBDT (Gradient Boosting Decision Tree), and the tuning strategy includes at least adjusting the hyperparameters and the number of iterations, which can be set according to actual needs. For example, the hyperparameters are set to a learning rate of 0.01, a maximum tree depth of 7, and a minimum number of samples of leaf nodes of 25.

[0060] In one embodiment of the present application, the tuned tree model is tested based on the fall test sample; if the test result does not meet the preset fourth test indicator, the tree model is retrained; until the test result meets the preset fourth test indicator, it is used as the second fall detection model.

[0061] In one embodiment of the present application, the test indicators of any of the above-mentioned models, as well as the verification indicators that can be used for each model when tuning, are set according to the actual needs of the model, such as any one of the average precision (AP), recall rate, and F1 score, and compared with the corresponding preset values ​​(such as the preset first, second, third, and fourth test indicators).

[0062] In one embodiment of the present application, the ratio between the training set, the validation set and the test set divided by any of the above data sets can be set according to actual needs, for example, the ratio is set to 8:1:1.

[0063] Step S150: performing fall detection by combining the first fall detection model and the second fall detection model, and performing weighted fusion on the two detection results as the final fall detection result.

[0064] In one embodiment of the present application, the two fall detection models do not have a sequential order in execution. In essence, the two models are processed in parallel. By running the first fall detection model and the second detection model in parallel on the edge device for fall detection, the local detection efficiency is improved, the detection time is shortened, and the detection results of the two fall detection models are weighted and fused, which greatly avoids the false alarm of a single fall detection model due to the real-time requirement, thereby improving the accuracy of fall detection. For example, since the first detection model is a spatiotemporal graph convolution model, if the number of its temporal features is reduced in order to ensure the real-time performance of the detection, that is, the number of frames of frame data collected by the edge device required for the detection is reduced, the accuracy of the detection will be affected. Based on this, a second fall detection model of machine learning is introduced to compensate for the inaccuracy of the spatiotemporal graph convolution model detection within a short frame. In addition, the weights can be adjusted according to actual needs. For example, the weights of the first fall detection model and the second fall detection model are both set to 0.5.

[0065] Before step S150, the above models need to be deployed on the edge device to achieve efficient reasoning on the edge device, including: quantizing the first fall detection model, the target detection model and the skeleton point detection model respectively; converting the model formats of the first fall detection model, the target detection model and the skeleton point detection model respectively, and the model format conversion includes at least ONNX model conversion, MLIR model conversion and bmodel model conversion, and the quantization processing is performed after the ONNX model conversion; deploying the second fall detection model, as well as the first fall detection model, the target detection model and the skeleton point detection model on the edge device.

[0066] In one embodiment of the present application, the quantization process adopts INT8 (8-bit integer) quantization to convert the parameters in the model (such as weights and activation values) from high-precision floating point numbers (such as FP32) to low-precision 8-bit integers (INT8).

[0067] See also Figure 2 , is a schematic diagram of a model quantization shown in an exemplary embodiment of the present application. Figure 2 The YOLOv5s model, HRNet model, and STGCN model in the training are the target detection model, the skeleton point detection model, and the first fall detection model respectively. Assuming that these three neural network models are built under the PyT orch framework, they are usually stored in the Pt model format. Of course, if they are built under other frameworks, they may also be stored in other formats. In order to better deploy the model on the edge device, it needs to be converted into a format that is easy to adapt to the edge computing environment, such as Figure 2Therefore, first convert the target detection model, the skeleton point detection model, and the first fall detection model into a model converted into ONNX format, and then convert the ONNX model into MLIR format, generate a quantization table, and convert the MLIR format model file into bmodel format according to the quantization table. Among them, ONNX (Open Neural Network Exchange) is an open deep learning model exchange format that allows models to be converted from one deep learning framework to another for deployment and reasoning on different platforms and devices. MLIR (Multi-Level Intermediate Representation) is a framework or representation method for building reusable and extensible compilers.

[0068] In one embodiment of the present application, the target detection model, the skeleton point detection model, and the first fall detection model can all be quantized and model converted by a highly integrated deep learning vision processor (Tensor Processing Unit, TPU), such as a computing platform, to improve their reasoning performance.

[0069] In step S150, fall detection is performed, including: obtaining a video frame sequence of the area to be detected currently collected by the edge device; inputting the video frame sequence into a preset detection model to determine the skeleton points of at least one target in the video frames corresponding to different time nodes, and the preset detection model is used to detect the skeleton points of the target; tracking the skeleton points of the same target at consecutive time nodes to obtain the skeleton points of the same target in each preset time period; inputting the skeleton points of the same target in each preset time period into the first fall detection model and the second fall detection model, respectively.

[0070] In one embodiment of the present application, it is necessary to take the skeleton points of the same target within each preset time period as an input item, calculate the target feature corresponding to the input item, that is, calculate the target feature according to the coordinates of the skeleton points of the same target within each preset time period, input the input item into the first fall detection model, and input the target feature corresponding to the input item into the second fall detection model.

[0071] In one embodiment of the present application, the target tracking adopts a multi-target tracking algorithm, such as BOTSO RT, which can be applied to the situation where multiple targets appear at the same time, and can also improve the efficiency and reliability of detection. In addition, in addition to being used after the skeleton point detection model, the target tracking can also be used before the skeleton point detection model to track the target images output by the target detection model at different time nodes to determine the same target at; The target image corresponding to the continuous time node is then input into the skeleton point detection model to determine the position of the skeleton points of the same target at the continuous time nodes. When the continuous time nodes meet the preset time period, that is, the number of frames of the video frame where the skeleton points of the same target are detected meets the preset number of frames, the first and second fall detection models are input. In other words, the first and second fall detection models can detect multiple targets at the same time, thereby improving the efficiency of fall detection.

[0072] In one embodiment of the present application, BOTSORT (Bottom-Up Tracking by Sorting) is an algorithm for multi-target tracking, which improves the accuracy of multi-target tracking by combining appearance information and motion information. First, BOTSORT detects all targets in each video frame through a target detection model, obtains the bounding box and confidence of each target, and selects high-confidence targets according to a preset confidence threshold as candidate objects for subsequent association. For the existing trajectory, that is, the position, speed and other information of the tracked target, a Kalman filter is used to perform state prediction to predict the position of the target in the current frame. Then, a two-step association method is used to determine the object: the first step of association is to associate the high-confidence target with the predicted trajectory using the Hungarian algorithm, which is mainly based on the overlap degree (IoU) between the targets and the cosine distance association of the appearance features; the second step of association is to use low-confidence targets for further association for the high-confidence targets and trajectories that are not matched in the first step of association, so as to restore the targets that may be missed due to some reasons (such as occlusion, light changes, etc.). Finally, according to the association results, the trajectory information of the matched target is updated, including position, speed and appearance features. For the unmatched target, whether to add a new trajectory or delete the trajectory is determined according to its continuous frame loss status. For the unmatched tracked target, its unmatched frame count is increased; for the unmatched new target, a new trajectory is initialized.

[0073] In one embodiment of the present application, if the preset time period is set to 30 seconds, the edge device collects one frame of video every second, and fall detection is performed on the target within this preset time period, then the length of the video frame sequence within the preset time period is 30. For example, at least one pedestrian is detected in the video frame, and the number of frames is calculated from the time the pedestrian is detected. The video frame sequence constructed by the first 30 frames of video frames is used as the first target sequence, and fall detection is performed on the target sequence. Then, the sequence of video frames from the i-th frame to the 30+i-th frame is recorded as the i-th target sequence. If there is no pedestrian in the middle, the sequence is recalculated.

[0074] In one embodiment of the present application, if the same target is lost before the preset time period is reached, fall detection is performed directly based on the skeleton points of the same target at consecutive time nodes. For example, the skeleton points of the same target are detected every 30 seconds, but the target only appears in 26 seconds, then target detection is performed directly based on the skeleton points of the same target within 26 seconds.

[0075] In one embodiment of the present application, if the final fall detection result is a fall, the final fall detection result is organized into an event, and then the edge device visualizes the event, for example, using the edge device to push the time to the AI ​​capability warehouse (i.e., a storage or processing unit with artificial intelligence capabilities) for centralized display. Among them, the event includes at least the discovery time, event location, event type, event source, device name, original material, annotation information, review status and other fields. For example, the edge device B located in area A detects the occurrence of a fall at 11:43, then the discovery time of the event is 11:43, the event location is area A, the name of the edge device B, the video frame sequence collected by the edge device B about the fall, the annotation information and review status annotated on the video frame sequence.

[0076] See also Figure 3 , is a schematic diagram of a fall detection process shown in an exemplary embodiment of the present application. Figure 3 As shown in Figure 1, the overall reasoning logic after linking all models deployed on edge devices is shown in Figure 1, where: Figure 3 The RTSP stream in the text refers to the streaming media data transmitted via RTSP (Real Time Streaming Protocol). Figure 3The target features in are not equivalent to the target features in this application, and represent the skeleton point features, which are used to describe the positions of the skeleton points. First, the edge device reads the video frame from the streaming data, and then detects whether there is a person in the current video frame through the target detection model. If there is no person, the next frame is judged; the detected person is tracked by the target tracking algorithm; then the target image of the person is input into the skeleton point detection model to obtain its skeleton point features, and the skeleton point features of the same target (person) are stored and collected. If the skeleton point features of any target are collected for 30 frames or the tracked target is lost, the fall detection model is performed, and the skeleton point features of 30 frames or possibly less than 30 frames are input into the STGCN model, that is, the first fall detection model for fall detection; and, after feature preprocessing and feature engineering of the collected 30 frames of skeleton point data, the corresponding vector features and bending angles are obtained and input into the LightGbm model, that is, the second fall detection model, and the detection results of the two fall models are weighted fused to determine whether there is a fall. For example, each weight is set to 0.5. If the final fall detection result after weighted fusion is greater than the preset fall threshold of 0.5, it means that a fall has occurred. Finally, the fall event is uploaded.

[0077] See also Figure 4 , is a schematic diagram of a fall detection application shown in an exemplary embodiment of the present application. The fall detection method of the present application is applied to a smart park, such as Figure 4As shown in the figure, firstly, the personnel data in the park is collected, that is, the video frame sequence of falling or non-falling in the park, and then it is made into a data set for human detection, and the YOLOv5s model is trained, verified and tested to obtain the target detection model; then, according to the target image output by the target detection model, a data set for identifying the skeleton points of human postures in the park is made, and the HRNet model is trained, verified and tested to obtain a posture rule skeleton point recognition model, that is, a skeleton point detection model; then a data set about pedestrian falls in the park is constructed, and it is input into the skeleton point detection model, and a human fall data set is made according to the output of the skeleton point detection model. The posture skeleton point dataset, that is, the sample in this application includes a dataset of the positions of the skeleton points that change continuously in a preset time period. The STGCN model is trained, verified and tested based on the dataset to obtain the first fall detection model (STGCN fall detection model); in addition, the posture skeleton point dataset is structured, data preprocessed and feature engineered to determine the target features of each sample in the dataset, and then the LightGBM model is trained, verified and tested based on the dataset, that is, the target features are input into the LightGBM model for training, verification and testing to obtain the second fall detection model. Finally, the target detection model, the skeleton point recognition model (equivalent to the skeleton point detection model) and the first fall detection model are uniformly quantized to INT8 on the computing platform to improve the reasoning performance, and then the target detection model, the skeleton point detection model and the two fall detection models are deployed in the edge device, the two fall detection models are inferred in parallel and the reasoning results are integrated, and the entire reasoning is realized by adding a target tracking algorithm to achieve real-time detection between frames. In this way, the edge device can receive real-time videos collected by designated cameras in the park for inter-frame reasoning, collect a specified number of frames for the same target to detect fall behavior, and organize the detection results into events and push them to the AI ​​capability warehouse for centralized display of the events.

[0078] Through the above method, edge devices are used to complete fall detection locally, effectively reducing data transmission delays while ensuring data privacy and security, and adapting to the multi-point and multi-angle monitoring needs of complex campus environments.

[0079] The fall detection method and device provided in the present application have the following advantages: First, by integrating deep learning and machine learning technologies, that is, combining spatiotemporal graph convolution and tree models for fall detection, the accuracy and real-time performance of fall detection are both improved in the edge computing environment, and the requirements for the hardware conditions of edge devices are reduced, so that both detection accuracy and efficient reasoning can be met even in resource-limited situations; second, multiple targets can be detected at the same time, thereby improving the efficiency of fall detection; third, edge computing devices are used to complete video data processing and model reasoning locally, effectively reducing data transmission delays while ensuring data privacy and security.

[0080] See also Figure 5 , is a block diagram of a fall detection device shown in an exemplary embodiment of the present application. Figure 5 As shown, in an exemplary embodiment, the fall detection device includes at least a data acquisition module 510, a first construction module 520, a feature extraction module 530, a second construction module 540 and a fall detection module 550, which are described in detail as follows:

[0081] A data acquisition module 510 is used to obtain a data set, the data set includes a plurality of samples representing a target falling process and a non-falling process, each sample includes a position of a skeleton point that changes continuously in a preset time period;

[0082] A first construction module 520 is used to construct a spatiotemporal graph convolution model, and train the spatiotemporal graph convolution model based on the data set to obtain a first fall detection model;

[0083] A feature extraction module 530 is used to extract features from each sample according to a preset feature processing strategy to determine target features of the sample. The preset feature processing strategy includes extracting vector features of bone points and bending angles of bone points.

[0084] A second construction module 540 is used to construct a tree model, and train the tree model based on the target feature to obtain a second fall detection model;

[0085] The fall detection module 550 is used to perform fall detection by combining the first fall detection model and the second fall detection model, and perform weighted fusion on the two detection results as the final fall detection result.

[0086] It should be noted that the fall detection device provided in the above embodiment and the fall detection method provided in the above embodiment belong to the same concept, wherein the contents of the operations performed by each module have been described in detail in the method embodiment and will not be repeated here.

[0087] The present application also provides an electronic device, comprising: a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the fall detection method as in the above embodiment.

[0088] See also Figure 6 , shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. It should be noted that, Figure 6 The computer system 600 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0089] like Figure 6As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 to the random access memory (RAM) 603, such as executing the method in the above embodiment. In the RAM 603, various programs and data required for system operation are also stored. The CPU 601, ROM 602 and RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0090] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 609 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read therefrom is installed into the storage section 608 as needed.

[0091] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 609, and / or installed from a removable medium 611. When the computer program is executed by a central processing unit (CPU) 601, various functions defined in the system of the present application are executed.

[0092] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer executes the rule engine configuration method for early warning as described above. The computer-readable storage medium may be included in the electronic device described in the above embodiment, or may exist independently without being assembled into the electronic device.

[0093] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, system or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, system or device. A computer program contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0094] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0095] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.

[0096] The above embodiments are merely illustrative of the principles and effects of the present application, and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.

Claims

1. A fall detection method, characterized in that: The method comprises: Acquire a data set, the data set comprising a plurality of samples respectively representing a falling process and a non-falling process of the target, each of the samples comprising a position of a skeleton point that changes continuously over a preset time period; Constructing a spatiotemporal graph convolution model, and training the spatiotemporal graph convolution model based on the data set to obtain a first fall detection model; Extracting features from each of the samples according to a preset feature processing strategy to determine target features of the sample, wherein the preset feature processing strategy includes extracting vector features of bone points and bending angles of bone points; Constructing a tree model, and training the tree model based on the target feature to obtain a second fall detection model; The first fall detection model and the second fall detection model are combined to perform fall detection, and the two detection results are weightedly fused as the final fall detection result.

2. The fall detection method according to claim 1, characterized in that: The extracting features of each sample according to a preset feature processing strategy to determine the target features of the sample includes: Performing structured transformation on each of the samples, respectively, determining the coordinates of the skeleton points in each of the samples within a preset time period, and using them as feature items; Performing feature processing on the feature items corresponding to each of the samples to obtain vector features of the skeleton points in the sample and using them as the target features, wherein the vector features are obtained by performing vector calculation on the coordinates of the same skeleton point within the preset time period; The bending angle of the skeleton point within the preset time period is calculated according to the feature items corresponding to each of the samples, and is used as the target feature of the sample, wherein at least the bending angles corresponding to the left leg joint point, the right leg joint point, the left hand joint point, the right hand joint point, the torso center point, the head point and the footstep center point are included.

3. The fall detection method according to claim 1, characterized in that: Constructing the first fall detection model and the second fall detection model includes: Dividing the data set into a fall training sample and a fall verification sample; Respectively setting the hyperparameters and tuning strategies of the spatiotemporal graph convolution model and the tree model; Training the spatiotemporal graph convolution model based on the fall training sample and the hyperparameter to obtain the trained spatiotemporal graph convolution model; Training the tree model based on the target features and the hyperparameters of the fall training samples to obtain the trained tree model, wherein the tree model is a LightGBM model; Tuning the trained spatiotemporal graph convolution model according to the tuning strategy and the fall verification sample to obtain the first fall detection model; The trained tree model is tuned according to the tuning strategy and the target features of the fall verification sample to obtain the second fall detection model.

4. The fall detection method according to claim 1, characterized in that: The performing fall detection by combining the first fall detection model and the second fall detection model includes: Obtain the video frame sequence of the area to be detected currently collected by the edge device; Inputting the video frame sequence into a preset detection model to determine the skeleton points of at least one of the targets in the corresponding video frames at different time nodes, wherein the preset detection model is used to detect the skeleton points of the target; Tracking the skeleton points of the same target at consecutive time nodes to obtain the skeleton points of the same target within each preset time period; The skeleton points of the same target within each of the preset time periods are input into the first fall detection model and the second fall detection model respectively.

5. The fall detection method according to claim 4, characterized in that: The step of inputting the video frame sequence into a preset detection model to determine the skeleton points of at least one target in the video frames corresponding to different time nodes includes: Input the video frame sequence into a target detection model to obtain a target image at a continuous time node, wherein the preset detection model includes a target detection model for detecting the target and a skeleton point detection model for detecting skeleton points connected in series, and the target image takes a human body as a target; The target image is input into the skeleton point detection model to obtain the position of the skeleton point in each target image.

6. The fall detection method according to claim 5, characterized in that: Before performing fall detection by combining the first fall detection model and the second fall detection model, the method further includes: Performing quantization processing on the first fall detection model, the target detection model and the skeleton point detection model respectively; Performing model format conversion on the first fall detection model, the target detection model, and the skeleton point detection model respectively, wherein the model format conversion includes at least ONNX model conversion, MLIR model conversion, and bmodel model conversion, and the quantization processing is performed after the ONNX model conversion; The second fall detection model, as well as the first fall detection model, the target detection model and the skeleton point detection model are deployed on the edge device.

7. The fall detection method according to claim 4, characterized in that: The acquiring of the data set comprises: Acquire an initial data set, wherein the initial data set includes a video frame sequence when the target falls and does not fall, and the video frame sequence is historically collected and customized by the edge device; Dividing the initial data set according to the preset time period to obtain a plurality of initial samples, wherein the initial samples include the video frame sequence within the preset time period; Each of the initial samples is input into the preset detection model to obtain the skeleton points of the same target within each preset time period and use them as the samples to construct the data set.

8. A fall detection device, characterized in that: The device comprises: A data acquisition module, used to acquire a data set, wherein the data set includes a plurality of samples respectively representing a target falling process and a non-falling process, and each of the samples includes a position of a skeletal point that changes continuously over a preset time period; A first construction module is used to construct a spatiotemporal graph convolution model, and train the spatiotemporal graph convolution model based on the data set to obtain a first fall detection model; A feature extraction module, used to extract features from each of the samples according to a preset feature processing strategy to determine target features of the sample, wherein the preset feature processing strategy includes extracting vector features of bone points and bending angles of bone points; A second construction module is used to construct a tree model, and train the tree model based on the target feature to obtain a second fall detection model; A fall detection module is used to perform fall detection in combination with the first fall detection model and the second fall detection model, and perform weighted fusion of the two detection results as a final fall detection result.

9. An electronic device, characterized in that: include: processor, memory, and communication bus; The communication bus is used to connect the processor and the memory; The processor is configured to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to make a computer execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Fall risk detection method based on multi-modal health data

    CN121682501A