Unsupervised tumble detection method and device based on variational auto-encoder

Through the unsupervised training method of the variational autoencoder, a sequence reconstruction model is constructed using the data of the elderly's normal behavior, and the falls are detected in real time, which solves the problems of poor wear compliance and high deployment costs in the existing technology, and achieves high-precision and low-cost fall detection.

CN120496192AInactive Publication Date: 2025-08-15HANGZHOU INSHINE INTELLIGENT TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510994219.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing fall detection technology has problems such as poor wear compliance, high deployment cost, insensitive to local abnormalities and insufficient generalization ability in the care scenarios of the elderly, especially in the home environment. It is difficult to achieve efficient fall detection.

Method used

Unsupervised training is performed using a variational autoencoder. By obtaining the normal behavior data of the elderly at home, a sequence reconstruction model is constructed, and the fall behavior is detected using reconstruction errors, including the encoder layer, the reparameter sampling layer, the regional attention mechanism layer and the decoder layer. The video is monitored in real time and the reconstruction error is calculated to determine the fall.

Benefits of technology

It realizes fall detection that adapts to different indoor environments and individual characteristics without labeling data, has high accuracy and low deployment costs, and can detect fall events in a timely manner and reduce damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496192A_ABST
    Figure CN120496192A_ABST
Patent Text Reader

Abstract

The invention provides an unsupervised tumble detection method and device based on a variational auto-encoder, and the method comprises the following steps: obtaining a monitoring video which does not contain a tumble behavior in a monitoring region, carrying out the detection of the coordinates of human body key points of each video frame of the monitoring video, and obtaining a plurality of groups of key point sequences, taking the multiple groups of key point sequences as a training data set; training a preset sequence reconstruction framework by using the training data set to obtain a sequence reconstruction model; the method comprises the following steps: acquiring a to-be-detected monitoring video in a monitoring area in real time, acquiring a to-be-detected key point sequence based on the to-be-detected monitoring video, reconstructing the to-be-detected key point sequence by using a sequence reconstruction model to obtain a to-be-detected reconstructed key point sequence, and if a reconstruction error between the to-be-detected reconstructed key point sequence and the to-be-detected key point sequence is greater than a set threshold value, determining that the to-be-detected key point sequence and if so, determining that falling occurs. According to the scheme, daily behaviors of old people are reconstructed in real time through the trained variational auto-encoder, and abnormal behaviors such as tumble are detected by calculating reconstruction errors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to an unsupervised fall detection method and device using a variational autoencoder. Background Art

[0002] With the global aging population, falls among the elderly have become a major public health challenge. According to the World Health Organization, approximately 300,000 elderly people die each year from falls worldwide. In China, the annual fall rate for people over 65 years old exceeds 30%. Complications such as fractures and brain damage caused by falls not only significantly reduce quality of life but also result in an annual medical burden of hundreds of billions of yuan. This is especially true in the care of elderly people living alone. Due to the lack of real-time monitoring, falls are often difficult to detect in a timely manner, further exacerbating the consequences of injuries. This harsh social reality has driven the urgent need for fall detection technology.

[0003] The current mainstream fall detection technologies can be roughly divided into two major camps: hardware sensing and vision. In the hardware sensing camp, although wearable devices based on accelerometers have high detection accuracy in laboratory environments, in real elderly care scenarios, user compliance is less than 15%. The elderly often refuse to use the devices due to the sense of restraint, cumbersome charging or aesthetic issues, resulting in a significant reduction in monitoring coverage. Mattress or floor solutions based on pressure sensing have defects such as fuzzy spatial positioning and insensitivity to slow falls, making it difficult to meet the time requirements for emergency intervention.

[0004] Traditional approaches in the field of vision rely on supervised learning frameworks, such as pose classifiers based on HOG features combined with support vector machines (SVMs), or direct fall posture identification using object detection networks like YOLO. The core drawback of these approaches lies in their strong reliance on labeled data. Firstly, real falls are sporadic and dangerous, making it difficult to collect a large and diverse sample of falls on a large scale. Secondly, fall postures (such as forward falls, side falls, and backward falls) vary significantly between individuals, and furniture occlusion and lighting variations in home environments can further exacerbate data distribution bias. Previous research has shown that the detection accuracy of a supervised model deployed in a nursing home plummeted from 85% to 42% during cross-scenario testing, highlighting the critical generalization shortcomings of traditional approaches.

[0005] Therefore, there is an urgent need for an unsupervised fall detection method that is individual adaptable, has low deployment cost, does not require fall samples, and is sensitive to local anomalies. Summary of the Invention

[0006] The embodiments of the present application provide an unsupervised fall detection method and device based on a variational autoencoder, which performs unsupervised training on the variational autoencoder using the normal behavior of the elderly at home. The trained variational autoencoder reconstructs the elderly's daily behavior in real time, and detects falls by calculating the reconstruction error.

[0007] In a first aspect, an embodiment of the present application provides an unsupervised fall detection method based on a variational autoencoder, the method comprising: Obtain surveillance videos that do not contain falls in the surveillance area, perform human key point coordinate detection on each video frame of the surveillance video to obtain multiple sets of key point sequences, and use the multiple sets of key point sequences as training data sets; A sequence reconstruction model is obtained by training a preset sequence reconstruction architecture using a training data set, wherein the sequence reconstruction architecture includes an encoder layer, a reparameter sampling layer, a regional attention mechanism layer, and a decoder layer. The encoder layer extracts features from each key point in the key point sequence to obtain a coded feature sequence. The reparameter sampling layer performs random noise sampling on the coded feature sequence to obtain a latent feature sequence. The regional attention mechanism layer assigns a weight to each key point in the latent feature sequence to obtain a weighted feature sequence. The decoder layer decodes and reconstructs the weighted feature sequence to obtain a reconstructed feature sequence. The reconstruction error between the reconstructed feature sequence and the corresponding key point sequence and the KL divergence of the coded feature sequence are used as the loss function of the sequence reconstruction architecture. The monitoring video to be tested is acquired in real time within the monitoring area, the coordinates of human body key points in the monitoring video to be tested are detected to obtain a key point sequence to be tested, and the key point sequence to be tested is reconstructed using a sequence reconstruction model to obtain a reconstructed key point sequence to be tested. If the reconstruction error between the reconstructed key point sequence to be tested and the key point sequence to be tested is greater than a set threshold, it is determined that a fall has occurred.

[0008] In a second aspect, an embodiment of the present application provides an unsupervised fall detection device based on a variational autoencoder, comprising: An acquisition module is used to obtain surveillance videos that do not contain falls in the surveillance area, perform human key point coordinate detection on each video frame of the surveillance video to obtain multiple sets of key point sequences, and use the multiple sets of key point sequences as training data sets; A training module is used to use a training data set to train a preset sequence reconstruction architecture to obtain a sequence reconstruction model, wherein the sequence reconstruction architecture includes an encoder layer, a reparameter sampling layer, a regional attention mechanism layer, and a decoder layer. The encoder layer extracts features from each key point in the key point sequence to obtain a coded feature sequence, the reparameter sampling layer performs random noise sampling on the coded feature sequence to obtain a latent feature sequence, the regional attention mechanism layer assigns a weight to each key point in the latent feature sequence to obtain a weighted feature sequence, and the decoder layer decodes and reconstructs the weighted feature sequence to obtain a reconstructed feature sequence, and uses the reconstruction error between the reconstructed feature sequence and the corresponding key point sequence and the KL divergence of the coded feature sequence as the loss function of the sequence reconstruction architecture; The detection module is used to obtain the monitoring video to be tested in real time within the monitoring area, perform human body key point coordinate detection on the monitoring video to be tested to obtain a key point sequence to be tested, and reconstruct the key point sequence to be tested using a sequence reconstruction model to obtain a reconstructed key point sequence to be tested. If the reconstruction error between the reconstructed key point sequence to be tested and the key point sequence to be tested is greater than a set threshold, it is determined that a fall has occurred.

[0009] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform an unsupervised fall detection method based on a variational autoencoder.

[0010] In a fourth aspect, an embodiment of the present application provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process. When the program code is executed by a processor, an unsupervised fall detection method based on a variational autoencoder is implemented.

[0011] The main contributions and innovations of the present invention are as follows: Embodiments of the present application.

[0012] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 is a flow chart of an unsupervised fall detection method according to an embodiment of the present application; Figure 2 is a structural diagram of a sequence reconstruction architecture according to an embodiment of the present application; Figure 3 is a schematic diagram of determining a fall according to an embodiment of the present application; Figure 4 is a structural block diagram of an unsupervised fall detection device according to an embodiment of the present application; Figure 5 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0014] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.

[0015] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0016] Example 1 The embodiment of the present application provides an unsupervised fall detection method based on a variational autoencoder. This solution uses the normal behavior of the elderly at home to perform unsupervised training on the variational autoencoder. The trained variational autoencoder reconstructs the elderly's daily behavior in real time and detects falls by calculating the reconstruction error. Specifically, refer to Figure 1 , the method comprising: Obtain surveillance videos that do not contain falls in the surveillance area, perform human key point coordinate detection on each video frame of the surveillance video to obtain multiple sets of key point sequences, and use the multiple sets of key point sequences as training data sets; A sequence reconstruction model is obtained by training a preset sequence reconstruction architecture using a training data set, wherein the sequence reconstruction architecture includes an encoder layer, a reparameter sampling layer, a regional attention mechanism layer, and a decoder layer. The encoder layer extracts features from each key point in the key point sequence to obtain a coded feature sequence. The reparameter sampling layer performs random noise sampling on the coded feature sequence to obtain a latent feature sequence. The regional attention mechanism layer assigns a weight to each key point in the latent feature sequence to obtain a weighted feature sequence. The decoder layer decodes and reconstructs the weighted feature sequence to obtain a reconstructed feature sequence. The reconstruction error between the reconstructed feature sequence and the corresponding key point sequence and the KL divergence of the coded feature sequence are used as the loss function of the sequence reconstruction architecture. The monitoring video to be tested is acquired in real time within the monitoring area, the coordinates of human body key points in the monitoring video to be tested are detected to obtain a key point sequence to be tested, and the key point sequence to be tested is reconstructed using a sequence reconstruction model to obtain a reconstructed key point sequence to be tested. If the reconstruction error between the reconstructed key point sequence to be tested and the key point sequence to be tested is greater than a set threshold, it is determined that a fall has occurred.

[0017] In some specific embodiments, in order to ensure that this solution has all-weather, continuous monitoring capabilities, this solution uses an embedded IPC camera to obtain surveillance video within the monitoring area. The optional resolution is 1280×960 and the frame rate is set to 25FPS to ensure that continuous human movement details are captured. The camera installation position should be arranged in combination with the indoor layout and installed in the indoor corner or wall corner. The recommended height from the ground is 1.8 to 2.2 meters to form a certain bird's-eye view. The lens focal length can be selected as a wide-angle lens to enhance the coverage.

[0018] In some specific embodiments, the present solution uses the pose estimation model YOLOv8n-pose to detect the coordinates of human key points in each video frame of the surveillance video. The pose estimation model can extract the two-dimensional key point coordinates of the human body in real time on an image with a standard input size of 480×640, supports detecting the key point positions of a single person or multiple people, and outputs the detection frame of the human body in each frame image, a set of human key point coordinates and the confidence score of each key point.

[0019] Specifically, the key points detected by this solution include 17 key point positions such as head, shoulders, elbows, wrists, hips, knees, and ankles.

[0020] Furthermore, in the step of detecting the coordinates of human key points in each video frame of the surveillance video to obtain original key point data, the original key point data are sequentially normalized and multiple groups of key point sequences are obtained through screening and segmentation, wherein the original key point data of each video frame are linearly mapped to a preset range through normalization, screening is used to remove video frames whose average confidence of key points in the video frames is less than the screening threshold and video frames with human body occlusion, and segmentation is used to segment continuous video frames into multiple groups of key point sequences in chronological order.

[0021] Specifically, when performing normalization preprocessing on the original key points, the key point coordinates in each video frame are first normalized according to the image width and height, and linear interpolation or inter-frame completion is performed on the key points with confidence lower than the completion threshold or the lost key points to enhance the continuity of the original key point data, and then the original key point data of each video frame is linearly mapped to the preset range.

[0022] Specifically, during the screening process, video frames in which the average confidence of key points is less than the screening threshold and video frames with human body occlusion are removed, and only video frames in which the human body is not obstructed in the picture and the key point confidence is relatively high overall are retained, thereby improving the accuracy of subsequent model training.

[0023] Specifically, in the segmentation, a fixed time window method is used to segment the surveillance video in chronological order to obtain multiple groups of key point sequences. In this scheme, 30 consecutive video frames are collected as a group of key point sequences through a fixed time window at the original key point. Each group of key point sequences includes 17 key points, and each key point is a two-dimensional coordinate, so a group of key point sequences is a 30-frame × 34-dimensional = 1020-dimensional time series vector.

[0024] This solution first uses a camera to collect surveillance videos for a period of time in the monitoring area, and then generates a training data set by performing human key point detection and normalization, screening, segmentation and other pre-processing on the surveillance video, and uses the training data set to train the preset sequence reconstruction architecture. In other words, before the fall detection method of this solution is applied, it is necessary to re-acquire the training data set according to different monitoring areas and different monitoring personnel, and re-train the sequence reconstruction architecture, rather than the traditional artificial intelligence model that completes the training of all scenarios. This is because the application scenarios of this solution are generally indoor scenes such as living rooms and bedrooms, and the interior decoration of each family will be different, and the health conditions of the elderly in each family are also different. Therefore, before performing fall detection, this solution will first collect surveillance videos that do not contain fall behaviors to fully collect the normal behavior data of the monitoring personnel as training data. The collection cycle of this solution is 48 hours.

[0025] In other words, this solution collects the normal behavior sequences of the monitored personnel in an unlabeled manner, so that the modeling samples can be automatically adjusted according to the physical characteristics and living habits of the monitored personnel to perform unsupervised training on the sequence reconstruction architecture, so that the trained sequence reconstruction model can adapt to different indoor environment layouts and camera angles.

[0026] In some embodiments, the structure diagram of the sequence reconstruction architecture is as follows Figure 2 As shown, in this scheme, the training of the sequence reconstruction architecture is carried out in an unsupervised manner. The encoder layer of the sequence reconstruction architecture includes an upper body branch, a lower body branch and a full body branch. The upper body branch is used to extract features of the upper body key points in the key point sequence, the lower body branch is used to extract features of the lower body key points in the key point sequence, and the full body branch is used to extract features of the full body key points in the key point sequence. The feature extraction results of the upper body branch, the lower body branch and the full body branch are weightedly fused to obtain a weighted fusion result, and the weighted fusion result is output using a fully connected layer to obtain a coded feature sequence.

[0027] Specifically, the upper body branch in this scheme is used to extract key features of the upper body, such as the head, shoulders, and arms; the lower body branch is used to extract key features of the lower body, such as the hips, knees, and ankles; and the whole body branch is used to extract key features of the whole body, in order to respectively model the whole body's motion coordination. The upper body branch, lower body branch, and whole body branch in this scheme have the same structure, all consisting of a three-layer one-dimensional convolutional network. In each branch, the first layer is used to preliminarily extract a small change model of each key point in the time series, and the output is (batch, 32, T). The convolution kernel size and stride of the first layer are as follows:

[0028] The second layer is used to enhance the perception of temporal context. The output is (batch, 64, T / 2). The convolution kernel size and stride of the second layer are as follows:

[0029] The third layer is used to extract long-term dependent motion trend information. The output is: The convolution kernel size and step size of the third layer are as follows:

[0030] Among them, the number of convolution channels in the first layer is 32, the number of convolution channels in the second layer is 64, and the number of convolution channels in the third layer is 128.

[0031] Furthermore, the formula for parallel feature extraction of the upper body branch, lower body branch, and whole body branch is expressed as follows:

[0032]

[0033]

[0034] in, Features extracted for the upper body branch, Features extracted for the lower body branches, The formula for weighted fusion of the features extracted from the whole body branch, the upper body branch, the lower body branch, and the whole body branch is as follows:

[0035] in, is the weighted fusion result, Features extracted for the upper body branch The weight of Features extracted for the lower body branches The weight of Features extracted for the whole body branch The weights of the weights, where the weight coefficients α, β, and γ can be set as learnable parameters or fixed values according to the actual situation. The weighted fusion results are obtained by using the fully connected layer Linear (128 * T / 4 → 64). The output is the encoded feature sequence h.

[0036] Furthermore, the encoded feature sequence is input into two different fully connected layers to obtain the μ vector and logσ² vector, which can be expressed as follows: μ = Linear_μ( ) logσ² = Linear_logvar( ); Among them, Linear_μ is a fully connected layer that outputs the mean of the latent variable, Linear_logvar is a fully connected layer that outputs the logarithm of the variance of the latent variable, and μ and logσ² together define the Gaussian distribution of the encoded feature sequence h in the latent space.

[0037] The formula for random noise sampling of μ vector and logσ² vector is as follows: z = μ + σ * ε Where z is the potential feature sequence, ε is the noise vector that follows the standard normal distribution, σ = exp(0.5 * logσ²), μ is the mean vector obtained based on the encoded feature sequence, and logσ² is the logarithm vector of the variance obtained based on the encoded feature sequence.

[0038] In some specific embodiments, the inter-frame velocity estimation and inter-frame acceleration estimation are calculated for each key point in the potential feature sequence in the regional attention mechanism layer to obtain velocity features and acceleration features, the coordinates, velocity features, and acceleration features of each key point are input into the Transformer encoder to obtain the attention weight of each key point, and each key point in the potential feature sequence is weighted with the corresponding attention weight to obtain a weighted feature sequence.

[0039] Specifically, the coordinates of each key point in each frame of the potential feature sequence are obtained ( , ), the formula for inter-frame velocity estimation is expressed as:

[0040] in, is the inter-frame velocity of the i-th key point in the t-th video frame, is the coordinate of the i-th key point in the t-th video frame, is the coordinate of the i-th key point in the t-1-th video frame.

[0041] Specifically, the formula for inter-frame acceleration is expressed as:

[0042] in, is the inter-frame acceleration of key point i, is the inter-frame velocity of the i-th key point in the t-th video frame, is the inter-frame velocity of the i-th key point in the t-1-th video frame.

[0043] Furthermore, the velocity feature and acceleration feature of each key point are normalized and then concatenated with the key point coordinates to obtain the weighted input feature. , and then input the weight into the feature Input into the Transformer encoder to obtain the attention weight of each key point. The structure of the Transformer encoder is expressed as:

[0044] Specifically, the attention weight of each key point By assigning attention weights, the model can pay more attention to important key points, such as the head and buttocks, which helps to model the potential behavioral changes of different key points.

[0045] Furthermore, an inverse weighted reconstruction error is introduced at the decoder output: , , that is, high-weight key points, whose reconstruction error accounts for a larger proportion of the total loss, guiding the model to focus on the accurate restoration of key parts.

[0046] In some embodiments, the loss function for the sequence architecture is:

[0047] in, is the loss function of the sequence architecture, is the reconstruction error between the reconstructed feature sequence and the corresponding key point sequence, , , that is, high-weight key points, whose reconstruction error accounts for a larger proportion of the total loss, guiding the model to focus on the accurate restoration of key parts, among which, is the acceleration of the i-th key point, is the key point in the key point sequence, and K is the total number of key points. is the KL divergence of the potential feature sequence, , where k is the dimension of the potential feature sequence, μ is the mean vector obtained based on the encoded feature sequence, logσ² is the logarithm vector of the variance obtained based on the encoded feature sequence, and σ = exp(0.5 *logσ²).

[0048] Specifically, during the training process of the sequence reconstruction architecture, small batches of data are input, and the above-mentioned total loss function is iteratively optimized through the optimizer. In each round of training, the convergence trend of the total loss value is monitored to ensure that the model gradually learns the normal behavioral distribution characteristics of the input data.

[0049] In some specific embodiments, when the sequence reconstruction architecture completes training, it enters the detection loop stage. In the detection stage, the monitoring video to be tested in the monitoring area is obtained in real time, and the coordinates of human key points in the monitoring video to be tested are detected to obtain the key point sequence to be tested. It is worth mentioning that the number of video frames in the key point sequence to be tested in this scheme is the same as the number of video frames in the key point sequence used during training, that is, the key point sequence to be tested obtained in this scheme includes 30 consecutive video frames.

[0050] Specifically, in this scheme, the reconstructed key point sequence to be tested obtained by the sequence reconstruction model is compared with the key point sequence to be tested, and the reconstruction error is calculated. If the reconstruction error is greater than the set threshold, it means that the motion pattern of the current key point sequence cannot be well restored by the model, which may indicate that the action is an abnormal behavior, such as falling.

[0051] Specifically, since the sequence reconstruction model in this solution is built based on a variational autoencoder and only uses the user's normal activity data during the training phase, the encoder learns the probability distribution of normal behavior in the latent space, and the decoder learns to sample from this distribution and reconstruct the original key point sequence. At this time, the model's reconstruction error for normal behavior is extremely small. When a fall occurs, the human body posture differs greatly from the latent distribution of normal behavior. At this time, the input key point data cannot be effectively encoded and decoded by the trained sequence reconstruction model, resulting in a significant deviation between the reconstructed key point coordinates and the original input, and the reconstruction error will far exceed the preset threshold.

[0052] In some specific embodiments, the schematic diagram for determining a fall is as follows: Figure 3 As shown, when a fall is determined to have occurred, the fall alarm mechanism is immediately triggered. The fall alarm mechanism includes sending text messages, APP notifications or server push notifications to relevant personnel, local voice broadcasts to remind people around the user, etc.

[0053] Example 2 Based on the same concept, refer to Figure 4 , this application also proposes an unsupervised fall detection device based on variational autoencoder, comprising: An acquisition module is used to obtain surveillance videos that do not contain falls in the surveillance area, perform human key point coordinate detection on each video frame of the surveillance video to obtain multiple sets of key point sequences, and use the multiple sets of key point sequences as training data sets; A training module is used to use a training data set to train a preset sequence reconstruction architecture to obtain a sequence reconstruction model, wherein the sequence reconstruction architecture includes an encoder layer, a reparameter sampling layer, a regional attention mechanism layer, and a decoder layer. The encoder layer extracts features from each key point in the key point sequence to obtain a coded feature sequence, the reparameter sampling layer performs random noise sampling on the coded feature sequence to obtain a latent feature sequence, the regional attention mechanism layer assigns a weight to each key point in the latent feature sequence to obtain a weighted feature sequence, and the decoder layer decodes and reconstructs the weighted feature sequence to obtain a reconstructed feature sequence, and uses the reconstruction error between the reconstructed feature sequence and the corresponding key point sequence and the KL divergence of the coded feature sequence as the loss function of the sequence reconstruction architecture; The detection module is used to obtain the monitoring video to be tested in real time within the monitoring area, perform human body key point coordinate detection on the monitoring video to be tested to obtain a key point sequence to be tested, and reconstruct the key point sequence to be tested using a sequence reconstruction model to obtain a reconstructed key point sequence to be tested. If the reconstruction error between the reconstructed key point sequence to be tested and the key point sequence to be tested is greater than a set threshold, it is determined that a fall has occurred.

[0054] Example 3 This embodiment also provides an electronic device, referring to Figure 5 , includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.

[0055] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0056] Memory 404 may include a large-capacity memory 404 for data or instructions. By way of example, and not limitation, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to the data processing device. In certain embodiments, memory 404 is non-volatile memory. In certain embodiments, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. In appropriate circumstances, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory 404 (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0057] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .

[0058] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any one of the unsupervised fall detection methods based on variational autoencoders in the above embodiments.

[0059] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .

[0060] Transmission device 406 can be used to receive or transmit data via a network. Specific examples of such networks may include wired or wireless networks provided by the electronic device's communications provider. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0061] The input and output devices 408 are used to input or output information. In this embodiment, the input information may be a training data set, a surveillance video to be tested, etc., and the output information may be a fall detection result, etc.

[0062] Optionally, in this embodiment, the processor 402 may be configured to execute the following steps through a computer program: Obtain surveillance videos that do not contain falls in the surveillance area, perform human key point coordinate detection on each video frame of the surveillance video to obtain multiple sets of key point sequences, and use the multiple sets of key point sequences as training data sets; A sequence reconstruction model is obtained by training a preset sequence reconstruction architecture using a training data set, wherein the sequence reconstruction architecture includes an encoder layer, a reparameter sampling layer, a regional attention mechanism layer, and a decoder layer. The encoder layer extracts features from each key point in the key point sequence to obtain a coded feature sequence. The reparameter sampling layer performs random noise sampling on the coded feature sequence to obtain a latent feature sequence. The regional attention mechanism layer assigns a weight to each key point in the latent feature sequence to obtain a weighted feature sequence. The decoder layer decodes and reconstructs the weighted feature sequence to obtain a reconstructed feature sequence. The reconstruction error between the reconstructed feature sequence and the corresponding key point sequence and the KL divergence of the coded feature sequence are used as the loss function of the sequence reconstruction architecture. The monitoring video to be tested is acquired in real time within the monitoring area, the coordinates of human body key points in the monitoring video to be tested are detected to obtain a key point sequence to be tested, and the key point sequence to be tested is reconstructed using a sequence reconstruction model to obtain a reconstructed key point sequence to be tested. If the reconstruction error between the reconstructed key point sequence to be tested and the key point sequence to be tested is greater than a set threshold, it is determined that a fall has occurred.

[0063] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.

[0064] In general, various embodiments may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0065] The embodiments of the present invention may be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components that are configured to perform an embodiment when the program is run. One or more computer executable components may be at least one software code or a portion thereof. In addition, it should be noted at this point that, for example, Figure 5 Any block of the logic flow in the program may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. Physical media are non-transitory media.

[0066] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0067] The above embodiments merely illustrate several embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. An unsupervised fall detection method based on variational autoencoder, characterized in that: The following steps are involved: Obtain surveillance videos that do not contain falls in the surveillance area, perform human key point coordinate detection on each video frame of the surveillance video to obtain multiple sets of key point sequences, and use the multiple sets of key point sequences as training data sets; A sequence reconstruction model is obtained by training a preset sequence reconstruction architecture using a training data set, wherein the sequence reconstruction architecture includes an encoder layer, a reparameter sampling layer, a regional attention mechanism layer, and a decoder layer. The encoder layer extracts features from each key point in the key point sequence to obtain a coded feature sequence. The reparameter sampling layer performs random noise sampling on the coded feature sequence to obtain a latent feature sequence. The regional attention mechanism layer assigns a weight to each key point in the latent feature sequence to obtain a weighted feature sequence. The decoder layer decodes and reconstructs the weighted feature sequence to obtain a reconstructed feature sequence. The reconstruction error between the reconstructed feature sequence and the corresponding key point sequence and the KL divergence of the coded feature sequence are used as the loss function of the sequence reconstruction architecture. The monitoring video to be tested is acquired in real time within the monitoring area, the coordinates of human body key points in the monitoring video to be tested are detected to obtain a key point sequence to be tested, and the key point sequence to be tested is reconstructed using a sequence reconstruction model to obtain a reconstructed key point sequence to be tested. If the reconstruction error between the reconstructed key point sequence to be tested and the key point sequence to be tested is greater than a set threshold, it is determined that a fall has occurred.

2. The unsupervised fall detection method based on variational autoencoder according to claim 1 is characterized in that In the step of detecting the coordinates of human key points on each video frame of the surveillance video to obtain original key point data, the original key point data are sequentially normalized and multiple groups of key point sequences are obtained through screening and segmentation, wherein the original key point data of each video frame are linearly mapped to a preset range through normalization, screening is used to remove video frames whose average confidence of key points in the video frames is less than a screening threshold and video frames with human body occlusion, and segmentation is used to segment continuous video frames into multiple groups of key point sequences in chronological order.

3. The unsupervised fall detection method based on variational autoencoder according to claim 1 is characterized in that The encoder layer of the sequence reconstruction architecture includes an upper body branch, a lower body branch, and a full body branch. The upper body branch is used to extract features of the upper body key points in the key point sequence, the lower body branch is used to extract features of the lower body key points in the key point sequence, and the full body branch is used to extract features of the full body key points in the key point sequence. The feature extraction results of the upper body branch, the lower body branch, and the full body branch are weightedly fused to obtain a weighted fusion result, and the weighted fusion result is output using a fully connected layer to obtain a coded feature sequence.

4. The unsupervised fall detection method based on variational autoencoder according to claim 1, characterized in that The formula for random noise sampling of the coding feature sequence is as follows: z = μ + σ * ε; Where z is the potential feature sequence, ε is the noise vector that follows the standard normal distribution, σ = exp(0.5 * logσ²), μ is the mean vector obtained based on the encoded feature sequence, and logσ² is the logarithm vector of the variance obtained based on the encoded feature sequence.

5. The unsupervised fall detection method based on variational autoencoder according to claim 1, characterized in that In the regional attention mechanism layer, the inter-frame velocity estimation and inter-frame acceleration estimation are calculated for each key point in the potential feature sequence to obtain the velocity feature and acceleration feature. The coordinates, velocity feature and acceleration feature of each key point are input into the Transformer encoder to obtain the attention weight of each key point. Each key point in the potential feature sequence is weighted with the corresponding attention weight to obtain a weighted feature sequence.

6. The unsupervised fall detection method based on variational autoencoder according to claim 1, characterized in that The loss function for the sequence architecture is: ; in, is the loss function of the sequence architecture, is the reconstruction error between the reconstructed feature sequence and the corresponding key point sequence, , where d is the total number of key points, i is the key point, is the key point in the key point sequence, To reconstruct the key points in the feature sequence, is the KL divergence of the potential feature sequence, , where k is the dimension of the potential feature sequence, μ is the mean vector obtained based on the encoded feature sequence, logσ² is the logarithm vector of the variance obtained based on the encoded feature sequence, and σ = exp(0.5 *logσ²).

7. The unsupervised fall detection method based on variational autoencoder according to claim 1, characterized in that The number of video frames in the key point sequence to be tested is the same as the number of video frames in the key point sequence used during training.

8. An unsupervised fall detection device based on variational autoencoder, characterized in that: include: An acquisition module is used to obtain surveillance videos that do not contain falls in the surveillance area, perform human key point coordinate detection on each video frame of the surveillance video to obtain multiple sets of key point sequences, and use the multiple sets of key point sequences as training data sets; A training module is used to use a training data set to train a preset sequence reconstruction architecture to obtain a sequence reconstruction model, wherein the sequence reconstruction architecture includes an encoder layer, a reparameter sampling layer, a regional attention mechanism layer, and a decoder layer. The encoder layer extracts features from each key point in the key point sequence to obtain a coded feature sequence, the reparameter sampling layer performs random noise sampling on the coded feature sequence to obtain a latent feature sequence, the regional attention mechanism layer assigns a weight to each key point in the latent feature sequence to obtain a weighted feature sequence, and the decoder layer decodes and reconstructs the weighted feature sequence to obtain a reconstructed feature sequence, and uses the reconstruction error between the reconstructed feature sequence and the corresponding key point sequence and the KL divergence of the coded feature sequence as the loss function of the sequence reconstruction architecture; The detection module is used to obtain the monitoring video to be tested in real time within the monitoring area, perform human body key point coordinate detection on the monitoring video to be tested to obtain a key point sequence to be tested, and reconstruct the key point sequence to be tested using a sequence reconstruction model to obtain a reconstructed key point sequence to be tested. If the reconstruction error between the reconstructed key point sequence to be tested and the key point sequence to be tested is greater than a set threshold, it is determined that a fall has occurred.

9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to perform the unsupervised fall detection method based on a variational autoencoder according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process. When the program code is executed by a processor, an unsupervised fall detection method based on a variational autoencoder according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Abnormal action determination method and device, electronic equipment and storage medium

    CN113392742A

  • Fall detection method, system and equipment based on millimeter wave radar

    CN116840835A

  • Track prediction method based on space-time attention mechanism and mental differential equation

    CN117077727A

  • Radar interference identification method based on LSTM and variational auto-encoder

    CN119226928A