Fall detection method and system based on video semantics
By employing a video semantics-based approach, this method utilizes Transformer networks and efficient channel attention mechanisms to extract structural semantic information, and combines this with support vector machines for prediction. This addresses the shortcomings of existing fall detection systems in terms of adaptability and feature selection, achieving a more accurate fall risk assessment.
Patent Information
- Application Number
- CN202411821804.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-11
AI Technical Summary
Existing fall detection systems have poor versatility when adapting to people of different ages and physical conditions, insufficient feature selection, and significant limitations of traditional methods. They cannot fully utilize semantic information about human body structure, resulting in limited predictive performance.
We employ a video semantics-based approach, utilizing Transformer networks and efficient channel attention mechanisms to extract structural semantic information, and combining this with support vector machines for prediction, thereby improving the accuracy of fall detection.
It achieves more accurate fall risk prediction, improves the system's versatility and prediction speed, adapts to the differences between different individuals, and enhances the sensitivity and predictive ability of fall risk.
Smart Images

Figure CN119649461B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of fall detection technology, and in particular relates to a fall detection method and system based on video semantics. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Fall prediction and prevention has become a key research topic in the field of public health. Although various fall prediction systems have emerged on the market, these systems still have several limitations in practical applications, such as insufficient adaptability to specific populations, such as the elderly, and inaccurate feature selection.
[0004] Existing fall detection solutions suffer from several problems: 1. Poor universality: Current fall detection methods often struggle to adapt to different age groups and physical conditions. Significant physiological differences exist between the elderly and young people, but current fall prediction systems often perform poorly in distinguishing these key differences. 2. Insufficient feature selection: Many detection methods rely on traditional feature selection methods, which may fail to fully identify and utilize features that have a significant impact on fall detection. Therefore, predictive performance may be limited, and features cannot be extracted globally or their actual utility evaluated. 3. Significant limitations of traditional methods: Traditional timed mobility tests lack accuracy. Using laboratory equipment (the gold standard) to collect fall-related data has limitations, such as inconvenience, relatively high cost, and the need for improvement in some indicators and other monitoring data. Traditional IMU devices typically cannot measure the general semantic information of fall behavior, which is crucial for fall prediction. Furthermore, the structural semantic information of the human body is a key indicator for assessing fall risk; however, the application of this important information is often neglected in current fall prediction systems. Since structural semantic information of the human body reflects the stability and coordination of human form and movement, incomplete utilization of this information may lead to inaccurate assessments of individual fall risk, thereby affecting the effective implementation of preventive measures. Therefore, developing a system capable of comprehensively utilizing structural semantic information for accurate fall risk assessment is particularly important. This can not only improve the accuracy of fall prediction but also help to adaptively develop more suitable preventive strategies for high-risk individuals.
[0005] In summary, how to effectively extract and utilize structural semantic information reflecting human morphology and movement for fall prediction, and achieve more accurate predictions, is a problem that needs to be solved. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a fall detection method and system based on video semantics. It utilizes an initial structural semantic vector and a visual input vector, and performs feature extraction, aggregation, and prediction based on a Transformer network, an efficient channel attention mechanism, and a support vector machine, thereby improving the accuracy of fall detection.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides a fall detection method based on video semantics, comprising:
[0009] Obtain video clips of the behavior of the individual to be tested;
[0010] The frame images corresponding to the behavioral video segments of the individual under test are divided into grid regions and linearly transformed to obtain the visual input vector.
[0011] The initial structural semantic vector learned through training is input together with the visual input vector into the Transformer network to obtain multi-channel structural semantic information;
[0012] By using an efficient channel attention mechanism, structural semantic information from multiple channels is assigned corresponding weights and then temporally aggregated to obtain temporally aggregated structural semantic information.
[0013] Based on the temporal aggregated structural semantic information, a support vector machine is used for prediction to obtain the detection result of whether the individual under test has fallen.
[0014] Secondly, the present invention provides a fall detection system based on video semantics, comprising:
[0015] The acquisition unit is configured to acquire video clips of the behavior of the individual to be tested.
[0016] The first feature extraction unit is configured to: divide the frame image corresponding to the behavioral video segment of the individual to be tested into a grid region and perform a linear transformation to obtain a visual input vector;
[0017] The second feature extraction unit is configured to: input the learned initial structural semantic vector obtained through training and the visual input vector together into the Transformer network to obtain multi-channel structural semantic information;
[0018] The temporal aggregation unit is configured to: use an efficient channel attention mechanism to assign corresponding weights to the multi-channel structural semantic information and perform temporal aggregation to obtain temporally aggregated structural semantic information;
[0019] The prediction unit is configured to: based on the temporal aggregated structural semantic information, use a support vector machine to predict whether the individual to be tested has fallen.
[0020] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0021] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0022] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0023] The above one or more technical solutions have the following beneficial effects:
[0024] In this invention, visual input vectors corresponding to behavioral video segments of the individual under test are extracted. A Transformer deep neural network is used to extract multi-channel structural semantic information, resulting in multi-channel structural semantic information for the individual. This multi-channel structural semantic information is then processed using a temporal aggregation method based on an efficient channel attention mechanism, followed by fall prediction using a support vector machine. Compared to traditional methods and CNN-based deep learning methods, the Transformer deep neural network method requires less computation and is faster in predicting fall risk. It also improves the efficiency, convenience, coherence, and globality of information acquisition, effectively solving the problem of lack of connection between left and right foot monitoring data and the lack of uniformity in measurement data in traditional methods. Through the efficient channel attention mechanism, structural semantic information from different times interacts, learning the correlation between structural semantic information at different times, enhancing the ability to identify predicted behaviors, and thus providing more accurate fall risk prediction.
[0025] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0026] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0027] Figure 1 This is a schematic diagram of the Transformer deep neural network structure in Embodiment 1 of the present invention;
[0028] Figure 2 This is a flowchart of the fall detection method based on video semantics in Embodiment 1 of the present invention. Detailed Implementation
[0029] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0030] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0031] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0032] Example 1
[0033] This embodiment discloses a fall detection method based on video semantics, including:
[0034] Obtain video clips of the behavior of the individual to be tested;
[0035] The frame images corresponding to the behavioral video segments of the individual under test are divided into grid regions and linearly transformed to obtain the visual input vector.
[0036] During training, a set of initial structural semantic vectors that are independent of any specific image and beneficial to the model's prediction performance are learned. The initial structural semantic vectors and the corresponding visual input vectors are input into the Transformer network to obtain multi-channel structural semantic information.
[0037] By using an efficient channel attention mechanism, structural semantic information from multiple channels is assigned corresponding weights and then temporally aggregated to obtain temporally aggregated structural semantic information.
[0038] Based on the temporal aggregated structural semantic information, a support vector machine is used for prediction to obtain the detection result of whether the individual under test has fallen.
[0039] This embodiment uses a Transformer deep neural network to extract structural semantic information from video clips of the individual's behavior. This structural semantic information is then processed using a temporal aggregation method based on an efficient channel attention mechanism, and finally, a prediction model is used for fall prediction. The semantic information obtained through video analysis not only includes parameter information found in traditional methods but also extracts hidden information that traditional methods cannot obtain. Compared to traditional methods and CNN-based deep learning methods, this embodiment uses a Transformer deep neural network, which requires less computation and is faster in predicting fall risk. It also improves the efficiency, convenience, coherence, and globality of information acquisition, effectively solving the problem of lack of connection between left and right foot monitoring data and the lack of uniformity in measurement data in traditional methods, thus providing more accurate fall risk prediction. This innovative method not only improves the reliability of fall prediction but also provides strong technical support for the safety monitoring of the elderly, further enhancing the accuracy, speed, and overall performance of the system.
[0040] The following is combined with Figures 1-2 This embodiment provides a detailed description of a fall detection method based on video semantics, specifically including:
[0041] Step 1: Obtain video clips of the behavior of the individual to be tested, divide the frame images corresponding to the video clips of the individual to be tested into grid regions, and perform linear transformation to obtain visual input vectors; input the initial structural semantic vectors learned during training and the corresponding visual input vectors into the Transformer network to obtain multi-channel structural semantic information.
[0042] In this embodiment, the method also includes preprocessing the behavioral video of the individual to be tested. The preprocessing includes video cropping, normalization, noise reduction, contrast enhancement, etc., in order to improve the accuracy of semantic information detection for further analysis and application.
[0043] In this embodiment, a visual input vector is extracted from the preprocessed behavioral video of the individual to be tested. Specifically, the two-dimensional image of the current frame is acquired at regular intervals (e.g., period T = 0.5s), with a size of [H, W]. The two-dimensional image of the current frame is divided into multiple grid regions of the same size. After linear transformation, the original video data is converted into an embedding vector in [1, d] format, which is used as the input of the Transformer deep neural network. The structural semantic information of the two-dimensional image is extracted using the Transformer deep neural network based on the self-attention mechanism.
[0044] Specifically, in the process of acquiring the visual input vector, the image is first divided into a grid of size [P_h, P_w]. The image information of each grid is flattened into a one-dimensional vector of size P_h*P_w, and the number of such vectors is:
[0045] L=(H×W) / (P_h×P_w ) (1)
[0046] L one-dimensional vectors are mapped to embedding vectors of dimension d through a linear projection function. Since human body structure semantic extraction is a position-sensitive task, positional encoding is used on the L embedding vectors to obtain the visual input vector.
[0047] The initial structural semantic vector is represented by N learnable d-dimensional embedding vectors. The initial structural semantic vector is independent of any specific image and is a set of initial data that is beneficial to the model's prediction performance.
[0048] At the start of the first round of training, these N d-dimensional embedding vectors are randomly initialized according to a Gaussian distribution. During training, these N d-dimensional embedding vectors, along with the visual input vectors of the training samples, are input into the Transformer deep neural network; during training, these N d-dimensional embedding vectors are updated. After training, the updated N d-dimensional embedding vectors are the initial structural semantic vectors learned by the model.
[0049] The input to a Transformer deep neural network consists of two parts: a visual input vector and an initial structural semantic vector.
[0050] The visual input vector is concatenated with the initial structural semantic vector and input into the Transformer deep neural network. This embodiment uses only the encoder part of the standard Transformer deep neural network, learning the structural semantic representation by stacking M encoder modules. Each module includes two components: a multi-head self-attention component and a feedforward neural network component. The output of each component is followed by residual connections and layer normalization (LayerNorm) to improve model stability and enhance generalization ability.
[0051] The output of the Transformer deep neural network is a structural semantic vector based on the original two-dimensional image of the current frame. The structural semantic information extracted from the original two-dimensional image is represented by a matrix composed of N structural semantic vectors.
[0052] For a single two-dimensional image, its structural semantic matrix is obtained through a deep neural network in the manner described above. For a video data segment, the two-dimensional image of the current frame is processed and input into the Transformer deep neural network at regular intervals (e.g., period T = 0.5s) to obtain K structural semantic matrices. These structural semantic matrices are stored in the form of tensors to form multi-channel structural semantic information.
[0053] In this embodiment, a training sample dataset is constructed, which includes video data of known fall or non-fall states. The video data is collected by capturing videos with a camera. The Transformer deep neural network is trained using this training sample dataset. During training, the structural semantic information output by the Transformer deep neural network is mapped through a linear projection layer to obtain a two-dimensional heatmap with shape [H, W]. The Transformer deep neural network is optimized by predicting the difference between the two-dimensional heatmap and the actual two-dimensional heatmap, using the mean squared error (MSE) as the loss function.
[0054] To accurately assess and prevent fall risk, comprehensive analysis using multi-channel structural semantic information is crucial. Furthermore, multi-channel structural semantic information can effectively reflect the stability of an individual's gait, thus more accurately capturing their biological structural information. By utilizing multi-channel structural semantic information, sensitivity and predictive ability regarding potential fall risk can be enhanced, providing a scientific basis for fall prevention.
[0055] Step 2: Use an efficient channel attention mechanism to assign corresponding weights to the structural semantic information of multiple channels and perform temporal aggregation to obtain the temporally aggregated structural semantic information; based on the temporally aggregated structural semantic information, use a support vector machine to make predictions to obtain the detection result of whether the individual to be tested has fallen.
[0056] The machine learning model in this embodiment consists of two parts: a temporal aggregation part based on an attention mechanism and a support vector machine classification part.
[0057] In this embodiment, the first part is temporal aggregation based on an efficient channel attention mechanism. The channel attention mechanism is used to integrate multi-channel structural semantic information extracted from video data and assign corresponding weights to the structural semantic information of different channels.
[0058] Specifically, multi-channel structural semantic information is temporally aggregated based on an attention mechanism. An efficient channel attention mechanism (ECA) module is used to learn the correlations between channels of the input feature map and assign different weights to each channel. The specific method is as follows:
[0059] Step 201: First, perform global average pooling on the multi-channel structural semantic information with K channels to obtain channel features with a scale of 1x1xK.
[0060] Step 202: Then, feature extraction is performed using a convolutional kernel with 1 channel and an adaptive size to achieve cross-channel interaction.
[0061] During convolution, appropriate padding is used to ensure that the input and output sizes are consistent.
[0062] Step 203: The output of the convolutional layer participates in the Sigmoid activation function operation, resulting in different weights assigned to each channel, thus obtaining weighted features.
[0063] Step 204: Multiply the weighted features with the multi-channel structural semantic information predicted by the Transformer deep neural network to obtain the structural semantic information aggregated by attention mechanism based on temporal aggregation.
[0064] The Efficient Channel Attention (ECA) module is particularly well-suited for processing initially extracted multi-channel structural semantic information because it effectively considers the inter-channel correlations within the input multi-channel structural semantic information. The channels of the multi-channel structural semantic information represent the time dimension, and the structural semantic matrices of different channels represent the structural semantic information at different discrete moments within a continuous time period.
[0065] This embodiment employs an efficient channel attention mechanism (ECA) module to perform temporal aggregation of multi-channel structural semantic data based on an attention mechanism. This allows structural semantic information from different time points to interact and learn the correlations between them, significantly enhancing the model's ability to identify difficult-to-predict behaviors. This, in particular, significantly improves the model's generalization ability on diverse data and its overall fall prediction performance. This approach enables the model to not only learn a wide range of data features but also to precisely adjust to cope with various complex situations, ensuring the accuracy and reliability of predictions.
[0066] The second part employs a Support Vector Machine (SVM). The temporal aggregated structural semantic information is first flattened into a vector, and then input into a trained SVM to obtain a prediction of whether a fall has occurred.
[0067] The temporally aggregated structural semantic information is flattened into a vector, which is then used as input samples into a trained Support Vector Machine (SVM). The SVM output is the final fall prediction result. The SVM maps the input samples from the input space to a new feature space using a Gaussian kernel function (RBF). Using the penalty parameter C and the kernel function parameter g, it searches for a hyperplane that separates the semantic information, thus classifying the semantic information.
[0068] Mathematical expression model of SVM model:
[0069]
[0070] sty i (ω T x i +b)≥1-ξ i ξ i ≥0 (2)
[0071] Where ω is the normal vector of the hyperplane, b is the offset of the hyperplane, and x i It is the i-th data point in the training dataset, y i It is x i The tag, ξ i is a slack variable used to handle noise or outliers in the data, and C is a penalty parameter used to balance the complexity of the model and the fit of the training data.
[0072] The mathematical expression for the Gaussian kernel function is as follows:
[0073] k(‖x1-x2‖)=exp(-g‖x1-x2‖ 2 (3)
[0074] Where g is the kernel function parameter, and g>0.
[0075] In this embodiment, various gait parameters are input into a machine learning model for predicting fall risk for training, thereby evaluating the predictive performance index of each combination and finally obtaining the trained machine learning model.
[0076] Specifically, the multi-channel structural semantic data of all training samples are input into a machine learning model for predicting fall risk to train the model. Training is stopped when the number of training iterations exceeds a set number or the total loss function value is less than a set threshold, and the trained machine learning model is obtained.
[0077] By analyzing video data and preprocessing it to extract visual input vectors, these vectors, along with initial structural semantic vectors, are fed into a Transformer deep neural network to obtain preliminary structural semantic information. This information provides crucial insights into an individual's behavioral state. Temporal aggregation based on an attention mechanism is then applied to this preliminary structural semantic information to analyze fluctuations and changes in an individual's behavior over continuous time. This analysis reveals the individual's motion characteristics and helps assess the risk of falls.
[0078] This invention extracts individual structural semantic information and performs temporal aggregation on this information based on an attention mechanism. The result of temporal aggregation reflects the stability of an individual's gait. By comprehensively utilizing this information, the system can more comprehensively capture individual behavior and improve its sensitivity to fall risk. This method of comprehensively utilizing structural semantic information and attention-based temporal aggregation makes the system more advantageous in fall prediction.
[0079] To effectively learn and adapt to individual differences and accurately predict fall risk, this invention employs a Transformer-based deep learning neural network model for extracting structural semantic information and a Support Vector Machine-based machine learning model for predicting fall risk. The deep learning neural network model adaptively adjusts the model's weights and biases, enhancing its learning ability for individuals who are difficult to predict. By introducing this deep learning model, the system of this invention can better adapt to diversity, improving overall fall prediction performance.
[0080] This embodiment provides a real-time, effective, and highly adaptable fall prediction method applicable to multiple fields such as sports biomechanics and rehabilitation medicine. This method incorporates a deep learning model, significantly improving the intelligence and accuracy of fall prediction, and providing a new technical means for research and practical applications in these fields.
[0081] This embodiment is applicable to individuals of different ages and physiological states, significantly improving the system's versatility. It has broad application prospects, particularly in areas such as elderly health management and assistive medical devices.
[0082] Example 2
[0083] The purpose of this embodiment is to provide a fall detection system based on video semantics, including:
[0084] The acquisition unit is configured to acquire video clips of the behavior of the individual to be tested.
[0085] The first feature extraction unit is configured to: divide the frame image corresponding to the behavioral video segment of the individual to be tested into a grid region and perform a linear transformation to obtain a visual input vector;
[0086] The second feature extraction unit is configured to: learn a set of initial structural semantic vectors that are independent of any specific image and are beneficial to the model's prediction effect during the training process; input the initial structural semantic vectors and the corresponding visual input vectors into the Transformer network to obtain multi-channel structural semantic information.
[0087] The temporal aggregation unit is configured to: use an efficient channel attention mechanism to assign corresponding weights to the multi-channel structural semantic information and perform temporal aggregation to obtain temporally aggregated structural semantic information;
[0088] The prediction unit is configured to: based on the temporal aggregated structural semantic information, use a support vector machine to predict whether the individual to be tested has fallen.
[0089] In further embodiments, the following is also provided:
[0090] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0091] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0092] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0093] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0094] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0095] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0096] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0097] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0098] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0099] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0100] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A fall detection method based on video semantics, characterized in that, include: Obtain video clips of the behavior of the individual to be tested; The frame images corresponding to the behavioral video segments of the individual under test are divided into grid regions and linearly transformed to obtain the visual input vector. The initial structural semantic vector learned through training is input together with the visual input vector into the Transformer network to obtain multi-channel structural semantic information; By using an efficient channel attention mechanism, structural semantic information from multiple channels is assigned corresponding weights and then temporally aggregated to obtain temporally aggregated structural semantic information. Based on the temporal aggregated structural semantic information, a support vector machine is used for prediction to obtain the detection result of whether the individual to be tested has fallen. The input to the Transformer network consists of two parts. The Transformer network only uses the encoder part of the standard Transformer network. It learns the structural semantic representation by stacking M encoder modules. Each module includes two components: a multi-head self-attention component and a feedforward neural network component. The output of each component is followed by a residual connection and layer normalization structure. Specifically, the initial structural semantic vector obtained from the learning is as follows: N learnable d-dimensional embedding vectors are randomly initialized according to a Gaussian distribution to represent the initial structural semantic vector; During the training of the fall detection model, training samples and initial structural semantic vectors are input into the fall detection model. After training, the initial structural semantic vectors learned by the fall detection model are obtained. By employing an efficient channel attention mechanism to assign appropriate weights to the multi-channel structural semantic information and performing temporal aggregation, the temporally aggregated structural semantic information is obtained, specifically: Global average pooling is applied to the structural semantic information of multiple channels to obtain channel features; The channel features are extracted using convolution to obtain cross-channel interactive features; The cross-channel interaction features are calculated using an activation function to obtain weighted features; Multiplying the weighted features with the multi-channel structural semantic information yields temporally aggregated structural semantic information; Based on the temporal aggregated structural semantic information, a support vector machine is used for prediction to obtain the detection result of whether the individual to be tested has fallen, specifically: The structural semantic information of the temporal aggregation is mapped to the feature space using a Gaussian kernel function; Using penalty parameters and kernel function parameters, an optimal hyperplane is constructed to separate data information, classify the temporal aggregated structural semantic information, and obtain the detection result of whether the individual under test has fallen.
2. The fall detection method based on video semantics as described in claim 1, characterized in that, The frame images corresponding to the behavioral video segments of the individual under test are divided into grid regions, and a linear transformation is performed to obtain the visual input vector, specifically: The frame images corresponding to the behavioral video segments of the individual under test are divided into grid regions, and the image information of each grid region is flattened into a one-dimensional vector. The one-dimensional vector corresponding to each grid region is mapped to an embedding vector through a linear projection function; The embedding vector is encoded to obtain the visual input vector.
3. The fall detection method based on video semantics as described in claim 1, characterized in that, The initial structural semantic vector and the corresponding visual input vector are input into the Transformer network to obtain multi-channel structural semantic information, specifically: The structural semantic matrix is obtained by using a Transformer encoder to extract features from the visual input vector corresponding to the frame image and the learned initial structural semantic vector. The structural semantic matrix of each frame image is stored in tensor form to form multi-channel structural semantic information.
4. A fall detection system based on video semantics, characterized in that, include: The acquisition unit is configured to acquire video clips of the behavior of the individual to be tested. The first feature extraction unit is configured to: divide the frame image corresponding to the behavioral video segment of the individual to be tested into a grid region and perform a linear transformation to obtain a visual input vector; The second feature extraction unit is configured to: input the learned initial structural semantic vector obtained through training and the visual input vector together into the Transformer network to obtain multi-channel structural semantic information; the input of the Transformer network includes two parts, the Transformer network only uses the encoder part of the standard Transformer network, and learns the structural semantic representation by stacking M encoder modules, each module includes two components: a multi-head self-attention component and a feedforward neural network component; the output of each component is followed by a residual connection and layer normalization structure; The initial structural semantic vector obtained from the learning is as follows: N learnable d-dimensional embedding vectors are randomly initialized according to a Gaussian distribution to represent the initial structural semantic vector; During the training of the fall detection model, training samples and initial structural semantic vectors are input into the fall detection model. After training, the initial structural semantic vectors learned by the fall detection model are obtained. The temporal aggregation unit is configured to: aggregate multi-channel structural semantic information temporally by assigning appropriate weights to the multi-channel structural semantic information using an efficient channel attention mechanism, thereby obtaining temporally aggregated structural semantic information; specifically: Global average pooling is applied to the structural semantic information of multiple channels to obtain channel features; The channel features are extracted using convolution to obtain cross-channel interactive features; The cross-channel interaction features are calculated using an activation function to obtain weighted features; Multiplying the weighted features with the multi-channel structural semantic information yields temporally aggregated structural semantic information; The prediction unit is configured to: based on the temporally aggregated structural semantic information, use a support vector machine to predict whether the individual to be tested has fallen; specifically: The structural semantic information of the temporal aggregation is mapped to the feature space using a Gaussian kernel function; Using penalty parameters and kernel function parameters, an optimal hyperplane is constructed to separate data information, classify the temporal aggregated structural semantic information, and obtain the detection result of whether the individual under test has fallen.
5. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-3.
7. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-3.
Citation Information
Patent Citations
Video content description method, system and device based on multi-modal attention mechanism
CN111079601A
Video object positioning method and system based on mixed attention mechanism
CN113971208A