Fall detection method and system based on multi-direction gabor convolution network

By combining dual-view fusion and multi-directional Gabor convolutional fusion with YOLOv12 Neck and Decoupled Head modules, the false alarm and false negative problems of fall detection in complex scenes are solved, and higher detection accuracy is achieved.

CN120877389BActive Publication Date: 2025-12-09GENERAL HOSPITAL OF NUCLEAR IND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511383349.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-09
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing fall detection methods have a high rate of false alarms or false alarms in complex scenarios such as diverse target postures, changing lighting, and obstacle occlusion, making it difficult to achieve effective real-time monitoring.

Method used

A method based on multi-directional Gabor convolutional networks is adopted. Through a dual-view fusion module and a multi-directional Gabor convolutional fusion module, combined with intra-view attention and inter-view attention, multi-directional Gabor convolutional kernels are constructed for feature fusion and detection. YOLOv12 Neck and Decoupled Head modules are used for object detection.

Benefits of technology

It significantly improves the accuracy of fall detection, reduces false alarms and false negatives, and enhances detection performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877389B_ABST
    Figure CN120877389B_ABST
Patent Text Reader

Abstract

The application discloses a fall detection method and system based on a multi-direction Gabor convolution network, relates to the technical field of image processing, receives a first-view image and a second-view image, inputs the first-view image and the second-view image into an encoder of a pre-established multi-direction Gabor convolution network model respectively, outputs first L-level features and second L-level features, and fuses the first L-level features and the second L-level features to obtain fused double-view features; the fused double-view features are subjected to multi-direction Gabor convolution fusion based on the pre-established multi-direction Gabor convolution network model to obtain multi-direction Gabor convolution fusion features, the multi-direction Gabor convolution fusion features are subjected to fall detection, and fall detection results of the first-view image and the second-view image are output; and thus the accuracy of target detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a fall detection method and system based on a multi-direction Gabor convolution network. BACKGROUND

[0002] With the intensification of global aging, the fall of people over 65 years old has become one of the main causes of death of this group. Hospitalized patients have a higher risk of falling than the general population due to factors such as physical function decline, drug side effects (such as sedatives, antihypertensive drugs), etc. In hospital, nursing home and other scenarios, the ratio of nursing staff to patients usually cannot meet the "all-weather one-to-one monitoring" requirement. Especially at night or during the peak of nursing, manual patrol has the problems of long time interval and high risk of omission, and it is difficult to discover the fall event in real time.

[0003] The camera combined with the deep learning algorithm (such as convolutional neural network) can recognize the falling action (such as sudden change of human posture, rapid falling), which is suitable for ward, corridor and other scenes, and the deep learning algorithm (such as YOLO target detection, Transformer time sequence analysis) can extract falling features from complex data, far exceeding traditional rule-based algorithms (such as threshold judgment). However, in scenes with diverse target postures, light changes, and obstacle occlusions, the algorithm is prone to false positives or false negatives, resulting in the fact that existing methods are difficult to effectively detect falls and have low accuracy. SUMMARY

[0004] To solve the problems mentioned in the background, the purpose of the present application is to provide a fall detection method and system based on a multi-direction Gabor convolution network.

[0005] In a first aspect, the purpose of the present application can be achieved by the following technical solution: a fall detection method based on a multi-direction Gabor convolution network, the method comprising the following steps:

[0006] Receiving a first view image and a second view image, inputting the first view image and the second view image into an encoder of a pre-established multi-direction Gabor convolution network model respectively, outputting first L-level features and second L-level features, and obtaining a fused dual-view feature by fusing the first L-level features and the second L-level features;

[0007] Performing multi-direction Gabor convolution fusion on the fused dual-view feature based on the pre-established multi-direction Gabor convolution network model to obtain a multi-direction Gabor convolution fusion feature, performing fall detection on the multi-direction Gabor convolution fusion feature, and outputting a fall detection result of the first view image and a fall detection result of the second view image.

[0008] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the processing of the first viewpoint image and the second viewpoint image includes:

[0009] Extract continuous sequences from the first-person perspective video. Frames are stacked together on channels to form a channel number of Image , No. Frame labels as images tags Extract consecutive frames from the second-person video. Frames are stacked together on channels to form a channel number of Image , No. Frame labels as images tags .

[0010] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the process of obtaining the first L-level feature and the second L-level feature:

[0011] Image The input is fed into the encoder, which is composed of a YOLOv12 backbone network. Output Level features ,in , to image The input is fed into the encoder, which is composed of a YOLOv12 backbone network. Output Level features ;

[0012] The encoder, which is composed of the yolov12 backbone network, is the encoder of the pre-established multi-directional Gabor convolutional network model.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the feeding of the first L-level features and the second L-level features into the dual-view fusion module includes the calculation of inter-view attention and intra-view attention.

[0014] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the calculation process of the inter-view attention, comprising:

[0015] Will Send in three Depthwise separable convolutional layers produce three components: the first Query components between level feature perspectives ,First Level feature perspective key components , the first level feature inter-view value component , i.e. , , wherein, denotes a depth separable convolution layer with the convolution kernel being , the input is sent into three depth separable convolutions to obtain three components: the second level feature inter-view query component , the second level feature inter-view key component , and the second level feature inter-view value component , i.e. , , ; the inter-view attention is calculated as:

[0016]

[0017]

[0018] wherein, denotes a matrix transposition;

[0019] The calculation process of the intra-view attention includes:

[0020] The input is sent into three deformable convolution layers to obtain three components: the first level feature intra-view query component , the first level feature intra-view key component , and the first level feature intra-view value component , i.e. , , wherein, denotes a deformable convolution layer. The input is sent into three deformable convolution layers to obtain three components: the second level feature intra-view query component , the second level feature intra-view key component , and the second level feature intra-view value component , i.e. , , , the intra-view attention is calculated as:

[0021]

[0022]

[0023] Intra-view attention and inter-view attention are fused:

[0024]

[0025]

[0026] wherein, represents a convolutional layer with a convolution kernel of , , wherein, represents a learnable parameter, and a fused dual-view feature is obtained: wherein, represents a convolutional layer with a convolution kernel of , represents that the features are stacked together in the channel.

[0027] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the pre-established multi-directional Gabor convolution fusion module comprises:

[0028] According to the fused dual-view feature , a query component of the multi-directional Gabor convolution feature is calculated , and Gabor kernels with a size of need to be constructed:

[0029]

[0030] wherein, , represent the horizontal and vertical coordinates of the Gabor kernel; represents a rotation angle, ; , , represents the standard deviation of the Gaussian factor; represents the frequency of the cosine function; represents the phase of the cosine function, and learnable convolution kernels with a size of are constructed , the product of the corresponding positions between the convolution kernel and the Gabor kernel is calculated, and a multi-directional Gabor convolution kernel is obtained: Then, multi-directional Gabor convolution is performed:

[0031] wherein, represents the first of channel, representing the first channel, representing a convolution operation;

[0032] According to the fusion of the two-view features , the healthy components of the multi-direction Gabor convolution features are calculated , and Gabor kernels with a size of need to be constructed:

[0033]

[0034] wherein, a standard deviation of a Gaussian factor is represented; a frequency of a cosine function is represented; a phase of the cosine function is represented, and learnable convolution kernels with a size of are constructed , the product of the corresponding positions between the convolution kernel and the Gabor kernel is calculated, and the multi-direction Gabor convolution kernel is obtained: , and then the multi-direction Gabor convolution is performed:

[0035] wherein, representing the first channel;

[0036] According to the fusion of the two-view features , the value components of the multi-direction Gabor convolution features are calculated , and Gabor kernels with a size of need to be constructed:

[0037]

[0038] wherein, a standard deviation of a Gaussian factor is represented; a frequency of a cosine function is represented; a phase of the cosine function is represented, and learnable convolution kernels with a size of are constructed , the product of the corresponding positions between the convolution kernel and the Gabor kernel is calculated, and the multi-direction Gabor convolution kernel is obtained: , and then the multi-direction Gabor convolution is performed:

[0039] wherein, denotes the first channel.

[0040] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the process of fusing the dual-view features based on the pre-established multi-directional Gabor convolution fusion module to perform multi-directional Gabor convolution fusion to obtain multi-directional Gabor convolution fusion features, including:

[0041] fusing , , each of the image blocks , image blocks , image blocks , wherein denotes the height of , , denotes the width of , , denotes the number of channels of , , denotes the width of the image block, and each image block , , is flattened, transposed, and then projected after using a linear transformation:

[0042] , ,

[0043] wherein , , denotes a linear transformation weight matrix, and , , is divided into components , , ​​​ wherein, , represents , , the number of channels; calculate the correlation between each component:

[0044]

[0045] wherein, represents a learnable position encoding, and components are combined and projected to obtain multi-direction Gabor convolution fusion features:

[0046]

[0047] wherein, represents a linear transformation weight matrix, and is transposed and shape-transformed to return to the original image block size , and then the multi-direction Gabor convolution fusion feature image blocks are spliced according to the original positions to form multi-direction Gabor convolution fusion features with the original feature map size .

[0048] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: the target detection of the multi-direction Gabor convolution fusion features is composed of two YOLOv12 Neck modules and two YOLOv12 Decoupled Head modules, and the multi-direction Gabor convolution fusion features are sequentially sent into a YOLOv12 Neck module and a YOLOv12 Decoupled Head module to output the predicted class, confidence, center coordinates and width and height of the target bounding box in the image ; the multi-direction Gabor convolution fusion features are sequentially sent into another YOLOv12 Neck module and another YOLOv12 Decoupled Head module to output the predicted class, confidence, center coordinates and width and height of the target bounding box in the image .

[0049] The second aspect, in order to achieve the above object, the application discloses a fall detection system based on multi-direction Gabor convolution network, comprising:

[0050] ​The feature processing module is configured to receive the first-view image and the second-view image, input the first-view image and the second-view image into an encoder of a pre-established multi-direction Gabor convolution network model respectively, output first L-level features and second L-level features, and obtain fused double-view features by fusing the first L-level features and the second L-level features.

[0051] The fall detection module is configured to perform multi-direction Gabor convolution fusion on the fused double-view features based on the pre-established multi-direction Gabor convolution network model to obtain multi-direction Gabor convolution fusion features, and perform fall detection on the multi-direction Gabor convolution fusion features to output fall detection results of the first-view image and the second-view image.

[0052] In another aspect of the present application, in order to achieve the above-mentioned purpose, a terminal device is disclosed, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, the fall detection method based on the multi-direction Gabor convolution network is adopted.

[0053] The present application has the following beneficial effects:

[0054] The present application designs a double-view fusion module, models the correlation between two views by fusing the intra-view attention and the inter-view attention of two views, reduces the difference between the features of the two views, enhances the representation consistency of the fused features, and strengthens the feature representation, thereby reducing the false alarm or missed alarm probability caused by obstacle occlusion. BRIEF DESCRIPTION OF DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced as follows, and obviously, other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings;

[0056] Figure 1 is a schematic diagram of the method of the present application;

[0057] Figure 2Fig. 1 is a schematic diagram of a multi-direction Gabor convolution network of the present application;

[0058] Figure 3 Fig. 6 is a schematic diagram of identifying a fall patient condition in different frames of two views according to the present application;

[0059] Figure 4 Fig. 7 is a schematic diagram of identifying a fall patient condition in different frames of two views according to the present application using a bidirectional long short-term memory model;

[0060] Figure 5 Fig. 8 is a schematic diagram of identifying a fall patient condition in different frames of two views according to the present application using an instance position-related network;

[0061] Figure 6 Fig. 9 is a schematic diagram of identifying a fall patient condition in different frames of two views according to the present application using a transformer-based fall detection model;

[0062] Figure 7 Fig. 10 is a schematic diagram of identifying a fall patient condition in different frames of two views according to the present application using a YOLO fall detection model;

[0063] Figure 8 Fig. 11 is a schematic diagram of identifying a fall patient condition in different frames of two views according to the present application using an inflated spatio-temporal convolutional autoencoder;

[0064] Figure 9 Fig. 12 is a schematic diagram of a system structure according to the present application. DETAILED DESCRIPTION

[0065] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.

[0066] Embodiment one:

[0067] As shown in Fig. 1, the fall detection method based on a multi-direction Gabor convolution network includes the following steps: Figure 1

[0068] S101: receiving a first-view image and a second-view image, inputting the first-view image and the second-view image into an encoder of a pre-established multi-direction Gabor convolution network model, outputting first L-level features and second L-level features, and performing double-view fusion on the first L-level features and the second L-level features to obtain fused double-view features;

[0069] The processing process of the first-view image and the second-view image includes:

[0070] ​Extract continuous sequences from the first-person perspective video. Frames are stacked together on channels to form a channel number of Image , No. Frame labels as images tags Extract consecutive frames from the second-person video. Frames are stacked together on channels to form a channel number of Image , No. Frame labels as images tags .

[0071] The process of obtaining the first L-level feature and the second L-level feature:

[0072] Image The input is fed into the encoder, which is composed of a YOLOv12 backbone network. Output Level features ,in , to image The input is fed into the encoder, which is composed of a YOLOv12 backbone network. Output Level features ;

[0073] The encoder, which is composed of the yolov12 backbone network, is the encoder of the pre-established multi-directional Gabor convolutional network model.

[0074] The dual-view fusion of the first L-level features and the second L-level features includes the calculation of inter-view attention and intra-view attention.

[0075] The calculation process for attention between viewpoints includes:

[0076] Will Send in three Depthwise separable convolutional layers produce three components: the first Query components between level feature perspectives ,First Level feature perspective key components ,First Inter-value components of level feature perspective ,Right now , , ,in, Indicates that the convolution kernel is Depth-separable convolutional layers, Send in three The deep separable convolution obtains three components: second inter-view query component of second-level feature inter-view key component of second-level feature inter-view value component of second-level feature inter-view attention is calculated as follows:

[0077]

[0078]

[0079] wherein, denotes matrix transposition;

[0080] The calculation process of intra-view attention includes:

[0081] The three components: first intra-view query component of first-level feature intra-view key component of first-level feature intra-view value component of first-level feature are obtained by inputting into three deformable convolution layers, i.e. , , , wherein, denotes a deformable convolution layer. The three components: second intra-view query component of second-level feature intra-view key component of second-level feature intra-view value component of second-level feature are obtained by inputting into three deformable convolution layers, i.e. , , , , intra-view attention is calculated as follows:

[0082]

[0083]

[0084] The intra-view attention and the inter-view attention are fused as follows:

[0085]

[0086] ​​​​​​

[0087] wherein, represents a convolution layer with a convolution kernel of , , wherein, , represents a convolution layer with a convolution kernel of , wherein,

[0088] S102: Perform multi-direction Gabor convolution fusion on the fused dual-view features based on a pre-established multi-direction Gabor convolution network model to obtain multi-direction Gabor convolution fusion features, perform fall detection on the multi-direction Gabor convolution fusion features, and output fall detection results of the first-view image and fall detection results of the second-view image.

[0089] The pre-established multi-direction Gabor convolution network model comprises:

[0090] According to the fused dual-view features , a query component of the multi-direction Gabor convolution feature is calculated , and Gabor kernels with a size of need to be constructed:

[0091]

[0092] wherein, , represent the horizontal and vertical coordinates of the Gabor kernel; represents a rotation angle, ; , , represents a standard deviation of a Gaussian factor; represents a frequency of a cosine function; represents a phase of the cosine function, and learnable convolution kernels with a size of are constructed, the convolution kernel is multiplied with the Gabor kernel at corresponding positions to obtain a multi-direction Gabor convolution kernel: , and then multi-direction Gabor convolution is performed:

[0093] wherein, represents the channel of , represents The aisle, This represents the convolution operation;

[0094] Based on the fusion of dual perspective features Calculate the key components of multi-directional Gabor convolution features. , needs to be constructed The size is Gabor core:

[0095]

[0096] in, The standard deviation represents the Gaussian factor; This represents the frequency of the cosine function; To represent the phase of the cosine function, construct... The size is Learnable convolutional kernels Calculate the convolution kernel Multiplying the corresponding positions of the Gabor kernel yields a multi-directional Gabor convolution kernel: Then perform multi-directional Gabor convolution:

[0097] in, express The aisle;

[0098] Based on the fusion of dual perspective features Calculate the value components of multi-directional Gabor convolution features. , needs to be constructed The size is Gabor core:

[0099]

[0100] in, The standard deviation represents the Gaussian factor; This represents the frequency of the cosine function; To represent the phase of the cosine function, construct... The size is Learnable convolutional kernels Calculate the convolution kernel Multiplying the corresponding positions of the Gabor kernel yields a multi-directional Gabor convolution kernel: Then perform multi-directional Gabor convolution:

[0101] in, express The Channel.

[0102] The fusion bi-view features are fused based on a pre-established multi-direction Gabor convolution network model to obtain multi-direction Gabor convolution fusion features, and the process includes:

[0103] The , , Each is divided into image blocks , image blocks , image blocks , wherein represents , , height of represents , , width of represents , , channel number of represents image block width, each image block , , is flattened, transposed and projected after linear transformation:

[0104] , ,

[0105] wherein , , represents a linear transformation weight matrix, and , , is divided into components , , , wherein , represents , , the number of channels; calculate the correlation between each component:

[0106]

[0107] wherein, represents a learnable position encoding, and merges and projects the components to obtain multi-directional Gabor convolution fusion features:

[0108]

[0109] wherein, represents a linear transformation weight matrix, and is transposed and shape-transformed back to the original image block size , and then the multi-directional Gabor convolution fusion feature image blocks are spliced according to the original positions into multi-directional Gabor convolution fusion features of the original feature map size .

[0110] The multi-directional Gabor convolution fusion features are detected by two YOLOv12 Neck modules and two YOLOv12 Decoupled Head modules, and the multi-directional Gabor convolution fusion features are sequentially sent into a YOLOv12 Neck module and a YOLOv12 Decoupled Head module to output the predicted class, confidence, center coordinates and width and height of the target bounding box in the image ; the multi-directional Gabor convolution fusion features are sequentially sent into another YOLOv12 Neck module and another YOLOv12 Decoupled Head module to output the predicted class, confidence, center coordinates and width and height of the target bounding box in the image .

[0111] The entire network is optimized using the following dual-view loss:

[0112] wherein, is the loss between the predicted class, confidence, center coordinates and width and height of the target bounding box in the image predicted by the multi-directional Gabor convolution network and the real class, confidence, center coordinates and width and height of the bounding box in the label:

[0113] ​

[0114] wherein, represents the number of grid cells of image division, represents the number of target bounding boxes in each grid cell, represents the th target bounding box in the th cell, represents the target in the th cell, the th background bounding box in the th cell, represents the horizontal center coordinate of the bounding box predicted by the network, represents the vertical center coordinate of the bounding box predicted by the network, represents the width of the bounding box predicted by the network, represents the height of the bounding box predicted by the network, represents the confidence of the prediction by the network, represents the probability of the th class predicted by the network, represents the horizontal center coordinate of the bounding box of the detection label, represents the vertical center coordinate of the bounding box of the detection label, represents the width of the bounding box of the detection label, represents the height of the bounding box of the detection label, represents the confidence of the detection label, represents the probability of the th class of the detection label, , represents the weight coefficient, represents the total number of classes of the detection label.

[0115] wherein, is the loss between the class, confidence, center coordinate and width-height of the bounding box predicted by the multi-direction Gabor convolution network of the detection target in the image and the class, confidence, center coordinate of the bounding box and width-height of the bounding box in the real label:

[0116]

[0117] wherein, represents the horizontal center coordinate of the bounding box predicted by the network, represents the vertical center coordinate of the bounding box predicted by the network, represents the width of the bounding box predicted by the network, represents the height of the bounding box predicted by the network, represents the confidence of the prediction by the network, a center horizontal coordinate of a bounding box of the detection label, a probability of the first class, a center horizontal coordinate of a bounding box of the detection label, a center vertical coordinate of a bounding box of the detection label, a width of a bounding box of the detection label, a height of a bounding box of the detection label, a confidence of the detection label, a probability of the first class, a probability of the first class.

[0118] Specifically, the scheme of the present application is further described below through examples:

[0119] In this embodiment, 2 different camera video data collected by the unit are used, which are composed of 15 patient 15 room 30 segment 320x240 size color video and corresponding detection labels. Randomly divided into 5 folds for cross-validation, patients are different between each fold.

[0120] In this embodiment, the pytorch deep learning framework is used, and the torch version is 1.13.0. All training and verification processes are completed on an NVIDIA GeForce RTX 3090 graphics card with a 24G size of display memory. In the training process, the neural network adopts a small batch method to read data, and the batch size is set to 2. The optimizer selects the stochastic gradient descent (Stochastic Gradient Descent, SGD) method, the initial learning rate is set to 0.01, the optimizer momentum is set to 0.9, and the optimizer regularization coefficient is set to 0.0001. The learning rate adjustment strategy selects the Poly method.

[0121] In order to quantitatively evaluate the performance of the method proposed in the present application, two different evaluation indexes are selected to evaluate the performance of the network in the fall detection task: recall rate, precision rate and F1 score.

[0122] The experimental results are as follows:

[0123] The detection results of the present application are compared with the existing methods of bidirectional long short-term memory model, instance position related network, fall detection model based on transformer, YOLO fall detection and inflation spatio-temporal convolution autoencoder method, as shown in Table 1, and Table 1 is a quantitative comparison of the detection results of the present application and the existing methods.

[0124] Table 1 Comparison of experimental results of the method of the present application and the existing method

[0125]

[0126] Referring to Table 1, compared with other existing methods, the method proposed in the application is superior to other existing methods in recall rate, precision rate and F1 score. Compared with the optimal bidirectional long short-term memory model in other methods, the recall rate of the method proposed in the application is increased by 2.21, and the precision rate is increased by 5.16; the F1 score of the method proposed in the application is increased by 2.99, and the F1 score reaches 94.75%.

[0127] In order to verify the effectiveness of each part of the application, an ablation experiment is performed, the baseline model is YOLOv12; model 1 is 2 feature extractors and , and The output features are spliced in the channel, and then 2 YOLOv12 Neck modules, 2 YOLOv12 Decoupled Head modules, and loss functions are added; model 2 is 2 feature extractors and , a dual-view fusion module, 2 YOLOv12 Neck modules, 2 YOLOv12 Decoupled Head modules, and loss functions. The ablation experiment results are shown in Table 2:

[0128] Table 2 Ablation experiment results table

[0129]

[0130] From the ablation experiment in Table 2 above, it can be seen that gradually adding the view, the dual-view fusion module, the multi-directional Gabor convolution fusion module, the recall rate, the precision rate and the F1 score are gradually improved, thereby indicating that each part of the method proposed in the application helps to improve the detection accuracy.

[0131] The detection results of the fall detection method of the application and the existing method are compared as shown in Figures 3-8 , the first row in each figure is the video shot by the first camera, the second row is the video shot by the second camera, and the three columns represent the identification of three different frames, the red box represents the fall patient identified by the algorithm, and the green box represents the fall patient labeled.

[0132] As shown in Figure 2 , the image is input into an encoder composed of a yolov12 backbone network outputs level features , wherein , the image is input into an encoder composed of a yolov12 backbone network output level feature ; the first L level feature and the second L level feature are input into a dual-view fusion module to obtain a fused dual-view feature; the fused dual-view feature is input into a multi-direction Gabor convolution fusion module for multi-direction Gabor convolution fusion to obtain a multi-direction Gabor convolution fusion feature; the multi-direction Gabor convolution fusion feature is subjected to fall detection, and is sequentially input into a YOLOv12 Neck module and a YOLOv12 Decoupled Head module to output a fall detection result of the first-view image, and is sequentially input into another YOLOv12 Neck module and another YOLOv12 Decoupled Head module to output a fall detection result of the second-view image.

[0133] It can be seen that, Figure 3 The method of the application shown in the figure is very close to the label. Other methods have a lot of false positives and false negatives of falls. Figure 4 The bidirectional long short-term memory model has fall false positives in both columns of the two views. Figure 5 The instance position related network has fall false positives in the first view of the first column and the two views of the third column. Figure 6 The fall detection model based on transformer has fall false positives in the second view of the first column and the two views of the third column, and the method has a fall false negative in the first view of the second column. Figure 7 YOLO fall detection has fall false positives in both views of the first column and both views of the third column, and the method has a fall false negative in the first view of the second column. Figure 8 The inflated spatio-temporal convolutional autoencoder has fall false positives in the first view of the first column and the first view of the third column, and the method has fall false negatives in both views of the second column.

[0134] Embodiment Two: In order to achieve the above purpose, as Figure 9 shown, based on the basis of embodiment one, the application discloses a fall detection system based on a multi-direction Gabor convolution network, comprising:

[0135] The feature processing module 11 is used for receiving the first-view image and the second-view image, inputting the first-view image and the second-view image into the encoder of the pre-established multi-direction Gabor convolution network model respectively, and outputting the first L level feature and the second L level feature, and performing dual-view fusion on the first L level feature and the second L level feature to obtain a fused dual-view feature.

[0136] The fall detection module 12 is configured to perform multi-directional Gabor convolution fusion on the fused dual-view features based on a pre-established multi-directional Gabor convolution network model to obtain multi-directional Gabor convolution fusion features, and perform fall detection on the multi-directional Gabor convolution fusion features to output fall detection results of the first-view image and the second-view image.

[0137] Based on the same inventive concept, the present application further provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the program comprises program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are configured to implement one or more instructions, and are specifically configured to load and execute one or more instructions in the computer storage medium to implement the above method.

[0138] It needs to be further explained that, based on the same inventive concept, the present application further provides a computer storage medium, which stores a computer program, and the computer program is executed by the processor to perform the above method. The storage medium can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of the computer readable storage medium include: electrical connections with one or more conductive wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, device or component.

[0139] In the description of the specification, the description of the terms "one embodiment", "an example", "a specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.

[0140] The basic principles, main features and advantages of the present disclosure are shown and described above. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and the above embodiments and descriptions in the specification are only to illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, various changes and improvements of the present disclosure can be made, which all fall within the scope of the claimed present disclosure.

Claims

1. A fall detection method based on a multi-directional Gabor convolution network, characterized by, The method comprises the following steps: The method comprises the following steps: The pre-established multi-direction Gabor convolution network model comprises: According to the fusion of dual-view features , the query component of multi-direction Gabor convolution features is calculated , and it is necessary to construct Gabor kernels with a size of ​ wherein, , denote the horizontal and vertical coordinates of the Gabor kernel; denote the rotation angle, ; , , denote the standard deviation of the Gaussian factor; denote the frequency of the cosine function; denote the phase of the cosine function, construct learnable convolution kernels , calculate the convolution kernel , multiply the corresponding positions between the convolution kernel and the Gabor kernel, get the multi-direction Gabor convolution kernel: , then perform multi-direction Gabor convolution: wherein denotes the first passage, denotes the first passage, denotes a convolution operation; According to the fusion of dual-view features , the robust components of multi-direction Gabor convolution features are calculated , it is necessary to construct Gabor kernels with a size of ​ wherein, denotes the standard deviation of the Gaussian factor; denotes the frequency of the cosine function; denotes the phase of the cosine function, which is constructed as a size of learnable convolutional kernels , the convolutional kernels are multiplied with the corresponding positions between the Gabor kernels, resulting in multi-direction Gabor convolutional kernels: Then, the multi-direction Gabor convolution is performed: wherein represents the first channel; According to the fusion of dual-view features , the value component of the multi-direction Gabor convolution feature is calculated , it is necessary to construct Gabor kernels with a size of ​ wherein, denotes the standard deviation of the Gaussian factor; denotes the frequency of the cosine function; denotes the phase of the cosine function, which is constructed of size learnable convolutional kernels , the convolutional kernels are multiplied with the corresponding positions between the Gabor kernels, resulting in multi-directional Gabor convolutional kernels: and then multi-directional Gabor convolutions are performed: wherein represents the first channel; The pre-established multi-direction Gabor convolution network model comprises: 2.The multi-directional Gabor convolution network based fall detection method according to claim 1, wherein, The processing process of the first view image and the second view image comprises: Extract consecutive frames from the first-person perspective video. Frames are stacked together on channels to form a channel number of Image , No. Frame labels as images tags Extract consecutive frames from the second-person video. Frames are stacked together on channels to form a channel number of Image , No. Frame labels as images tags . 3.The multi-directional Gabor convolution network based fall detection method according to claim 1, wherein, The acquisition process of the first L-level feature and the second L-level feature comprises: feeding the image into an encoder composed of a yolov12 backbone network output level features wherein feeding the image into an encoder composed of a yolov12 backbone network output level features ; The encoder composed of the yolov12 backbone network is the encoder of the pre-established multi-direction Gabor convolution network model. 4.The multi-directional Gabor convolution network based fall detection method according to claim 1, wherein, The double-view fusion of the first L-level feature and the second L-level feature comprises the calculation of inter-view attention and intra-view attention. 5.The multi-directional Gabor convolution network based fall detection method according to claim 4, wherein, The calculation process of the inter-view attention comprises: Will Send in three Depthwise separable convolutional layers produce three components: the first Query components between level feature perspectives ,First Inter-level feature perspective key components ,First Inter-value components of level feature perspective ,Right now , , ,in, Indicates that the convolution kernel is Depth-separable convolutional layers, Send in three Depthwise separable convolution yields three components: the second... Query components between level feature perspectives ,second Inter-level feature perspective key components ,second Inter-value components of level feature perspective ,Right now , , Computational attention between viewpoints: wherein denotes matrix transposition; The calculation process of the intra-view attention comprises: Will Three deformable convolutional layers are fed into the first layer to obtain three components: Query components within the perspective of level features ,First Intra-key components from the perspective of level features ,First Intrinsic components of level feature perspective ,Right now , , ,in, Represents a deformable convolutional layer; Three deformable convolutional layers are fed into the second layer to obtain three components: Query components within the perspective of level features ,second Intra-key components from the perspective of level features ,second Intrinsic components of level feature perspective ,Right now , , Calculate in-view attention: The intra-view attention and the inter-view attention are fused: wherein, represents a convolutional layer with a convolution kernel of , , represents a learnable parameter, and a fused dual-view feature is obtained: wherein, represents a convolutional layer with a convolution kernel of , represents that the features are stacked together in the channel. 6.The multi-directional Gabor convolution network based fall detection method according to claim 1, wherein, The process of performing multi-direction Gabor convolution fusion on the fused double-view feature based on the pre-established multi-direction Gabor convolution network model to obtain multi-direction Gabor convolution fusion features comprises: will be , , each divided into image blocks , image blocks , image blocks wherein denotes , , the height of denotes , , the width of denotes , , the number of channels of denotes the image block width, each image block , , is flattened, transposed and projected using a linear transformation: , , wherein, , , denotes a linear transformation weight matrix, which will be , , split into components , , wherein, , denotes , , the number of channels; the correlation between each component is calculated: wherein, denotes a learnable position encoding, and component merging projection, obtaining a multi-direction Gabor convolution fusion feature: wherein, represents a linear transformation weight matrix, and Transposing, shape transforming and returning to the original image block size , and then The multi-direction Gabor convolution fusion feature image block is spliced into the original feature size according to the original position Multi-direction Gabor convolution fusion features . 7.The multi-directional Gabor convolution network based fall detection method according to claim 1, wherein, The target detection of the multi-direction Gabor convolution fusion feature is composed of two YOLOv12 Neck modules, two YOLOv12 Decoupled Head modules, and the multi-direction Gabor convolution fusion feature is sequentially sent into a YOLOv12 Neck module and a YOLOv12 Decoupled Head module to output an image , and the prediction class, confidence, center coordinates and width and height of the target bounding box in the image are detected is sequentially sent into another YOLOv12 Neck module and another YOLOv12 Decoupled Head module to output an image , and the prediction class, confidence, center coordinates and width and height of the target bounding box in the image are detected 8. A fall detection system based on multi-directional Gabor convolution network, adopting the fall detection method based on multi-directional Gabor convolution network according to any one of claims 1 to 7, characterized in that, The pre-established multi-direction Gabor convolution network model comprises: The feature processing module is configured to receive the first view image and the second view image, input the first view image and the second view image into the encoder of the pre-established multi-direction Gabor convolution network model respectively, output the first L-level feature and the second L-level feature, and perform double-view fusion on the first L-level feature and the second L-level feature to obtain the fused double-view feature. The fall detection module is configured to perform multi-direction Gabor convolution fusion on the fused double-view feature based on the pre-established multi-direction Gabor convolution network model to obtain multi-direction Gabor convolution fusion features, perform fall detection on the multi-direction Gabor convolution fusion features, and output the fall detection result of the first view image and the fall detection result of the second view image.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program capable of running on the processor, and the processor loads and executes the computer program. When the processor loads and executes the computer program, the fall detection method based on the multi-direction Gabor convolution network in any one of claims 1 to 7 is adopted.

Citation Information

Patent Citations

  • View angle adaptive multi-target fall detection method based on graph convolutional neural network

    CN112966628A

  • Fall detection method based on SpA-YOLOv7

    CN118982863A