Lightweight pedestrian falling detection method and system for inspection robot
By applying a pedestrian fall detection model with lightweight convolution and dual attention mechanisms in patrol robots, the traditional method has solved the shortcomings in accuracy and real-time performance, and efficient and accurate fall detection is achieved, which is suitable for low-computing equipment.
Patent Information
- Application Number
- CN202510289917.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional fall detection methods have shortcomings in accuracy, real-timeness and applicability, which are difficult to meet the challenges of increasing aging population and public safety needs.
The lightweight convolution and dual attention mechanism are used to build an efficient pedestrian fall detection model. Through the lightweight YOLOv11 network and ST-GCN model, it realizes rapid detection of pedestrian key nodes and real-time identification of fall events.
It improves the accuracy and robustness of pedestrian fall detection, reduces calculation costs, is suitable for low-computing equipment, and ensures the real-time and efficient system.
Smart Images

Figure CN120148115A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and real-time human pose estimation, and relates to a lightweight pedestrian fall detection method and system for inspection robots. Background Art
[0002] Due to the decline of physical functions, the balance ability and reaction speed of the elderly are weakened, and the risk of falling is relatively high. Falls often may lead to serious consequences, such as fractures, brain injuries, etc., and even endanger life. In addition, in scenarios such as construction sites, factories, power, etc., as well as public places such as shopping malls, subway stations, etc., fall incidents also occur from time to time, posing a threat to personnel safety.
[0003] The emergence of fall detection technology provides strong support for the prevention and handling of fall incidents. First of all, it can alarm in time when a fall incident occurs, notify relevant personnel to take rescue measures, thereby shortening the rescue time and reducing the harm consequences. Secondly, fall detection technology can monitor the activities of personnel in real time to prevent fall incidents. For example, in a home environment, a fall detection algorithm can be combined with a smart home system to detect fall incidents in real time; in a medical institution, a fall detection algorithm can be applied to a bedside monitoring system to provide timely rescue information for medical staff. In addition, fall detection technology also helps to improve the safety of public places and reduce disputes and losses caused by fall incidents.
[0004] The deficiencies of traditional fall detection methods in terms of accuracy, real-time performance and applicability have become the key factors restricting their wide deployment in practical applications. With the aggravation of the global population aging and the continuous growth of public safety needs, higher requirements are put forward for the accuracy and real-time performance of fall detection. However, traditional fall detection methods, such as manual monitoring and simple sensor detection, often fail to meet these needs.
[0005] Although the manual monitoring method can visually observe fall incidents, there are obvious response delay problems. The monitoring personnel need to stare at the monitoring screen for a long time, and it is easy to miss fall incidents due to fatigue or inattention. At the same time, the cost of manual monitoring is relatively high, requiring a large amount of human resources, and the monitoring range is limited, making it difficult to cover all areas that need to be monitored.
[0006] Simple sensor detection methods attempt to capture physical changes during a fall through sensors, such as acceleration, pressure, etc. However, such methods often have a high false alarm rate. Due to the sensitivity of sensors to the environment, some normal daily activities, such as walking fast, sitting down, lying down, etc., may be misjudged as fall incidents. In addition, the environmental adaptability of simple sensors is also poor, and it is easily affected by external factors, such as temperature, humidity, electromagnetic interference, etc., thus affecting the detection accuracy. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a method and system for rapid detection and processing of pedestrian falls for inspection robots. It provides an efficient solution for pedestrian fall detection through lightweight convolution and dual attention mechanisms. Lightweight convolution reduces the number of parameters and computational complexity, significantly reducing the computational cost of the model and being suitable for low-computing-power environments such as embedded devices. The dual attention mechanism combines channel attention and spatial attention, enabling the model to focus on important features in different dimensions, thereby improving the accuracy and robustness of object detection. Combining these two technologies, an efficient, accurate pedestrian fall detection model suitable for low-computing-power devices can be constructed.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] On the one hand, the present invention provides a lightweight pedestrian fall detection method for inspection robots, and the method includes the following steps:
[0010] S1. Obtain pedestrian pose image data through the inspection robot and preprocess the data;
[0011] S2. Construct a lightweight YOLOv11 network, which at least includes several ADown lightweight downsampling modules and C2PSA_DAT composite modules;
[0012] S3. Train the lightweight YOLOv11 network to obtain the optimal weights, thereby establishing a rapid detection model for pedestrian key nodes, and then obtain pedestrian key node data in real time through the rapid detection model for pedestrian key nodes;
[0013] S4. Use the ST-GCN model, a dynamic-skeleton-based action recognition method, to process the obtained pedestrian key node data and perform real-time pedestrian fall detection.
[0014] Further, in step S1, the data preprocessing includes data annotation and data partitioning. Among them, the LabelImg method is used to annotate the original pedestrian pose image data, the annotation category is pedestrians, and the annotation content is several key nodes of pedestrians, forming a data set required for training the human key node detection model; the data partitioning includes dividing the annotated human pose image data and label files into a training set and a test set.
[0015] Further, in step S2, the structure of the constructed lightweight ADAT-YOLOv11 model at least includes a backbone network and a head network. Among them, the backbone network at least includes several Adown modules, Conv modules, C3K2 modules, as well as at least one SPPF module and one C2PSA-DAT module. In the backbone network, multi-scale feature extraction is performed through the combination of several groups of Adown modules, C3K2 modules and C2PSA-DAT modules. Among them, the Adown module performs lightweight convolution operations, and the C3K2 module abstracts the feature information; the highest-scale features extracted are input into the SPPF module and the C2PSA-DAT module to extract key information;
[0016] The head network at least includes several Upsample upsampling modules, Concat splicing modules, C3K2 modules, Adown modules, and Pose detection modules; first, upsampling is performed through the Upsample module, then the feature maps of different resolutions are spliced through the Concat module, and then the features are further processed using C3K2 and Adown to obtain human key node data.
[0017] Further, in the backbone network of the lightweight ADAT-YOLOv11 model, the input features are first preliminarily processed by the Conv module, and then the data is lightweight processed by the Adown module. The process is as follows:
[0018]
[0019] where X ∈ R H×W×C is the input feature map, is the learnable dimensionality reduction weight matrix, F(·) is the non-linear transformation function, is the feature representation output by the Adown module;
[0020] Then, the feature extraction ability is enhanced by the C3K2 module. Let the input feature map be X in , then its processing process is:
[0021] X 1 = σ(Conv 1×1 (X in ))
[0022] X 2 = σ(Conv 3×3 (X 1 )) X 3 = σ(DWConv 2×2 (X 1 ))
[0023] X Concat =(X 2 ,X3 )
[0024] X out = σ(Conv 1×1 (X concat ))
[0025] wherein, X 1 is the result of the first - layer convolution, X 2 , X 3 are two branches of the second - layer convolution, X Concat is the third - layer processing, which connects and fuses the two branches, X out is the output feature representation after adjusting the number of channels through convolution; Conv 1×1 (·) is a 1×1 convolution kernel, Conv 3×3 (·) represents a 3×3 convolution kernel, DWConv 2×2 (·) represents a 2×2 convolution kernel, σ(·) represents the fusion function;
[0026] After alternating processing of multiple Adown modules and C3K2 modules, the SPPF module performs key feature extraction and fusion of multi - scale features. Among them, assuming the feature representation input to the SPPF module is X i ′ n , then its processing process is expressed as:
[0027] X o ′ ut = Concat(MaxPool(X i ′ n , C 1 ), MaxPool(X i ′ n , C 2 ), MaxPool(X i ′ n , C 3 ))
[0028] wherein, MaxPool(·) is the max - pooling operation, C 1 , C 2 , C 3 are different scales of the input features; Concat(·) is the concatenation function, X′ out is the feature representation output by the SPPF module;
[0029] Finally, the C2PSA - DAT module performs cross - stage partial feature selection and spatial and channel attention adjustment, and its process is:
[0030]
[0031] wherein, represent the feature inputs of the m-th layer and the (m - 1)-th layer respectively, α represents the feature weight, represents the feature output of the (m + 1)-th layer, A s and A c represent the spatial and channel attention weights respectively, the symbol represents element-wise weighting; represents the final output representation after spatial and channel attention adjustment;
[0032] In the head network of the lightweight ADAT-YOLOv11 model, the feature processing methods of the C3K2 module and the Adown module are the same as those of the C3K2 module and the Adown module in the backbone network;
[0033] The Concat module enhances the expression ability of the model through cross-layer feature connection and fusion, and retains the low-level and high-level feature information. From the high feature map X high ∈R H×W×C and the low feature map X low ∈R H×W×C perform alignment and fusion, and adjust the size consistency by upsampling or downsampling; assume that the low-level feature map is adjusted to the same spatial dimension as the high-level feature map by upsampling, and then the two are concatenated and fused as the input for subsequent feature extraction:
[0034] X low_upsamp =Upsample(X low ,H',W')
[0035] X fusion =Concat(X low_upsamp ,X high )
[0036] In the formula, X low_upsamp represents the adjusted spatial dimension of the low-level feature map, H′,W' represent the height and width of the output image respectively, and X fusion is the fused feature representation after concatenation.
[0037] Furthermore, the Adown module receives input image data, time series signals or other sensor data; the Adown module divides the input data into multiple small blocks or segments, where the division methods include at least through sliding window and segmentation processing methods; when processing image data, the image will be cut into multiple small regions, and each region will be downsampled separately; when processing time series data, the data is cut according to time periods, and each segment of time series data is downsampled separately, where,
[0038] Suppose the size of the input features of the ADown module is H×W×C, where H and W represent the height and width of the feature map, and C is the number of channels. After the ADown module reduces the dimension of the input by learning the matrix weight W, the number of parameters for convolving the same feature map is as follows:
[0039] P ADown = C in ×C out ×R
[0040] In the formula, R represents the number of channels after dimension reduction.
[0041] Furthermore, the C2PSA-DAT module combines the cross-stage partial attention selection module C2PAS and the dynamic attention mechanism DAT for feature processing. Among them, C2PAS passes part of the feature map to the next layer through a cross-stage connection method and performs weighted selection on the feature map; DAT automatically adjusts the attention weight according to the different characteristics of the input data. Among them:
[0042] The C2PSA-DAT module contains the C2PSA module, which passes part of the features through a cross-stage connection method and performs weighted selection on the features. Suppose the input feature map of the m-th layer is X (m) ∈R H×W×C , and the input feature map of the m+1-th layer is composed of a part of the feature map X s (m) of the current layer and a part of the feature map X s (m-1) of the previous layer. The selection matrix M s is used to transfer the cross-stage partial features, and the cross-stage feature selection is expressed as:
[0043]
[0044] Among them, M s ∈{0, 1} H×W×C is a binary selection matrix that selects the channels to be retained in the next stage, and the symbol · represents element-wise multiplication.
[0045] Furthermore, in step S3, the process of training the constructed lightweight YOLOv11 network is as follows:
[0046] S31. Perform data augmentation on the input data, including random scaling, random flipping, mixup, and Mosaic augmentation;
[0047] S32. Input the augmented data into the lightweight YOLOv11 network, extract features using the lightweight backbone network, and optimize the features by combining the cross-stage partial attention selection module and the dynamic attention mechanism;
[0048] S33. Generate multi-scale predictions through the target detection head, including class probabilities, object confidences, and bounding box parameters;
[0049] S34. Use the gradient descent method and optimization algorithm to update network parameters and gradually improve the detection accuracy of the model;
[0050] Among them, the multi-task weighted loss is used as the loss function, which is expressed as:
[0051] L = λ cls L cls + λ obj L obj + λ box L box
[0052] The classification loss L cls Adopt the cross-entropy loss to measure the prediction accuracy of the target class:
[0053]
[0054] Among them, y i is the true class label, is the predicted class probability;
[0055] The object confidence loss L obj Adopt the binary cross-entropy loss to measure the model's judgment on the existence of the object:
[0056]
[0057] Among them is the predicted object confidence;
[0058] The bounding box regression loss L box Adopt the CIoU loss to optimize the matching degree between the predicted box and the true box:
[0059] L box = 1 - CIoU
[0060] Among them, the CIoU calculation method takes into account IoU, the distance between the center points, and the aspect ratio.
[0061] Furthermore, in step S4, it specifically includes the following steps:
[0062] S41. Data preprocessing, output the skeletal points of each frame of the video and assemble them into a complete action; among them, adopt the graph partitioning strategy to establish multiple adjacency matrices reflecting different motion states; adopt the spatial configuration partitioning method to partition the human key node data, that is, take the distance between the root node and the center of gravity as the benchmark, and among all the distances from the adjacent nodes to the center of gravity, those less than the benchmark value are regarded as in-node points, and those greater than the benchmark value are regarded as out-node points;
[0063] S42. Use the processed video frames as the input to the YOLOv11-DAT key point detection network. The network detects the targets in the frames, identifies pedestrians and generates corresponding bounding boxes. At the same time, use the dynamic attention mechanism to optimize the key point feature extraction process, accurately locate the human body bone key points, and form complete human pose information;
[0064] S43. Input the obtained key point data information into the ST-GCN model. First, perform normalization processing through the normalization layer, where the normalization is carried out in the time and space dimensions, and normalize the position features (x, y, acc) of a joint in different frames;
[0065] S44. Adopt several ST-GCN units including spatial graph convolution modules and temporal graph convolution modules for feature extraction and spatio-temporal information fusion to obtain fused spatio-temporal information features;
[0066] S45. Input the fused spatio-temporal features into the fully connected layer, classify the pedestrian behavior through the classification layer, and judge whether the pedestrian has a falling behavior according to the output result; if the model output is "fall", then mark fall on the image frame, frame out the falling person, and trigger the corresponding alarm or subsequent processing process; if it is "normal walking", then continue with subsequent behavior monitoring.
[0067] Further, in step S43, the ST-GCN unit includes alternately arranged spatial graph convolution GCN, temporal graph convolution TCN and residual structure residual. Among them, let the input graph be G=(V, E), where V is the set of nodes, representing each joint point of the human body; E is the set of edges, representing the connection relationship between joint points; each node v i has a feature x i , and there is a corresponding adjacency matrix A in the graph, where A ij represents the weight connection between node v i and v j ;
[0068] Then the core formula of the spatial graph convolution is expressed as:
[0069]
[0070] X′ is the output feature matrix after convolution, X is the input feature matrix, with the shape of N×C, N is the number of nodes, and C is the feature dimension of each node; is the normalized adjacency matrix, which is expressed as:
[0071]
[0072] Among them, A is the original adjacency matrix, and D is the degree matrix of nodes, that is W is a trainable weight matrix used to learn the weights of each node feature;
[0073] The temporal graph convolution adopts a one-dimensional convolution operation, that is, the features are processed with a sliding window in the time dimension. The formula is as follows:
[0074] X″ = f(W t *X′ + b)
[0075] Among them, X′ ∈ R N×C×T is the input time series feature, N is the number of nodes, C is the feature dimension, and T is the time step; W t is the temporal convolution kernel used to learn the dependencies in time; * represents the one-dimensional convolution operation, that is, a sliding window is used for weighted summation in the time dimension T; b is the trainable bias term, and f(·) is the activation function, which is used for non-linear transformation;
[0076] Finally, the obtained output is expressed as:
[0077]
[0078] Among them represents performing spatial graph convolution to extract spatial relationships; W t *(·) performs temporal graph convolution to extract temporal dependencies, and X out is the output feature after fusing spatial information and temporal information.
[0079] On the other hand, a detection system for implementing the aforementioned lightweight pedestrian fall detection method for an inspection robot is also provided. The system includes a robot body, a housing, front wheels, a support frame, a movement control component, and various sensors;
[0080] The robot body is protected by the housing. The front wheels are used for movement, and the support frame is used to support the body. The movement control component and the support frame are respectively fixedly connected to the housing;
[0081] The core of the platform uses a NUC host as the computing processor and runs the ROS system as the robot operating system.
[0082] The beneficial effects of the present invention are as follows:
[0083] In the target detection module of the present invention, the improved YOLOv11 adopts the C2PSA-DAT structure. This module combines the C2PSA lightweight network and the DAtention (dynamic attention) mechanism, which not only effectively reduces the computational complexity of the model, but also enhances the feature extraction ability, enabling the detection network to operate efficiently with lower computing power on embedded devices and ensuring the real-time performance of the system.
[0084] Secondly, in order to improve the accuracy of fall detection, the present invention integrates ST-GCN (Spatial-Temporal Graph Convolutional Network), uses Graph Convolution (GCN) to model the topological structure between human body bone points (joint points), and combines Temporal Convolution (TCN) to capture the movement trends of joint points in the time dimension, enabling the system to not only recognize the static postures after falling but also analyze the complete fall process, thereby effectively distinguishing similar actions such as normal walking, sitting down, and lying down, and reducing false alarms.
[0085] The present invention features high precision, low latency, strong environmental adaptability, and high automation, and can be widely applied to scenarios such as smart homes, nursing homes, hospitals, and public security monitoring, providing a more intelligent and accurate fall detection solution for society.
[0086] Through the present invention, pedestrian fall detection will no longer rely on traditional manual inspections or single-sensor detections, but instead, with the help of artificial intelligence and deep learning technologies, achieve all-weather and efficient intelligent monitoring, thereby reducing the risks brought by fall accidents, improving the level of social public security, and providing a scientific and reasonable decision-making basis for relevant management departments, promoting the development of intelligent security technologies.
[0087] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail with preference below in conjunction with the drawings, where:
[0089] Figure 1 is a schematic structural diagram of the lightweight YOLOv11 network constructed in the embodiment of the present invention;
[0090] Figure 2 is a schematic diagram of the partitioning results of different graph partitioning strategies in the embodiment of the present invention, where, Figure 2 (a) is a key frame of the input skeleton, Figure 2 (b) is a schematic diagram of the partitioning result using Uni-labeling; Figure 2 (c) is a schematic diagram of the partitioning result using Distance partitioning; Figure 2 (d) is a schematic diagram of the partitioning result of Spatialconfiguration partitioning adopted in this embodiment;
[0091] Figure 3 Schematic diagram of the network structure of the ST-GCN model in the embodiments of the present invention;
[0092] Figure 4 Schematic diagram of the module connection of the ST-GCN model in the embodiments of the present invention;
[0093] Figure 5 Example of pedestrian fall detection output in the embodiments of the present invention;
[0094] Figure 6 Schematic diagram of the comparison of the accuracy index mAP and the computational complexity index GFLOPs of pedestrian fall detection between the lightweight model ADAT-YOLOv11 and the existing YOLOv11 model in the same dataset in the embodiments of the present invention. Among them, Figure 6 (a) Training results of the ADown-YOLOv11 model after only improving the ADown downsampling; Figure 6 (b) Training results of the original unimproved YOLOv11 model; Figure 6 (c) Training results of the ADAT-YOLOv11 model of the present invention. Specific implementation manners
[0095] The following illustrates the implementation manners of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0096] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0097] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0098] Please refer to Figures 1 to 6 , a lightweight pedestrian fall detection method and system for inspection robots.
[0099] Embodiment 1
[0100] This embodiment provides a detailed process of a lightweight pedestrian fall detection method for inspection robots, which specifically includes the following steps:
[0101] S1. Collect pedestrian pose image data through the inspection robot and preprocess the data;
[0102] S2. Construct a lightweight YOLOv11 network, which at least includes several ADown lightweight downsampling modules and C2PSA_DAT composite modules;
[0103] S3. Train the lightweight YOLOv11 network to obtain the optimal weights, thereby establishing a pedestrian key node rapid detection model, and then obtain pedestrian key node data in real time through the pedestrian key node rapid detection model;
[0104] S4. Use the ST-GCN model of the action recognition method based on dynamic skeleton to process the obtained pedestrian key node data and perform real-time pedestrian fall detection.
[0105] In step S1 of this embodiment, dynamic human pose sequence data is collected through a multi-modal visual sensor, and an adaptive light compensation and skeleton centering strategy is used for data preprocessing; the preprocessing process of the original pedestrian pose image data at least includes data annotation. Among them, the LabelImg method is used to annotate the original pedestrian pose image data, and the annotation category is pedestrians, with a total of 17 key nodes, forming a data set required for training the human key node detection model. After the annotation is completed, the annotated human pose image data and label files are divided into a training set and a test set.
[0106] In step S2 of this embodiment, as Figure 1As shown, the structure of the constructed lightweight ADAT-YOLOv11 model includes at least a backbone network and a head network. Among them, the backbone network includes at least several ADown, Conv, and C3K2 modules, as well as at least one SPPF module and one C2PSA-DAT module. In the backbone network, multi-scale feature extraction is performed through the combination of several groups of ADown modules, C3K2 modules, and C2PSA-DAT modules. Among them, the ADown module performs lightweight convolution operations, and the C3K2 module abstracts the feature information; the highest-scale features extracted are input into the SPPF and C2PSA-DAT modules to extract key information;
[0107] The head network includes at least several Upsample upsampling modules, Concat splicing modules, C3K2 modules, ADown modules, and Pose detection modules; first, the feature maps of different resolutions are upsampled by Upsample and spliced by Concat; then the features are further processed using C3K2 and ADown;
[0108] Finally, multiple Pose detection head layers detect targets of different scales.
[0109] The data processing process of ADown is as follows:
[0110] The ADown module receives input image data, time series signals, or other sensor data. Before the input data enters the downsampling module, some basic preprocessing operations need to be performed, including denoising, normalization, or standardization, etc. These operations can help improve the effect of the subsequent downsampling process and avoid unnecessary interference in the data.
[0111] The preprocessing steps include using filters, smoothing algorithms, etc. to remove noise and improve data quality; normalizing the range of the data to a certain interval, such as [0,1] or [-1,1], to ensure data consistency during subsequent processing; performing data conversion to ensure that the format of the input data meets the requirements of the downsampling algorithm.
[0112] Furthermore, the ADown module receives input image data, time series signals, or other sensor data; the ADown module divides the input data into multiple small blocks or segments, where the division methods include at least through sliding windows and segment processing methods; when processing image data, the image is cut into multiple small regions, and each region is downsampled separately; when processing time series data, the data is cut according to time periods, and each segment of time series data is downsampled separately. The structure of the C2PSA-DAT module is:
[0113] C2PAS passes the partial feature maps to the next layer through a cross-stage connection method, and performs weighted selection on these feature maps, so as to retain the most useful information and reduce unnecessary redundant calculations.
[0114] The DAttention mechanism is a dynamic attention mechanism designed to automatically adjust the attention weights according to the different characteristics of the input data.
[0115] The C2PSA-DAT module is an enhanced version based on the cross-stage partial attention selection module C2PAS (Cross-Stage Partial Attention Selection) in YOLOv11 and combined with the dynamic attention mechanism DAT (Dynamic Attention). By enhancing the flexibility of feature selection and attention, the C2PSA-DAT module can provide better feature representation when dealing with complex visual tasks, and improve the performance of the model in tasks such as object detection and image classification.
[0116] Thus, in the backbone network of the lightweight ADAT-YOLOv11 model in this embodiment, the input features are first preliminarily processed by the Conv module, and then the data is lightened by the Adown module. The process is as follows:
[0117]
[0118] where X ∈ R H×W×C is the input feature map, is the learnable dimensionality reduction weight matrix, F(·) is the non-linear transformation function, is the feature representation output by the Adown module; since W is trained and optimized to be able to compress information more efficiently, compared with fixed pooling or ordinary convolution, ADown can reduce the amount of computation while maintaining or even improving the accuracy.
[0119] The size of an input feature map is H×W×C, where C is the number of channels, and the number of parameters of its standard convolution is
[0120] P conv = C in ×C out ×K 2
[0121] where K is the size of the convolution kernel, and for the Adown down-convolution module, after learning the matrix weight W and reducing the dimension of the input, the number of parameters of its convolution on the same feature map is:
[0122] P ADown = C in ×C out ×R
[0123] Obviously, R << K 2 , that is, the weight learning method of ADown makes its parameter quantity much smaller than that of the standard convolution.
[0124] The key for ADown to reduce the parameter quantity while maintaining the accuracy lies in the transformation matrix W it uses, which can more effectively retain feature information, enabling the subsequent network layers to learn better feature representations.
[0125] Then, the C3K2 module enhances the feature extraction ability. The C3K2 module is an improvement of the C3 module in the YOLO series of networks. It combines lightweight design to enhance the feature extraction ability while reducing the computational complexity. Let the input feature map be X in , then its processing process is as follows:
[0126] X 1 = σ(Conv 1×1 (X in ))
[0127] X 2 = σ(Conv 3×3 (X 1 )) X 3 = σ(DWConv 2×2 (X 1 ))
[0128] X Concat = (X 2 , X 3 )
[0129] X out = σ(Conv 1×1 (X concat ))
[0130] In the formula, X 1 is the result of the first-layer convolution, X 2 , X 3 are the two branches of the second-layer convolution, X Concat is the third-layer processing, which connects and fuses the two branches, X out is the output feature representation after adjusting the number of channels through convolution; Conv 1×1 (·) is a 1×1 convolution kernel, Conv 3×3 (·) represents a 3×3 convolution kernel, DWConv 2×2 (·) represents a 2×2 convolution kernel, and σ(·) represents the fusion function.
[0131] After alternating processing of multiple Adown modules and C3K2 modules, the SPPF module performs key feature extraction and fusion of multiple scale features. Among them, the feature representation input to the SPPF module is set as X i ′ n , then its processing process is expressed as:
[0132] X′ out =Concat(MaxPool(X′ in ,C 1 ),MaxPool(X′ in ,C 2 ),MaxPool(X′ in ,C 3 ))
[0133] In the formula, MaxPool(·) is the max pooling operation, C 1 ,C 2 ,C 3 are different scales of the input features; Concat(·) is the concatenation function, and X′ out is the feature representation output by the SPPF module;
[0134] Finally, the C2PSA-DAT module performs cross-stage partial feature selection and spatial and channel attention adjustment. The process is as follows:
[0135]
[0136] In the formula, respectively represent the feature inputs of the m-th layer and the (m - 1)-th layer, α represents the feature weight, represents the feature output of the (m + 1)-th layer, A s and A c respectively represent A s and A c respectively represent the spatial and channel attention weights, and the symbol represents element-wise weighting; represents the final output representation after spatial and channel attention adjustment. In the C2PSA of the C2PSA-DAT module, it mainly passes partial features through a cross-stage connection method and performs weighted selection on the features to improve information utilization and reduce redundant calculations. Assume that the input feature map of the m-th layer is X (m) ∈R H×W×C , and the input feature map of the (m + 1)-th layer is composed of a part of the feature map X s (m) of the current layer and a part of the feature map X s (m-1) of the previous layer. The selection matrix M s is used to transfer cross-stage partial features, then the cross-stage feature selection can be expressed as:
[0137]
[0138] Among them, M s ∈ {0, 1} H×W×C is a binary selection matrix that controls which channels are retained for the next stage, and the symbol · represents element-wise multiplication.
[0139] In the head network of the lightweight ADAT-YOLOv11 model of this embodiment, the feature processing methods of the C3K2 module and the Adown module are the same as those of the C3K2 module and the Adown module in the backbone network;
[0140] The function of the Concat module is to enhance the expression ability of the model through cross-layer feature connection and fusion, retain low-level and high-level feature information, so as to better capture complex patterns.
[0141] The Concat module enhances the expression ability of the model through cross-layer feature connection and fusion, retains low-level and high-level feature information, from the high feature map X high ∈ R H×W×C and the low feature map X low ∈ R H×W×C are aligned and fused, and the size consistency is adjusted by upsampling or downsampling; assuming that the low-level feature map is adjusted to the same spatial dimension as the high-level feature map by upsampling, and then the two are concatenated and fused as the input for subsequent feature extraction:
[0142] X low_upsamp = Upsample(X low , H', W')
[0143] X fusion = Concat(X low_upsamp , X high )
[0144] In the formula, X low_upsamp represents the adjusted spatial dimension of the low-level feature map, H′, W' respectively represent the height and width of the output image, and X fusion is the fused feature representation after concatenation.
[0145] In step S3 of this embodiment, the process of training the constructed lightweight YOLOv11 network is as follows:
[0146] Data augmentation is performed on the input data, including random scaling, random flipping, mixup, and Mosaic augmentation, to improve the generalization ability of the model. Then, the augmented data is input into the lightweight YOLOv11 network, and a lightweight backbone network is used to extract features. The cross-stage partial attention selection module and the dynamic attention mechanism are combined for feature optimization. Multiscale predictions are generated through the object detection head, including class probabilities, object confidences, and bounding box parameters. Finally, the gradient descent method and optimization algorithms are used to update the network parameters, gradually improving the detection accuracy of the model. Among them, a multi-task weighted loss function is used as the loss function, which is expressed as:
[0147] L = λ cls L cls + λ obj L obj + λ box L box
[0148] The classification loss L cls Adopts the cross-entropy loss, which is used to measure the prediction accuracy of the target class:
[0149]
[0150] where y i is the true class label, is the predicted class probability.
[0151] The object confidence loss L obj Adopts the binary cross-entropy loss (BCE Loss), which is used to measure the model's judgment on the existence of the object:
[0152]
[0153] where is the predicted object confidence.
[0154] The bounding box regression loss L box Adopts the CIoU loss, which is used to optimize the matching degree between the predicted box and the true box:
[0155] L box = 1 - CIoU
[0156] where the CIoU calculation method takes into account IoU, the distance between the center points, and the aspect ratio, improving the regression stability.
[0157] When the loss index output by the training result gradually decreases and flattens out, and the mAP index gradually increases and also flattens out, it indicates that the deep learning network model has finished learning on this dataset and can no longer learn new things. At this time, it shows that the model has been fitted.
[0158] The lightweight model ADAT-YOLOv11 is trained based on the dataset labeled according to the steps in S1 and the improved network structure in S2 to obtain model weights, which are then used for human key point detection and then passed into the ST-GCN action recognition model for human fall detection. In this embodiment, the training environment is Python-3.9.19 torch-2.3.1 CUDA:0 (NVIDIA GeForce RTX 4060 Laptop GPU, 8188MiB), the number of training epochs is set to 300, the batch_size is set to 8, and finally a trained model is obtained. After the completion of the training of ADAT-YOLOv11, compared with YOLOv11 under the same training environment and parameter settings, both the parameters and GFLOPs have decreased by more than 16%, and the average detection accuracy mAP_0.5 has increased by 0.2%.
[0159] In step S4 of this embodiment, the processing and recognition process of the ST-GCN model is as follows:
[0160] The data processing and recognition process of ST-GCN mainly includes several key steps such as data preprocessing, graph construction, feature extraction, spatio-temporal information fusion, and classification. Specifically:
[0161] S41. Data preprocessing. For data preprocessing, the bone points of each frame of the video need to be output and assembled together to form a complete action. Therefore, the general input of the skeleton-based action recognition method is the time-continuous human skeleton key points. In this embodiment, the ADAT-YOLOv11 model weights trained by the previous improvement are used to label the human key nodes, and after connecting them to form a skeleton, they are input into the ST-GCN action recognition model for road pedestrian fall detection.
[0162] A graph partitioning strategy is adopted to establish multiple adjacency matrices reflecting different motion states; considering the characteristics of action recognition, instead of using a single convolution kernel, a graph partitioning strategy is used, that is, multiple adjacency matrices reflecting different motion states (such as static, centrifugal motion, and centripetal motion) are established. Figure 2 A comparison schematic diagram of graph partitioning for human key nodes using different strategies is shown, where Figure 2 (a) is a key frame of the input skeleton, Figure 2 (b) is a schematic diagram of the partitioning result using the unique partitioning method (Uni-labeling), that is, all nodes adjacent to the root node have the same label; Figure 2 (c) is a schematic diagram of the partitioning result using the distance partitioning method (Distance partitioning); that is, the label of the root node itself is set to 0, and its adjacent nodes are set to 1; Figure 2(d) is a schematic diagram of the division result of the spatial configuration partitioning method adopted in this embodiment. That is, based on the distance between the root node and the centroid (label = 0), among the distances from all adjacent nodes to the centroid, those less than the reference value are regarded as nodes towards the centroid (label = 1), and those greater than the reference value are regarded as centrifugal nodes (label = 2).
[0163] In Figure 2 (d), the first part connects neighbor nodes (yellow nodes) that are further away from the centroid of the entire skeleton in terms of spatial position than the current node, and contains the characteristics of centrifugal motion. The second part connects neighbor nodes (blue nodes) that are closer to the centroid, and contains the characteristics of centripetal motion. The third part connects the root node itself (green node), and contains the characteristics of rest.
[0164] Using such a decomposition method, one graph is decomposed into three subgraphs. The convolutional kernels also change from one to three, that is, (1, 18, 18) becomes (3, 18, 18). The convolutional results of the three convolutional kernels respectively represent action characteristics at different scales. To obtain the convolutional result, it is only necessary to perform convolution using each convolutional kernel respectively and then perform weighted averaging.
[0165] S42: Use the processed video frames as input to the YOLOv11-DAT key point detection network. The network first detects the targets in the frame, identifies pedestrians and generates corresponding bounding boxes. At the same time, it uses DAT (dynamic attention mechanism) to optimize the key point feature extraction process, accurately locates the key points of the human skeleton, and forms complete human pose information. The detected key point data will be used as the input of the ST-GCN model for further action recognition analysis.
[0166] S43: Input the obtained key point data into the ST-GCN model. First, perform normalization processing through the normalization layer (IN-BN). The network structure of the ST-GCN model is as Figure 3 shown. The joint positions of joints vary greatly in different frames. If normalization is not performed, it is not conducive to the convergence of the algorithm. The joint positions in different batches and different frames basically follow a random distribution, and will not cause the normalization results of different batches to vary too much, resulting in fluctuations in accuracy. Therefore, normalization processing (IN-BN) is performed on the input matrix, and the normalization is carried out in the time and space dimensions, that is, the position characteristics (x, y, acc) of a joint in different frames are normalized.
[0167] Figure 4Shows the schematic diagram of the module connection structure of the ST-GCN model, where residual is the residual structure, GCN is the spatial graph convolution module, and TCN is the temporal graph convolution module. By alternately using GCN and TCN to transform the temporal and spatial dimensions and then adding a residual structure, the ST-GCN cell is formed. The complete ST-GCN model consists of a normalization part, ten sequentially connected ST-GCN cells, and a feature classification part including average pooling and a fully connected layer.
[0168] S44. Use a number of ST-GCN cells including spatial graph convolution modules and temporal graph convolution modules for feature extraction and spatio-temporal information fusion; where GCN and TCN are spatial graph convolution and temporal graph convolution respectively. Spatial graph convolution is a convolution operation for processing graph data, aiming to capture the spatial relationships between graph nodes through a graph convolutional network. Its basic idea is that the feature of a node is a weighted combination of the features of its neighbor nodes.
[0169] Taking the input graph G=(V, E) as an example, V is the set of nodes, representing each joint point of the human body; E is the set of edges, representing the connection relationships between joint points; each node v i has a feature x i , and there is a corresponding adjacency matrix A in the graph, where A ij represents the weight connection between node v i and v j . The core formula of spatial graph convolution can be expressed as:
[0170]
[0171] X′ is the output feature matrix after convolution, X is the input feature matrix, with a shape of N×C, N is the number of nodes, and C is the feature dimension of each node. is the normalized adjacency matrix, which is obtained by preprocessing the adjacency matrix to ensure effective weighted aggregation of the features of each node with those of its neighbors.
[0172]
[0173] Among them, A is the original adjacency matrix, D is the degree matrix of nodes, that is W is the trainable weight matrix used to learn the weights of each node feature.
[0174] Temporal graph convolution uses a one-dimensional convolution (1D Convolution) operation, that is, sliding window processing of features in the temporal dimension. The formula is as follows:
[0175] X" = f(W t *X′ + b)
[0176] where X′∈R N×C×T is the input time - series feature, N is the number of nodes, C is the feature dimension, and T is the time step; W t is the Temporal Convolution Kernel, used to learn temporal dependencies; * represents a one - dimensional convolution operation, that is, sliding a window in the time dimension T for weighted summation; b is a trainable bias term, and f(·) is an activation function, which is used for non - linear transformation; for example, the ReLU function can be used as the activation function.
[0177] In summary, the complete calculation process of the ST - GCN model combining temporal - spatial graph convolution is as follows:
[0178]
[0179] where represents performing spatial graph convolution to extract spatial relationships; W t *(·) performs temporal graph convolution to extract temporal dependencies, and X out is the output feature after fusing spatial information and temporal information.
[0180] S45. Input the fused spatio - temporal features into the fully - connected layer, classify pedestrian behaviors through the classification layer, and determine whether a pedestrian has fallen according to the output result. If the model output is "fallen", then label "fall" on the image frame, frame the fallen person's portrait, and trigger the corresponding alarm or subsequent processing process; if it is "walking normally", then continue with subsequent behavior monitoring.
[0181] Use the trained lightweight ADAT - YOLOv11 combined with the ST - GCN motion detection model to detect the road pedestrian image data collected in step S1 on the TX2 computing processor of the patrol robot, and output the detected pictures; as Figure 5 is an example of the output of pedestrian fall detection in this embodiment. Figure 6 This is a comparison chart of the pedestrian fall detection indicators of the improved lightweight model ADAT - YOLOv11 of the present invention and the existing YOLOv11 model under the same dataset. Among them, Figure 6 is a comparison schematic diagram of the accuracy index mAP and the computational volume index GFLOPs of the training results of each model. Among them, Figure 6 (a) is the training result of the ADown - YOLOv11 model only after improving the ADown downsampling; Figure 6 (b) is the training result of the original unimproved YOLOv11 model; Figure 6(c) The training result of the brand-new ADAT-YOLOv11 architecture of this patent is formed by adding the improved C2PSA-DAT module to the ADown-YOLOv11 model. Obviously, there has been a certain improvement in both accuracy and the number of detection parameters, indicating the superiority of the architecture proposed in this invention for pedestrian fall detection on mobile robots, as shown in Table 1 below:
[0182] Table 1
[0183]
[0184] As can be seen from Table 1, ADAT-YOLOv11 has better lightweight performance and higher accuracy.
[0185] Embodiment 2
[0186] This embodiment provides a lightweight pedestrian fall detection system for inspection robots for performing the aforementioned lightweight pedestrian fall detection method for inspection robots, which at least includes a robot body, a housing, front wheels, a support frame, a movement control component, and various sensors. The robot body is protected by the housing, the front wheels are used for movement, the support frame is used to support the body, and the movement control component and the support frame are respectively fixedly connected to the housing. The platform core uses a NUC host as the computing processor and runs the ROS system as the robot operating system. To achieve the perception and monitoring of the environment, the platform is equipped with sensors such as a lidar and an Intel RealSense D435i camera. Through the ROS communication framework, real-time communication and data processing are achieved between the sensor data and the host. When the computer receives a control signal, the robot can select the automatic cruise or manual operation mode. The control signal is transmitted to the DC motor via the walking control drive module to achieve flexible control of various speeds and turning radii. During the cruise, the camera continuously collects pedestrian posture image data, and image processing and analysis are performed by the NUC processor. The detection result will be displayed on the display screen, thus completing the automatic detection task of the inspection robot for pedestrian fall situations.
[0187] In this embodiment, by improving the YOLOv11 object detection model to have a more lightweight network structure, the computational burden is reduced, enabling it to operate efficiently on resource-constrained embedded devices and ensuring the real-time nature of the detection process. At the same time, to improve the accuracy of fall detection, the ST-GCN model is introduced to make full use of spatio-temporal features to accurately model the dynamic postures of pedestrians, so that the system can not only detect the static postures after falling, but also capture the entire process of the fall occurrence, thereby effectively distinguishing normal walking, sitting down, lying down and other actions from real fall events.
[0188] The working process of the system covers video data acquisition, object detection, action recognition, as well as fall determination and alarm. When a pedestrian fall is detected, the system will automatically issue an alarm and transmit relevant information to the monitoring platform or management center in real time, so that relevant personnel can take rescue measures in a timely manner. This system has the characteristics of high precision, low latency, strong environmental adaptability, and high degree of automation, and can be widely applied to scenarios such as smart homes, nursing homes, hospitals, public safety monitoring, etc., providing a more intelligent and precise fall detection solution for society.
[0189] Through the present invention, pedestrian fall detection will no longer rely on traditional manual patrols or single-sensor detection, but will utilize artificial intelligence and deep learning technologies to achieve all-weather and efficient intelligent monitoring, thereby reducing the risks brought by fall accidents, improving the level of social public safety, and providing a scientific and reasonable decision-making basis for relevant management departments, promoting the development of intelligent security technologies.
[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A lightweight pedestrian fall detection method for an inspection robot, characterized in that: The method comprises the following steps: S1, inspect the pedestrian posture image data of the patrol robot and pre-process the data; S2. Build a lightweight YOLOv11 network, which includes at least several ADown lightweight downsampling modules and C2PSA_DAT composite modules; S3. Train the lightweight YOLOv11 network to obtain the optimal weights, thereby establishing a pedestrian key node fast detection model, and then obtain pedestrian key node data in real time through the pedestrian key node fast detection model; S4. The ST-GCN model, a dynamic skeleton-based action recognition method, is used to process the acquired pedestrian key node data and perform pedestrian fall detection in real time.
2. A lightweight pedestrian fall detection method for an inspection robot according to claim 1, characterized in that: In step S1, data preprocessing includes data labeling and data partitioning, wherein the LabelImg method is used to label the original pedestrian posture image data, the labeling category is pedestrian, and the labeling content is several key nodes of the pedestrian, forming a data set required for training the human key node detection model; data partitioning includes dividing the labeled human posture image data and label files into a training set and a test set.
3. The lightweight pedestrian fall detection method for an inspection robot according to claim 1, characterized in that: In step S2, the structure of the constructed lightweight ADAT-YOLOv11 model includes at least a backbone network and a head network, wherein the backbone network includes at least several Adown modules, Conv modules, C3K2 modules, and at least one SPPF module and a C2PSA-DAT module. In the backbone network, multi-scale feature extraction is performed by combining several groups of ADown modules, C3K2 modules and C2PSA-DAT modules, wherein the ADown module performs a lightweight convolution operation and the C3K2 module performs abstract processing on the feature information; the extracted highest scale feature is input into the SPPF module and the C2PSA-DAT module to extract key information; The head network includes at least several Upsample modules, Concat modules, C3K2 modules, ADown modules and Pose detection modules. Upsample is first performed through the Upsample module, and then feature maps of different resolutions are spliced through the Concat module. C3K2 and ADown are used to further process the features to obtain key node data of the human body.
4. The lightweight pedestrian fall detection method for an inspection robot according to claim 3, characterized in that: In the backbone network of the lightweight ADAT-YOLOv11 model, the Conv module first performs preliminary processing of the input features, and then the Adown module performs data lightweight processing. The process is as follows: Where X∈R H×W×C is the input feature map, is a learnable dimensionality reduction weight matrix, F(·) is a nonlinear transformation function, It is the feature representation output by the Adown module; The C3K2 module then enhances the feature extraction capability, assuming that its input feature map is X in , then the processing process is: X1=σ(Conv 1×1 (X in )) X2=σ(Conv 3×3 (X1)) X3=σ(DWConv 2×2 (X1)) X Concat =(X2,X3) X out =σ(Conv 1×1 (X concat )) In the formula, X1 is the result of the first layer of convolution, X2 and X3 are the two branches of the second layer of convolution, and X Concat The third layer is used to connect and merge the two branches. out Conv is the output feature representation after adjusting the number of channels through convolution; 1×1 (·) is a 1×1 convolution kernel, Conv 3×3 (·) represents a 3×3 convolution kernel, DWConv 2×2 (·) represents a 2×2 convolution kernel, σ(·) represents the fusion function; After multiple Adown modules and C3K2 modules are processed alternately, the SPPF module extracts key features and fuses multiple scale features. The feature representation of the input SPPF module is set as X i ' n , then the processing process is expressed as: X′ out =Concat(MaxPool(X′ in ,C1),MaxPool(X′ in ,C2),MaxPool(X′ in ,C3)) In the formula, MaxPool(·) is the maximum pooling operation, C1, C2, C3 are different scales of input features; Concat(·) is the concatenation function, X′ out It is the feature representation output by the SPPF module; Finally, the C2PSA-DAT module performs cross-stage partial feature selection and spatial and channel attention adjustment. The process is as follows: In the formula, They represent the feature inputs of the mth layer and the m-1th layer respectively, α represents the feature weight, and X s ' m+1 represents the feature output of the m+1th layer, A s and A c Represent the spatial and channel attention weights respectively, and the symbols represents element-wise weighting; represents the final output representation after spatial and channel attention adjustment; In the head network of the lightweight ADAT-YOLOv11 model, the feature processing method of the C3K2 module and the Adown module is the same as that of the C3K2 module and the Adown module in the backbone network; The Concat module enhances the model's expressiveness by connecting and fusing cross-layer features, retaining low-level and high-level feature information. high ∈R H×W×C and low feature map X low ∈R H×W×C Perform alignment fusion and adjust the size consistency by upsampling or downsampling. Assume that the low-level feature map is adjusted to the same spatial dimension as the high-level feature map by upsampling, and then the two are concatenated and fused as the input for subsequent feature extraction: X low_upsamp =Upsample(X low ,H',W') X fusion =Concat(X low_upsamp ,X high ) In the formula, X low_upsamp represents the spatial dimension of the adjusted low-level feature map, H′, W′ respectively represent the height and width of the output image, X fusion It is the fusion feature representation after splicing.
5. The lightweight pedestrian fall detection method for an inspection robot according to claim 3, characterized in that: The ADown module receives input image data, time series signals or other sensor data; the ADown module divides the input data into multiple small blocks or segments, wherein the division method includes at least a sliding window and a segmentation processing method; when processing image data, the image is cut into multiple small areas, and each area is downsampled separately; when processing time series data, the data is cut according to time periods, and each time series data is downsampled separately, wherein, Assume that the input feature size of the ADown module is H×W×C, where H and W represent the height and width of the feature map, and C is the number of channels. After the ADown module reduces the dimension of the input by learning the matrix weight W, the size of the convolution parameters for the same feature map is: P ADown =C in ×C out ×R Where R represents the number of channels after dimensionality reduction.
6. The lightweight pedestrian fall detection method for an inspection robot according to claim 3, characterized in that: The C2PSA-DAT module combines the cross-stage partial attention selection module C2PAS and the dynamic attention mechanism DAT for feature processing. C2PAS transfers part of the feature map to the next layer through a cross-stage connection method and performs weighted selection on the feature map; DAT automatically adjusts the attention weight according to the different characteristics of the input data; where: The C2PSA-DAT module contains the C2PSA module, which transfers some features through cross-stage connections and performs weighted selection on the features; let the input feature map of the mth layer be X (m) ∈R H×W×C , the input feature map of the m+1th layer is composed of a part of the feature map X of the current layer s (m) And some feature maps X of the previous layer s (m-1) Together, we select the matrix M s By transferring some features across stages, the cross-stage feature selection can be expressed as: Among them, M s ∈{0, 1} H×W×C It is a binary selection matrix that selects the channels to be retained to the next stage. The symbol · represents element-wise dot product.
7. The lightweight pedestrian fall detection method for an inspection robot according to claim 3, characterized in that: In step S3, the training process for the constructed lightweight YOLOv11 network is: S31, perform data enhancement on the input data, including random scaling, random flipping, mixup and Mosaic enhancement; S32, input the enhanced data into the lightweight YOLOv11 network, use the lightweight backbone network to extract features, and combine the cross-stage partial attention selection module and dynamic attention mechanism to perform feature optimization; S33, generating multi-scale predictions through the object detection head, including class probability, object confidence and bounding box parameters; S34, using gradient descent method and optimization algorithm to update network parameters and gradually improve the detection accuracy of the model; Among them, multi-task weighted loss is used as the loss function, which is expressed as: L=λ cls L cls +λ obj L obj +λ box L box Classification loss L cls Cross entropy loss is used to measure the prediction accuracy of the target category: where y i is the true category label, is the predicted category probability; Target confidence loss L obj The binary cross entropy loss is used to measure the model's judgment on the existence of the target: in is the confidence level of the predicted target; Bounding box regression loss L box CIoU loss is used to optimize the matching degree between the predicted box and the real box: L box =1-CIoU The CIoU calculation method takes into account IoU, center point distance and aspect ratio.
8. The lightweight pedestrian fall detection method for an inspection robot according to claim 7, characterized in that: In step S4, the following steps are specifically included: S41, data preprocessing, outputting the skeleton points of each frame of the video, and assembling them to form a complete action; wherein, a graph partitioning strategy is adopted to establish multiple adjacency matrices reflecting different motion states; the key node data of the human body is divided by a spatial configuration partitioning method, that is, the distance between the root node and the center of gravity is used as a reference, and among all the adjacent nodes to the center of gravity, those whose distances are less than the reference value are regarded as nodal centroids, and those whose distances are greater than the reference value are regarded as centrifugal nodes; S42, using the processed video frame as input to the YOLOv11-DAT key point detection network, the network detects the target in the frame, identifies the pedestrian and generates the corresponding bounding box, and uses the dynamic attention mechanism to optimize the key point feature extraction process, accurately locate the key points of the human skeleton, and form complete human posture information; S43, input the acquired key point data information into the ST-GCN model, firstly perform normalization processing through the normalization layer, wherein the normalization is performed in the time and space dimensions, and the position features (x, y, acc) of a joint in different frames are normalized; S44, using several ST-GCN unit feature extraction and spatiotemporal information fusion including a spatial graph convolution module and a temporal graph convolution module to obtain fused spatiotemporal information features; S45. Input the fused spatiotemporal features into the fully connected layer, classify the pedestrian behavior through the classification layer, and judge whether the pedestrian falls based on the output results; if the model output is "fall", mark the image frame with fall and frame the fallen person, trigger the corresponding alarm or subsequent processing flow; if it is "normal walking", continue with subsequent behavior monitoring.
9. The lightweight pedestrian fall detection method for an inspection robot according to claim 8, characterized in that: In step S44, the ST-GCN unit includes an alternately arranged spatial graph convolution GCN, a temporal graph convolution TCN, and a residual structure residual, wherein the input graph is G = (V, E), V is a node set, representing each joint point of the human body; E is an edge set, representing the connection relationship between the joint points; each node v i With feature x i , and the graph has a corresponding adjacency matrix A, where A ij Represents node v i and v j The weight connection between them; Then the core formula of spatial graph convolution is expressed as: X′ is the output feature matrix after convolution, X is the input feature matrix, the shape is N×C, N is the number of nodes, C is the feature dimension of each node; is the normalized adjacency matrix, which is expressed as: Among them, A is the original adjacency matrix, and D is the degree matrix of the node, that is, W is a trainable weight matrix used to learn the weight of each node feature; The temporal graph convolution uses a one-dimensional convolution operation, that is, a sliding window process is performed on the features in the time dimension. The formula is as follows: X″=f(W t *X′+b) Where X′∈R N×C×T is the input time series feature, N is the number of nodes, C is the feature dimension, and T is the time step; W t is a temporal convolution kernel, which is used to learn temporal dependencies; * represents a one-dimensional convolution operation, i.e., weighted summation of sliding windows on the time dimension T; b is a trainable bias term, and f(·) is an activation function, which is used for nonlinear transformation; The final output is expressed as: in Indicates spatial graph convolution to extract spatial relationships; W t *(·) Perform temporal graph convolution to extract temporal dependencies, X out It is the output feature after fusing spatial information and temporal information.
10. A detection system for executing the lightweight pedestrian fall detection method for a patrol robot as described in any one of claims 1 to 9, characterized in that: The system includes a robot body, a shell, front wheels, a support frame, a mobile control component and various sensors; The robot body is protected by a shell, the front wheels are used for movement, the support frame is used to support the body, and the mobile control component and the support frame are respectively fastened to the shell; The core of the platform uses the NUC host as the computing processor and runs the ROS system as the robot operating system.
Citation Information
Cited By
Campus safety analysis early warning method and system fused with multi-modal reasoning capability
CN120599772A
Campus safety analysis and early warning method and system fusing multi-modal reasoning capabilities
CN120599772B