Fall detection method and system based on robot vision

By installing a camera on a mobile robot to collect multi-angle human posture image data, combining the direction gradient histogram features, grayscale symbiosis matrix features and EMA mechanism, falling detection is performed using the YOLOv5 model, which solves the problems of environmental impact and wear dependence in the existing technology, and achieves high-accuracy fall detection.

CN120183034APending Publication Date: 2025-06-20GUANGXI LVFA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149354.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing fall detection methods based on wearable sensors are susceptible to environmental factors and rely on users' active wear, which has the disadvantage of forgetting to wear or improperly wearing it.

Method used

The fall detection method based on robot vision is adopted, and the camera equipped with a mobile robot collects multi-angle human posture image data, performs preprocessing, feature extraction, feature fusion and model training, and uses the YOLOv5 model to judge the fall posture.

Benefits of technology

Accurate detection of fall posture is achieved, environmental impact is reduced, detection accuracy is improved, and error detection and missed detection rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183034A_ABST
    Figure CN120183034A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a fall detection method and system based on robot vision, and the method comprises the steps: obtaining multi-angle human body posture image data collected by a camera carried by a mobile robot; preprocessing the human body posture image data; extracting a histogram of oriented gradient feature and a gray-level co-occurrence matrix feature of the processed human body posture image data; performing cross-dimension interaction and weight calibration on the histogram of oriented gradient features and the gray-level co-occurrence matrix features through an EMA mechanism to generate fusion features; inputting the fusion features into a YOLOv5 model for training to obtain a trained fall detection model; and inputting the fusion features of the to-be-detected image into the trained tumble detection model, and judging whether a tumble posture exists or not. According to the method, the histogram of oriented gradient features, the gray-level co-occurrence matrix features and the EMA mechanism are combined, the YOLOv5 model is used for training and prediction, and the detection rate of the falling posture is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a fall detection method and system based on robot vision. Background Art

[0002] With the aggravation of population aging, the safety issues of the elderly living alone have received increasing attention. As one of the common accidental injuries of the elderly, falling may cause serious physical injuries or even endanger life. Therefore, timely and accurately detecting the falling situation of the elderly is of great significance for ensuring the life safety of the elderly and improving the quality of life.

[0003] Regarding the fall detection of the elderly, some existing research is based on fall detection methods using wearable sensors. However, this method is easily affected by environmental factors such as temperature and humidity, and relies on the user to actively wear it, having the disadvantages of forgetting to wear or improper wearing. Summary of the Invention

[0004] Aiming at the problem that the fall detection method based on wearable sensors in the prior art is affected by environmental factors, the present invention provides a fall detection method and system based on robot vision, which can reduce the influence of the environment and achieve accurate detection of the fall posture. The specific technical solutions are as follows:

[0005] A fall detection method based on robot vision, comprising the following steps:

[0006] Obtain multi-angle human posture image data collected by a camera mounted on a mobile robot;

[0007] Preprocess the human posture image data;

[0008] Extract the histogram of oriented gradients features and gray-level co-occurrence matrix features of the processed human posture image data;

[0009] Perform cross-dimensional interaction and weight calibration on the histogram of oriented gradients features and gray-level co-occurrence matrix features through the EMA mechanism to generate fused features;

[0010] Input the fused features into the YOLOv5 model for training to obtain a trained fall detection model;

[0011] Input the fused features of the image to be detected into the trained fall detection model to determine whether there is a fall posture.

[0012] Preferably, the preprocessing of the human posture image data includes:

[0013] Perform graying, normalization, and background segmentation processing on the human posture image data.

[0014] Preferably, the acquisition of human body posture image data from multiple angles by the camera mounted on the mobile robot specifically refers to the acquisition of human body posture image data from different angles and scenarios by the dual cameras mounted on the mobile robot;

[0015] The dual cameras include a monocular camera and an infrared thermal camera; the human body posture image data includes normal posture image data and fall posture image data.

[0016] Preferably, the histogram of oriented gradients features of the processed human body posture image data includes:

[0017] Calculate the gradient magnitude and gradient direction of each image in the human body posture image data;

[0018] Limit the gradient direction within a preset range and divide it into multiple intervals to construct a histogram of the gradient direction of the cells;

[0019] Combine the histograms of the gradient directions of a preset number of adjacent cells into blocks, and normalize the histograms of the gradient directions of all cells within each block;

[0020] After connecting the normalized histograms of the gradient directions of each block in sequence, the histogram of oriented gradients features is obtained.

[0021] Preferably, the gray-level co-occurrence matrix features of the processed human body posture image data include:

[0022] Convert each image of the processed human body posture image data into a binary image;

[0023] Count the number of occurrences of the same pixels in the binary image to construct the gray-level co-occurrence matrix of each image;

[0024] Normalize the gray-level co-occurrence matrix;

[0025] According to the normalized gray-level co-occurrence matrix, calculate the texture features of contrast, energy, correlation and homogeneity to obtain the gray-level co-occurrence matrix features.

[0026] Preferably, the cross-dimensional interaction and weight calibration of the histogram of oriented gradients features and the gray-level co-occurrence matrix features through the EMA mechanism to generate fusion features includes:

[0027] Input the histogram of oriented gradients features and the gray-level co-occurrence matrix features into the EMA mechanism for channel and batch dimension recombination to obtain an initial fusion feature matrix;

[0028] Extract the global features of the initial fusion feature matrix through global average pooling operation;

[0029] Generate a weight vector by passing the global features through a fully connected layer;

[0030] Multiply the weight vector and the initial fused feature matrix element by element to generate the calibrated fused feature.

[0031] Preferably, the step of inputting the fused feature into the YOLOv5 model for training to obtain the trained fall detection model includes:

[0032] Input the fused feature into the YOLOv5 model, and the model performs forward propagation according to the input data to calculate the prediction result;

[0033] Calculate the loss value according to the prediction result and the preset loss function, and calculate the gradient through backpropagation;

[0034] According to the gradient, use the preselected optimizer to update the parameters of the model, and continuously repeat updating the parameters until the model performance reaches the preset index, forming the trained fall detection model.

[0035] Preferably, the prediction result includes the predicted bounding box and class probability.

[0036] Preferably, a fall detection method based on robot vision further includes:

[0037] When it is judged that there is a fall posture and after reaching the preset time, automatically send a distress signal.

[0038] A fall detection system based on robot vision, which is applied to the aforementioned fall detection method based on robot vision, includes:

[0039] A data acquisition unit, configured to acquire multi-angle human posture image data collected by a camera mounted on a mobile robot;

[0040] An image processing unit, configured to preprocess the human posture image data;

[0041] A feature extraction unit, configured to extract the histogram of oriented gradients feature and the gray-level co-occurrence matrix feature of the processed human posture image data;

[0042] A feature fusion unit, configured to perform cross-dimensional interaction and weight calibration on the histogram of oriented gradients feature and the gray-level co-occurrence matrix feature through the EMA mechanism to generate a fused feature;

[0043] A model training unit, configured to input the fused feature into the YOLOv5 model for training to obtain the trained fall detection model;

[0044] A judgment unit, configured to input the fused feature of the image to be detected into the trained fall detection model to judge whether there is a fall posture.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0046] The present invention proposes a fall detection method based on robot vision. By acquiring multi-angle human posture image data collected by a camera mounted on a mobile robot, and performing steps such as preprocessing, feature extraction, feature fusion, model training, and fall posture judgment, accurate detection of fall postures is achieved. This method combines features of histogram of oriented gradients, gray-level co-occurrence matrix, and the EMA mechanism, and uses the YOLOv5 model for training and prediction, improving the accuracy of fall detection while reducing the false detection rate and missed detection rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts do not necessarily draw according to the actual scale.

[0048] Figure 1 It is a flowchart of a fall detection method based on robot vision of the present invention.

[0049] Figure 2 It is a flowchart of an embodiment of a fall detection method based on robot vision of the present invention.

[0050] Figure 3 It is a flowchart of an embodiment of a fall detection method based on robot vision of the present invention.

[0051] Figure 4 It is a flowchart of an embodiment of a fall detection method based on robot vision of the present invention.

[0052] Figure 5 It is a flowchart of an embodiment of a fall detection method based on robot vision of the present invention.

[0053] Figure 6 It is a schematic diagram of a fall detection system based on robot vision of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0055] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0056] It should also be understood that the terms used in the specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0057] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0058] Please refer to the following embodiments Figures 1 to 6 .

[0059] The embodiment of the present application provides a fall detection method based on robot vision, including the following steps:

[0060] Step S1, acquiring multi-angle human posture image data collected by a camera mounted on a mobile robot;

[0061] Collecting human posture image data from different angles and scenarios through dual cameras mounted on a mobile robot;

[0062] The dual cameras include a monocular camera and an infrared thermal camera; the human posture image data includes normal posture image data and fall posture image data.

[0063] The monocular camera mounted on the mobile robot can capture human posture images under visible light and obtain rich texture and color information; the infrared thermal camera mounted on the mobile robot can sense the infrared radiation emitted by the human body, is not restricted by the light conditions, and can also form clear images at night or in low-light environments.

[0064] The mobile robot collects human posture image data at different angles and scenarios. For example, the human body is photographed from different perspectives such as the front, side, and back, and data is collected simultaneously in various scenarios such as indoors, outdoors, and different light intensities. The collected human posture image data is divided into normal posture image data and fall posture image data. To ensure the accuracy and diversity of the data, a large number of normal posture and fall posture images of different individuals and different actions need to be collected. Fall posture images can be obtained by simulating fall scenarios, collecting fall events in actual surveillance videos, etc., while normal posture images of people's daily activities are collected.

[0065] Step S2: Preprocess the human body pose image data;

[0066] Perform grayscale conversion, normalization, and background segmentation on the human body pose image data. Convert the captured color image into a grayscale image. The weighted average method can be used. According to the sensitivity of the human eye to different colors, the pixel values of the red, green, and blue channels are weighted and summed to obtain the grayscale value. Use the linear normalization method to map the pixel values of the grayscale image to a specific range. Separate the background and foreground (human body) in the image. For example, a threshold-based segmentation method can be used. By setting an appropriate threshold, the part with pixel values greater than the threshold is regarded as the foreground, and the part less than the threshold is regarded as the background; or a segmentation method based on the Gaussian mixture model (GMM) can be used to model and update the background in a dynamic scene to achieve more accurate background segmentation.

[0067] Step S3: Extract the histogram of oriented gradients (HOG) features and gray-level co-occurrence matrix (GLCM) features of the processed human body pose image data;

[0068] The histogram of oriented gradients (HOG) features can well describe the shape and contour information of the human body and have high sensitivity to the pose changes of the human body; the gray-level co-occurrence matrix (GLCM) features can capture the texture features of the image and reflect the local and global texture information of the human body pose image. The combination of the two can more comprehensively describe the human body pose and improve the accuracy of fall detection.

[0069] The histogram of oriented gradients (HOG) features and gray-level co-occurrence matrix (GLCM) features have strong robustness to illumination changes, small pose changes, etc., and can maintain good performance under different environmental conditions.

[0070] Step S4: Perform cross-dimensional interaction and weight calibration on the histogram of oriented gradients (HOG) features and gray-level co-occurrence matrix (GLCM) features through the EMA mechanism to generate fused features;

[0071] The EMA mechanism can effectively fuse the histogram of oriented gradients (HOG) features and gray-level co-occurrence matrix (GLCM) features. Through cross-dimensional interaction and weight calibration, different features can complement and enhance each other. At the same time, the EMA mechanism can adaptively adjust the weight of each dimension according to the historical information of the features, highlight the important feature dimensions, suppress noise and irrelevant information, thereby improving the performance of the fall detection model.

[0072] Step S5: Input the fused features into the YOLOv5 model for training to obtain a trained fall detection model;

[0073] By dividing the fused feature dataset into a training set, a validation set, and a test set, for example, in the proportions of 70%, 20%, and 10%. Set the pre-trained YOLOv5 model weights and initialize the model parameters. Set the hyperparameters for model training, such as the learning rate, batch size, number of training epochs, etc. The learning rate controls the step size of model parameter updates, the batch size determines the number of samples used in each training, and the number of training epochs represents the number of times the model trains on the entire dataset.

[0074] Input the fused features of the training set into the YOLOv5 model for training. Use the backpropagation algorithm to calculate the gradient of the loss function and update the model parameters through an optimizer (such as SGD, Adam, etc.). During training, use the validation set to evaluate the model and adjust the hyperparameters according to the evaluation results to prevent overfitting or underfitting of the model. After training, use the test set to evaluate the trained model and calculate metrics such as the accuracy, recall, and F1 value of the model to evaluate the performance of the model.

[0075] After the division and evaluation of the training set, validation set, and test set, the model can maintain good performance on different datasets and can adapt to fall detection tasks in different scenarios.

[0076] Step S6: Input the fused features of the image to be detected into the trained fall detection model to determine whether there is a fall posture.

[0077] Input the fused features of the image to be detected into the trained fall detection model. The model makes inferences based on the learned feature patterns, outputs the detection results, and determines whether the human body in the image is in a fall posture. According to the output results of the model, output the judgment results, such as "fall" or "normal", and the detected human body and fall state can be marked on the image by combining visualization techniques.

[0078] It should be noted that the specific methods for grayscaling and normalizing the human body posture image data are as follows:

[0079] For grayscaling, according to the different sensitivities of the human eye to different colors, different weights are assigned to the red, green, and blue color components for weighted averaging. For example, the color weights of red, green, and blue are assigned as 0.299, 0.587, and 0.114, and the calculation formula is as follows:

[0080] G ray = 0.299R + 0.587G + 0.114B

[0081] In the formula, G ray represents the grayscale value; R, G, and B represent the pixel values of the red, green, and blue channels respectively.

[0082] Normalization normalizes the pixel values of a grayscale image to [0,1] or [-1,1], and the calculation formula is as follows:

[0083]

[0084] Among them, X norm represents the normalized value; X is the original pixel value; X min and X max are the minimum and maximum pixel values in the image respectively.

[0085] The present invention proposes a fall detection method based on robot vision. By acquiring multi-angle human posture image data collected by a camera mounted on a mobile robot, and performing steps such as preprocessing, feature extraction, feature fusion, model training, and fall posture judgment, accurate detection of fall postures is achieved. This method combines features of histogram of oriented gradients, gray-level co-occurrence matrix, and the EMA mechanism, and uses the YOLOv5 model for training and prediction to improve the accuracy of fall detection while reducing the false detection rate and missed detection rate.

[0086] Specifically, in a preferred embodiment of the present application, the extraction of the histogram of oriented gradients features of the processed human posture image data includes:

[0087] Step S300: Calculate the gradient magnitude and gradient direction of each image in the human posture image data;

[0088] For the normalized image, use a gradient operator (such as the Sobel operator) to calculate the gradient of each pixel point. The Sobel operator contains templates in the horizontal and vertical directions, which are used to calculate the horizontal gradient G x and the vertical gradient G y ;

[0089] For each pixel point (i, j) in the image, the horizontal gradient G x (i, j) and the vertical gradient G y (i, j) are obtained through the convolution operation of the template with the pixel point and its neighboring pixels;

[0090] According to the calculated horizontal and vertical gradients, calculate the gradient magnitude G and direction θ of each pixel point;

[0091]

[0092] The gradient magnitude represents the intensity of the gray-level change at this pixel point, and the gradient direction reflects the direction of the gray-level change.

[0093] Step S301: Limit the gradient direction within a preset range and divide it into multiple intervals to construct a gradient direction histogram of the cells;

[0094] The gradient direction is divided into a preset number of bins in the range of [0, 180°) (for unsigned gradients) or [0, 360°) (for signed gradients). For example, [0, 180°) is divided into nine bins, with each bin covering 20°.

[0095] For each pixel point in each cell, the gradient magnitude is assigned to the corresponding bin according to its gradient direction. For example, if the gradient direction of a certain pixel point is 30°, its gradient magnitude is accumulated into the bin corresponding to 20° - 40°. Through the above operations, each cell can obtain a histogram of oriented gradients containing several bins, which reflects the distribution of gradient directions within the cell.

[0096] Step S302: Combine the histograms of oriented gradients of a preset number of adjacent cells into blocks, and normalize the histograms of oriented gradients of all cells within each block;

[0097] By combining adjacent cells into larger blocks, for example, a block consists of 2×2 cells. Normalize the histograms of oriented gradients of all cells within each block. The purpose of normalization is to extract consistent features from images under different lighting, contrast, and other conditions. Commonly used normalization methods include L1 norm normalization and L2 norm normalization. Taking L2 norm normalization as an example, let the vector composed of the histograms of oriented gradients of all cells within the block be V, then the normalized vector V norm is:

[0098]

[0099] where ∈ is a very small constant used to prevent the denominator from being zero.

[0100] Step S303: After connecting the normalized histograms of oriented gradients of each block in sequence, obtain the histogram of oriented gradients feature.

[0101] Connect the normalized histograms of oriented gradients of all blocks in the image in sequence to form a one-dimensional feature vector, and this vector is the HOG feature representation of the image. The HOG feature vector contains the gradient direction distribution information of different regions in the image and can be used to describe the shape and contour features of the human body posture.

[0102] In this embodiment, by calculating the gradient magnitude and direction, the edge information of the human body contour in the image can be accurately captured, which helps to distinguish normal postures from fall postures. The gradient direction is limited within a preset range and divided into multiple intervals to construct a gradient direction histogram of cells, so that the texture information of the image is enhanced, highlighting the texture features of the human body posture, which is crucial for identifying the texture changes in the image caused by falls. The gradient direction histograms of adjacent cells are combined into blocks and normalized, so that the features have better robustness, can effectively resist interference factors such as noise and illumination changes in the image, and improve the detection performance of the model in complex environments.

[0103] Specifically, in a preferred embodiment of the present application, the extraction of the gray-level co-occurrence matrix features of the processed human body posture image data includes:

[0104] Step S310: Convert each image of the processed human body posture image data into a binary image;

[0105] According to a preset threshold, pixels with pixel values greater than the threshold in the grayscale image are set to a fixed value (set to 255, representing white), and pixels less than or equal to the threshold are set to another fixed value (usually 0, representing white) to obtain a binary image.

[0106] Converting the image into a binary image can simplify the image information, highlight the main contours and structures of the human body posture, reduce the complexity of subsequent calculations, and facilitate the construction of the subsequent gray-level co-occurrence matrix.

[0107] Step S311: Count the number of occurrences of the same pixels in the binary image and construct the gray-level co-occurrence matrix of each image;

[0108] By defining the spatial distance between two pixels, for example, taking values of 1, 2, etc. For example, when d = 1, adjacent pixel pairs are considered. Specify the direction of the pixel pair, and the directions can be set to 0°, 45°, 90°, 135°. Different directions and distances will affect the texture information reflected by the gray-level co-occurrence matrix.

[0109] Since it is a binary image with only 2 gray levels (0 and 255), the gray-level co-occurrence matrix is a 2×2 matrix. Initialize all elements of the matrix to 0. For each pixel (x, y) in the binary image, find the corresponding another pixel (x′, y′) according to the set distance d and direction θ. For example, when the direction is 0° and the distance is 1, (x′, y′) = (x + 1, y). Count the number of occurrences of the pixel pair (I(x, y), I(x′, y′)), where I(x, y) represents the gray value of the pixel (x, y). Accumulate the number of occurrences to the corresponding element position in the gray-level co-occurrence matrix. For example, if I(x, y) = 0 and I(x′, y′) = 255, then add 1 to the element value at the position (0, 255) in the matrix. After traversing the entire binary image, the complete gray-level co-occurrence matrix is obtained.

[0110] Step S312: Normalize the gray-level co-occurrence matrix;

[0111] Add up the values of all elements in the gray-level co-occurrence matrix to get the sum S. For each element P(i, j) in the gray-level co-occurrence matrix, calculate the normalized element value P′(i, j) using the formula:

[0112] P′(i, j) = P(i, j) / S

[0113] The sum of all elements in the normalized gray-level co-occurrence matrix is 1.

[0114] Step S313: Calculate the contrast, energy, correlation, and homogeneity texture features based on the normalized gray-level co-occurrence matrix to obtain the gray-level co-occurrence matrix features.

[0115] Contrast: Reflects the degree of local gray change in the image, and the calculation formula is:

[0116]

[0117] where P'(i, j) is the element of the normalized gray-level co-occurrence matrix. The greater the contrast, the clearer the texture in the image and the more obvious the gray change.

[0118] Energy: Also known as angular second moment, reflects the uniformity of the image gray distribution and the thickness of the texture, and the calculation formula is:

[0119]

[0120] The larger the energy value, the more concentrated the distribution of elements in the gray-level co-occurrence matrix, and the more regular and smoother the texture of the image.

[0121] Correlation: Measures the linear correlation of grayscale values in an image. The calculation formula is as follows:

[0122]

[0123] where μ i and μ j are the means of i and j respectively, and σ i and σ j are the standard deviations of i and j respectively. The higher the correlation, the stronger the linear relationship of grayscale values in the image in the specified direction and distance.

[0124] Homogeneity: Reflects the local uniformity of the image grayscale distribution. The calculation formula is as follows:

[0125]

[0126] The greater the homogeneity, the better the local uniformity of the image.

[0127] In this embodiment, by calculating texture features such as contrast, energy, correlation, and homogeneity, the information in the gray-level co-occurrence matrix can be further quantified and refined to obtain specific values that can represent the texture features of the human body pose image. These features can be used for subsequent classification, recognition, and other tasks. For example, in a fall detection model, the human body pose state can be judged based on these features.

[0128] Specifically, in a preferred implementation manner of the present application, the cross-dimensional interaction and weight calibration of the histogram of oriented gradients features and the gray-level co-occurrence matrix features through the EMA mechanism to generate fused features include:

[0129] Step S41: Input the histogram of oriented gradients features and the gray-level co-occurrence matrix features into the EMA mechanism for channel and batch dimension reorganization to obtain an initial fused feature matrix;

[0130] Taking the extracted HOG features and GLCM features as inputs, which are respectively represented as F HOG and F GLCM . These feature vectors usually have different dimensions. The HOG features mainly capture the gradient direction information of the image, while the GLCM features mainly capture the texture information of the image.

[0131] The EMA mechanism reorganizes the channel and batch dimensions to convert the feature vectors F HOG and F GLCM into a form suitable for cross-dimensional interaction, and splices the feature vectors F HOG and F GLCM along the channel dimension to form an initial fused feature matrix Fconcat, which is specifically represented as follows:

[0132] Fconcat = Concat(F HOG , F GLCM )

[0133] These feature vectors usually have different dimensions. The HOG feature mainly captures the gradient direction information of the image, while the GLCM feature mainly captures the texture information of the image.

[0134] Step S42: Extract the global features of the initial fusion feature matrix through global average pooling operation;

[0135] The EMA mechanism uses global information encoding to calibrate the weights of each channel in the parallel branches. Specifically, through the global average pooling operation, the global features of the feature matrix Fconcat are extracted:

[0136] H = GAP(Fconcat)

[0137] The global feature H is used to capture the global information of the entire image, providing a basis for subsequent weight calibration.

[0138] Step S43: Generate a weight vector from the global feature through a fully connected layer;

[0139] The EMA mechanism performs cross-dimensional interaction to interact the global feature G with the feature matrix Fconcat to generate calibrated feature weights. The specific steps are as follows:

[0140] First, generate a weight vector W from the global feature H through a fully connected layer (FullyConnectedLayer, FC): W = FC(H). Then, multiply the weight vector W element-wise with the feature matrix Fconcat to generate a calibrated feature matrix Fcalibrated:

[0141] Fcalibrated = Fconcat ⊙ W

[0142] where ⊙ represents the element-wise multiplication operation.

[0143] Step S44: Multiply the weight vector element-wise with the initial fusion feature matrix to generate a calibrated fusion feature.

[0144] Use the calibrated feature matrix Fcalibrated as the fusion feature for subsequent training and prediction of the fall detection model.

[0145] In this embodiment, the EMA mechanism reorganizes the HOG features and GLCM features in the channel and batch dimensions, enabling the two types of features to complement each other and form a more comprehensive feature representation. The HOG features mainly capture the edge and shape information of the image, while the GLCM features reflect the texture information of the image. Through cross-dimensional interaction, the fused features can contain both types of information simultaneously, thus more accurately describing the image content. The global average pooling operation is used to extract global features and generate a weight vector, which multiplies each element of the initial fused feature matrix. This enhances important features and suppresses unimportant features, improving the discriminative ability of the model. The EMA mechanism can effectively fuse feature information at different scales, enabling the model to have better detection capabilities for human targets of different sizes and poses. In the fall detection task, the human pose changes significantly, and multi-scale feature fusion can improve the model's adaptability to these changes, thereby enhancing the robustness of the detection. Through weight calibration, the EMA mechanism can suppress noise features and reduce the impact of environmental interference on the model. In practical applications, there may be various noise and interference factors in the image, such as cluttered backgrounds and lighting changes. The EMA mechanism can effectively reduce the impact of these factors on fall detection. The global features extracted by the global average pooling operation can provide the overall information of the image, helping the model better understand the image content. In the fall detection task, the global information can assist the model in distinguishing human targets in different poses and improving the detection accuracy. By performing cross-dimensional interaction and weight calibration on the HOG features and GLCM features through the EMA mechanism, the generated fused features can enhance the feature expression ability, improve the robustness and detection accuracy of the model, and at the same time optimize the model performance, enabling it to maintain good detection results in complex environments.

[0146] Specifically, in a preferred embodiment of the present application, training the fused features in the YOLOv5 model to obtain a trained fall detection model includes:

[0147] Step S51: Input the fused features into the YOLOv5 model, and the model performs forward propagation based on the input data to calculate the prediction results;

[0148] Input the adapted fused features into the YOLOv5 model. The YOLOv5 model is an object detection model based on a convolutional neural network (CNN), which includes multiple convolutional layers, pooling layers, and fully connected layers, etc. After the input data, the data will be calculated through these layers in sequence. Specifically, the convolutional layer extracts the features of the input data through convolutional kernels, the pooling layer downsamples the feature map to reduce the data volume, and the fully connected layer converts the feature map into the prediction results.

[0149] The prediction results include predicted bounding boxes and class probabilities. The bounding boxes are used to locate the position of the target (such as a human body) in the image, and are usually represented by the center point coordinates, width, and height of the bounding box. The class probabilities represent the likelihood that the target within each bounding box belongs to different classes. In the fall detection task, the classes mainly include "fall" and "normal".

[0150] Through forward propagation, the model can detect and classify the targets in the image based on the input fused features, and output the predicted bounding boxes and class probabilities.

[0151] Step S52: Calculate the loss value according to the prediction results and a preset loss function, and calculate the gradient through backpropagation;

[0152] It is used to measure the difference between the predicted target confidence (i.e., the probability of the existence of a target within the bounding box) and the actual situation. By using the cross-entropy loss function, according to the predicted bounding boxes, class probabilities, and actual bounding boxes, class labels, combined with the preset loss function, the total loss value is calculated. The loss value reflects the degree of difference between the model's prediction results and the actual situation. The smaller the loss value, the more accurate the model's prediction.

[0153] Using the backpropagation algorithm, starting from the loss value, propagate backward along the computational graph of the model to calculate the gradient of the loss function with respect to each parameter of the model. The gradient represents the rate of change of the loss function at the current parameter values.

[0154] By calculating the loss value and the gradient, the prediction error of the model can be quantified, and the direction and magnitude of the adjustment of the model parameters can be determined.

[0155] Step S53: Update the parameters of the model according to the gradient, and continuously repeat the parameter update until the model performance reaches the preset metrics, forming a trained fall detection model.

[0156] Update the parameters in the opposite direction of the gradient by using Stochastic Gradient Descent (SGD). The optimizer updates the parameters of the model according to the calculated gradient according to a certain update rule. In SGD, the parameter update formula is:

[0157]

[0158] where ρ represents the parameters of the model, α is the learning rate, is the gradient of the loss function with respect to the parameter ρ old of.

[0159] Steps S51 - S53 are continuously repeated, that is, forward propagation, loss calculation, backward propagation, and parameter update are continuously performed until the model performance reaches the preset metrics. The preset metrics can be accuracy, recall, F1 - value, etc. on the validation set. During the training process, the performance of the model can be evaluated on the validation set regularly to observe whether the model has overfitting or underfitting phenomena. If overfitting occurs, measures such as increasing data augmentation and reducing model complexity can be taken; if underfitting occurs, attempts can be made to increase the training data, adjust the model structure or hyperparameters, etc.

[0160] By inputting the fused features into the YOLOv5 model for training, and using steps such as forward propagation, loss calculation, backward propagation, and parameter update, the model can learn the mapping relationship between the fused features of the human pose image and the fall pose, continuously optimize its own parameters, thereby improving the accuracy and reliability of fall detection. The trained fall detection model can be applied to actual scenarios, such as nursing homes, smart homes, etc., to monitor the human pose in real - time, detect fall events in a timely manner and send out alarms, providing strong support for ensuring people's safety.

[0161] Specifically, in a preferred implementation manner of the present application, a fall detection method based on robot vision further includes:

[0162] When it is judged that there is a fall - pose situation and after reaching the preset time, a distress signal is automatically sent.

[0163] The preset time is an important parameter, and its setting needs to consider various factors comprehensively. If the preset time is too short, false distress signals may be sent due to some short - term accidental actions (such as squatting down to tie shoelaces, quickly getting up after a short fall, etc.), causing unnecessary interference; if the preset time is too long, the rescue opportunity may be delayed, and assistance cannot be provided in a timely manner to those who have truly fallen and need help. Therefore, the preset time can be set according to the actual application scenario and experience. For example, in scenarios such as nursing homes, the preset time can be set to 30 seconds.

[0164] The timer continuously records the time when the target object is in a falling posture. When the timing reaches the preset time, the system triggers the sending mechanism of the distress signal. Before triggering, the system will reconfirm the current judgment result to ensure that it is still judged as a falling posture, so as to further reduce the false alarm rate. There are various ways to send the distress signal. For example, the robot sends the distress information to a specified server or monitoring center through a wireless network (such as Wi-Fi, 4G, 5G, etc.). The distress information can include the time and location of the fall (obtained through the robot's positioning system), images or video clips of the target object, etc., so that relevant personnel can understand the situation in a timely manner; or communicate with preset contacts (such as family members, medical staff, etc.). The relevant information of the fall event can be informed by text message, or the preset phone number can be directly dialed to convey the distress information in voice; the robot itself can also emit light and sound signals, such as flashing lights, emitting loud alarm sounds, etc., to attract the attention of surrounding people and provide help in a timely manner.

[0165] An embodiment of this application also provides a fall detection system based on robot vision, which is applied to the aforementioned fall detection method based on robot vision, and includes:

[0166] A data acquisition unit, configured to acquire multi-angle human posture image data collected by a camera mounted on a mobile robot;

[0167] An image processing unit, configured to preprocess the human posture image data;

[0168] A feature extraction unit, configured to extract the histogram of oriented gradients features and gray-level co-occurrence matrix features of the processed human posture image data;

[0169] A feature fusion unit, configured to perform cross-dimensional interaction and weight calibration on the histogram of oriented gradients features and gray-level co-occurrence matrix features through the EMA mechanism to generate fusion features;

[0170] A model training unit, configured to input the fusion features into the YOLOv5 model for training to obtain a trained fall detection model;

[0171] A judgment unit, configured to input the fusion features of the image to be detected into the trained fall detection model to judge whether there is a falling posture.

[0172] The function explanations of the units in this embodiment are the same as those of a fall detection method based on robot vision, and the technical effects are the same, so they will not be repeated here.

[0173] Those of ordinary skill in the art will appreciate that the units of each example described in connection with the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components of each example have been generally described in terms of function in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0174] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored, etc.

[0175] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0176] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. A fall detection method based on robot vision, characterized in that: The following steps are involved: Obtain multi-angle human posture image data collected by the camera carried by the mobile robot; Preprocessing the human body posture image data; Extract the directional gradient histogram features and gray-level co-occurrence matrix features of the processed human posture image data; The directional gradient histogram feature and the gray-level co-occurrence matrix feature are subjected to cross-dimensional interaction and weight calibration through the EMA mechanism to generate a fusion feature; Input the fusion features into the YOLOv5 model for training to obtain a trained fall detection model; The fused features of the image to be detected are input into the trained fall detection model to determine whether there is a fall posture.

2. The robot vision-based fall detection method according to claim 1, characterized in that: The preprocessing of the human body posture image data comprises: The human body posture image data is gray-scaled, normalized and background segmented.

3. The robot vision-based fall detection method according to claim 2, characterized in that: The obtaining of multi-angle human posture image data collected by the camera carried by the mobile robot specifically refers to obtaining human posture image data collected by the dual cameras carried by the mobile robot from different angles and scenes; The dual cameras include a monocular camera and an infrared thermal camera; the human body posture image data include normal posture image data and falling posture image data.

4. The robot vision-based fall detection method according to claim 3, characterized in that: The directional gradient histogram features of the human body posture image data after the extraction process include: Calculate the gradient magnitude and gradient direction of each image in the human posture image data; The gradient direction is limited to a preset range and divided into multiple intervals to construct a gradient direction histogram of the cell; The gradient direction histograms of a preset number of adjacent cells are combined into blocks, and the gradient direction histograms of all cells in each block are normalized; The normalized gradient direction histograms of each block are connected in sequence to obtain the directional gradient histogram feature.

5. The robot vision-based fall detection method according to claim 3, characterized in that: The gray level co-occurrence matrix features of the extracted human posture image data include: Convert each image of the processed human posture image data into a binary image; Count the number of times the same pixels appear in the binary image and construct the gray-level co-occurrence matrix of each image; Normalizing the gray level co-occurrence matrix; According to the normalized gray-level co-occurrence matrix, the contrast, energy, correlation and homogeneity texture features are calculated to obtain the gray-level co-occurrence matrix features.

6. The robot vision-based fall detection method according to claim 3, characterized in that: The cross-dimensional interaction and weight calibration of the directional gradient histogram feature and the gray level co-occurrence matrix feature by the EMA mechanism to generate the fusion feature includes: Input the oriented gradient histogram features and the gray-level co-occurrence matrix features into the EMA mechanism, reorganize the channel and batch dimensions, and obtain an initial fusion feature matrix; Through the global average pooling operation, the global features of the initial fusion feature matrix are extracted; Generate a weight vector by passing the global features through a fully connected layer; The weight vector is multiplied element-wise with the initial fused feature matrix to generate the calibrated fused features.

7. The robot vision-based fall detection method according to claim 3, characterized in that: The inputting the fusion features into the YOLOv5 model for training to obtain a trained fall detection model comprises: The fused features are input into the YOLOv5 model, and the model performs forward propagation based on the input data to calculate the prediction results; Calculate the loss value based on the prediction result and the preset loss function, and calculate the gradient through back propagation; According to the gradient, the pre-selected optimizer is used to update the parameters of the model, and the parameters are updated repeatedly until the model performance reaches the preset indicators to form a trained fall detection model.

8. The robot vision-based fall detection method according to claim 7, characterized in that: The prediction result includes a predicted bounding box and a category probability.

9. The robot vision-based fall detection method according to claim 1, characterized in that: Also includes: When a fall is detected, a distress signal will be automatically sent out after a preset time.

10. A fall detection system based on robot vision, characterized in that: The robot vision-based fall detection method applied to any one of claims 1 to 9 comprises: A data acquisition unit, used to acquire multi-angle human posture image data collected by a camera carried by the mobile robot; An image processing unit, used for preprocessing the human body posture image data; A feature extraction unit, used to extract directional gradient histogram features and gray-level co-occurrence matrix features of the processed human posture image data; A feature fusion unit, used for performing cross-dimensional interaction and weight calibration on the directional gradient histogram feature and the gray-level co-occurrence matrix feature through an EMA mechanism to generate a fusion feature; A model training unit, used for inputting the fusion features into a YOLOv5 model for training to obtain a trained fall detection model; The judgment unit is used to input the fusion features of the image to be detected into the trained fall detection model to judge whether there is a fall posture.