A remote monitoring robot for the health monitoring of the elderly

By designing a remote monitoring robot that integrates deep learning algorithms, mobile robots and emergency safety assurance modules, the problem of falling and going out to monitor when living alone is solved, real-time monitoring and emergency response for the elderly are achieved, and the safety of the elderly is significantly improved.

CN119115977BActive Publication Date: 2025-07-01SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411331381.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-07-01
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

The elderly are prone to accidents such as falling when they live alone due to lack of family care. The fixed camera of the traditional family camera cannot monitor the elderly’s outings in real time, resulting in the inability to detect and deal with emergency events in a timely manner.

Method used

A remote monitoring robot is designed, integrating deep learning algorithm module, mobile robot module, voice module, abnormality detection module, positioning module and emergency safety assurance module, collect video streams in real time through USB cameras, use YOLOv5 algorithm to perform fall detection, and realize the interaction between the elderly and the robot through voice module and mobile APP. The positioning module and emergency safety assurance module ensure timely response in emergencies.

Benefits of technology

Real-time monitoring of the elderly and high-precision detection of fall incidents is achieved, ensuring timely response in emergencies, and improving the safety and quality of life of the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119115977B_ABST
    Figure CN119115977B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote monitoring robot for the health monitoring of the elderly. The present invention relates to the technical field of monitoring robots and includes a mobile robot module, which is communicatively connected to a deep learning algorithm module, a voice module, an anomaly detection module, a positioning module, and an emergency safety guarantee module. This remote monitoring robot for the health monitoring of the elderly provides a safe living space for the elderly by integrating deep learning algorithms and mobile robot technologies, monitors the activities of the elderly in real time, and ensures that help can be obtained in a timely manner when the elderly encounter a health crisis through a fall detection and emergency feedback mechanism. When the elderly are detected to have fallen by the YOLOv5 algorithm, a help request message is immediately sent to the preset guardian. In addition, the mobile following function of the robot provides continuous guardianship when the elderly go out. Whether at home or outdoors, the robot can closely follow to ensure that there is no dead angle in guardianship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of monitoring robots, and particularly to a remote monitoring robot for elderly health monitoring. Background Art

[0002] With the deepening of the aging of the population, the social pension burden has increased, and it is very important to handle the pension problems of the whole society. Globally, the physical health of the elderly is threatened by many fatal diseases. Due to the lack of family care in daily life and psychological comfort, the prevalence rate of the elderly living alone is higher. At the same time, accidents cannot be detected and rescued in time. Due to the busy work of young people in daily life, the care of the elderly and the real-time monitoring of the family environment have become emerging needs.

[0003] However, the traditional home camera is restricted by a fixed camera position, which limits the monitoring effect and cannot take into account the real-time monitoring when the elderly go out. From the perspective of caring about the health of the elderly, in view of the above deficiencies, a remote monitoring robot for elderly health monitoring is proposed. It can monitor the elderly in real time and has a mobile function to realize a system that continuously follows the elderly and monitors them. Summary of the Invention

[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: A remote monitoring robot for elderly health monitoring includes a mobile robot module, which is communicatively connected to a deep learning algorithm module, a voice module, an anomaly detection module, a positioning module, and an emergency safety guarantee module. Among them, the electrical signals are connected between the modules;

[0005] The deep learning algorithm module collects video streams in real time through a USB camera, uses the YOLOv5 algorithm for fall detection, intelligently analyzes the video content, and identifies the body posture of the elderly. Once a fall behavior is detected, subsequent safety guarantee measures will be immediately triggered, improving the detection accuracy and response speed of elderly fall events, and effectively preventing serious consequences that may be caused by falls;

[0006] The mobile robot module integrates CSRT_Tracker tracking technology. According to the target position information provided by the deep learning algorithm module, it autonomously adjusts the moving direction and speed, maintains a safe distance from the elderly and continuously monitors, realizing real-time monitoring when the elderly go out. Whether the elderly are at home or out for activities, the robot can closely follow to ensure no dead corners in guardianship. The mobile robot module includes a mobile robot, a development board, a USB camera, and a microphone;

[0007] The voice module is used to deploy a large language model on the development board of the mobile robot module, extract and recognize the voice of the elderly in the ambient sound, including dialect recognition, enabling the elderly to send instructions to the robot in the form of voice conversations, such as "start following", "stop following", etc. In addition, a corresponding mobile phone APP is provided to facilitate the elderly to issue instructions from the mobile phone APP, simplifying the operation process and improving the usage experience of the elderly. The elderly do not need to learn complex operation interfaces or buttons and can control the robot only by voice, realizing convenient human-computer interaction;

[0008] The anomaly detection module is used to monitor the behavior of the elderly by combining the results of sound recognition and visual recognition, and detect and give early warnings of the occurrence of the elderly falling and emergency events;

[0009] The positioning module realizes instant communication between the robot and the emergency contact person, as well as the precise positioning of the robot, ensuring the rapid transmission of help information and the accurate arrival of rescue personnel, improving the rescue efficiency, reducing the risks faced by the elderly, and ensuring that help can be obtained in a timely manner even if an accident occurs when the elderly are out;

[0010] The emergency safety guarantee module is used to combine the abnormal behavior detection results of the anomaly detection module, immediately send help and positioning information to the emergency contact person, and play a help voice around. In case of an emergency, it starts the safety guarantee mechanism to provide timely and effective assistance to the elderly. At the same time, through sound recognition technology, the intelligence level of the robot is enhanced, making it more adaptable to the actual needs of the elderly.

[0011] Preferably, in the deep learning algorithm module, the process of recognizing the body posture of the elderly includes:

[0012] Install a USB camera on the mobile robot module, and continuously collect the video stream of the elderly's activities through the USB camera, continuously monitor the activity area of the elderly, and capture dynamic images;

[0013] Pre-collect a batch of action samples of falling and non-falling, and build a fall detection model according to the ratio of training set: test set = 9:1, learn the unique body posture change pattern when falling, and deploy the fall detection model in the deep learning algorithm module;

[0014] Segment the collected video stream into individual video frames, each frame being a static image, and preprocess the segmented video frame images, including resizing to match the model input size, normalizing pixel values, etc., to improve the model processing efficiency and accuracy;

[0015] Send the preprocessed video frames into the YOLOv5 algorithm for object detection to detect the human contour of the elderly in the video frame image and distinguish the elderly from other objects or backgrounds;

[0016] The deep learning algorithm module further analyzes the human posture in the human contour, detects the positions and relative relationships of key points (such as the head, shoulders, knees, etc.) of the elderly's body, and combines a preset fall detection model to determine whether the elderly has fallen.

[0017] Preferably, the process of determining whether the elderly has fallen includes:

[0018] Run the YOLOv5 algorithm on the video frame, and according to the detection of the human body in the video frame by the deep learning algorithm, output the bounding box of the human body;

[0019] Within the output bounding box area of the human body, apply a pose estimation algorithm (such as OpenPose, HRNet, etc.) to identify the key points of the human body, such as the head, shoulders, knees, ankles, etc., and represent the key points in the image in the form of two-dimensional coordinates Let the coordinates of the th key point be , where represents the index of the key point (such as the head is 1, the left shoulder is 2,..., the right ankle is N);

[0020] Extract fall features from the detected key points. Among them, the fall features include the distance, angle, and relative position between key points. According to the preset fall detection model, analyze whether the fall features conform to the behavior pattern during a fall. Among them, the distance between key points is calculated using the Euclidean distance formula, and its expression is: , is the distance between key points and , is the coordinate of key point , is the coordinate of key point ; the angle between key points is calculated using the vector dot product formula, and its expression is: , is the angle formed by key points , and , is the coordinate of key point , is the coordinate of key point and ;

[0021] Set threshold values for different fall features. When the feature values are continuously detected to be lower than the preset threshold values 5 times at the fall detection node, combined with the preset fall detection model, it is determined that the fall features conform to the behavior pattern during a fall.

[0022] Preferably, in the mobile robot module, the process of real-time monitoring when the elderly go out includes:

[0023] Before going out, the elderly turn on the robot and start the mobile robot through voice control or mobile phone APP operation. The deep learning algorithm module processes the video frame data and outputs the human key points of the elderly.

[0024] Use the pose estimation algorithm to output key point information, initialize the CSRT_Tracker tracking algorithm, and use CSRT_Tracker to track the position of the elderly in consecutive video frames, and update its bounding box coordinates in real time.

[0025] Obtain the real-time position information of the elderly from CSRT_Tracker, including the center point coordinates of the bounding box. The mobile robot uses its own sensors (such as wheel encoders, IMUs, etc.) to determine its location.

[0026] After identifying the target, calculate the pixel center coordinates of the target. If the coordinates appear on the right side of the USB camera, control the mobile robot to rotate to the right. Similarly for the left side. By continuously fine-tuning the rotation direction and speed, the target is always kept in the center of the screen.

[0027] According to the real-time position of the elderly and the current position of the mobile robot, use the path planning algorithm to calculate the best path for the mobile robot to reach the target position, and according to the path planning result, calculate the speed and direction that the mobile robot needs to adjust to maintain a safe distance from the elderly.

[0028] The drive system of the mobile robot performs corresponding movement operations according to the calculation results, including adjusting speed and steering. During the movement of the mobile robot, continuously use the USB camera and CSRT_Tracker to track the position of the elderly. The mobile robot combines the movement speed and direction of the elderly and dynamically adjusts the speed and position to ensure the smoothness and continuity during the following process and avoid losing the target.

[0029] Preferably, in the voice module, the process of extracting and recognizing the voice of the elderly from the ambient sound includes:

[0030] Pre-enter the command vocabulary preset by the elderly and store the preset command vocabulary to match the voice commands of the elderly. The voice commands are such as "start following" or "stop following".

[0031] Use the microphone equipped in the mobile robot module to continuously collect the ambient sound, and preprocess the collected ambient sound, including preprocessing operations such as noise reduction, echo cancellation, and gain control, to improve the accuracy of speech recognition.

[0032] Through voice activity detection technology, the effective voice part in the ambient sound is identified, and the identified effective voice part is recognized using the large language model deployed on the development board, handling different dialects and accents to recognize the voice text, improving the recognition accuracy and ensuring that the elderly can communicate with the robot in a natural and habitual language;

[0033] Perform natural language processing on the voice text, including semantic understanding and command classification, to ensure correct understanding of the elderly's instructions, and match the recognized voice text with the preset command vocabulary to parse the elderly's intentions;

[0034] Send the parsed instructions to the development board of the mobile robot module, and the development board controls the mobile robot to execute corresponding actions. After the robot executes the instructions, feedback is provided to the elderly synchronously through voice and the display screen to confirm that the instructions have been executed.

[0035] Preferably, in the anomaly detection module, the process of monitoring the elderly's behavior includes:

[0036] Collect audio and video data of the elderly's normal and abnormal behaviors (such as falling, calling for help, etc.) under different environments, lighting conditions, and background noises, and preprocess the collected audio and video data to mark the abnormal behavior data;

[0037] Analyze the preprocessed audio and video data, extract sound features and visual features from it, and respectively mark the abnormal sound features and abnormal action features in the sound features and visual features to obtain a sound feature set and a visual feature set;

[0038] Utilize the abnormal sound features in the sound feature set to perform sound intensity and frequency analysis on the sound frequency data, calculate the decibel value and frequency deviation of the abnormal sound features, obtain a sound feature weighted summation function, calculate the weights of all sampling points, find the reference value in the long-term trend, obtain a reference adjustment exponential function, and perform weighted calculation on the abnormal sound to obtain an abnormal sound index to evaluate the degree of the overall abnormal sound. Utilize the abnormal action features in the visual feature set, preset the normal behavior reference value, calculate the deviation between a single abnormal action feature and the normal behavior reference value, obtain a behavior deviation weighted summation function, calculate the abnormal scores of all video frames at all time points, find the maximum value in the long-term trend, obtain a time series maximum weighted abnormal score function, and perform weighted calculation on multiple abnormal action features to obtain an abnormal behavior index to evaluate the degree of the overall abnormal behavior;

[0039] Construct a training data set using the correlation data of the sound feature set and the visual feature set, and train an anomaly detection model through a neural network model to identify the features of normal and abnormal behaviors;

[0040] Based on the abnormal behavior data in the sound feature set and the visual feature set combined with the anomaly detection model, comprehensively analyze the abnormal sound index and the abnormal behavior index to obtain the abnormal warning coefficient, analyze the detected abnormal behavior, set the warning threshold, and determine the node that triggers the warning mechanism;

[0041] Use the microphone equipped with the mobile robot module to continuously collect ambient sounds, and use the USB camera installed on the robot to collect video streams in real time. Extract the sound features and visual features from them and input them into the trained anomaly detection model. According to the output of the anomaly detection model and the abnormal warning coefficient, determine whether to trigger the warning mechanism.

[0042] Preferably, the abnormal sound index is obtained based on the sound feature weighted summation function and the reference adjustment exponential function, and its expression is: Where, is the abnormal sound index, is the sound feature weighted summation function, is the reference adjustment exponential function, is the decibel value of the th sampling point, representing the intensity of the sound, is the frequency deviation of the th sampling point, representing the deviation of the sound frequency from normal speech, is the total number of sampling points, is the weight of the th sample, and the value of is between 0 and 1. 0 indicates no abnormality, and 1 indicates a high degree of abnormality. As and increase (i.e., the sound intensity and frequency deviation increase),

[0043] increases, indicating an increased likelihood of abnormality. Where, is the abnormal behavior index, is the behavior deviation weighted summation function, is the time series maximum weighted anomaly scoring function, is the measurement value of the rd action feature in the th video frame, is the normal behavior reference value of the th action feature, is the total number of frames of a single action feature, is the The standard deviation of the action features, For the The standard deviation of normal behavior of action features, is the number of action features, For the The weight of the video frame, For the The anomaly score of each video frame in the first evaluation, The value of is between 0 and 1, where 0 indicates no abnormality and 1 indicates high abnormality. As the action characteristics deviate from normal behavior, Increases, indicating an increased likelihood of an abnormality.

[0044] Preferably, the expression of the abnormal warning coefficient is: in, Abnormal warning coefficient, For in time The abnormal behavior index, is the abnormal sound index at time t, T is the total number of time points, and is the weight coefficient, which is used to adjust the influence of abnormal behavior index and abnormal sound index in the abnormal warning coefficient. is a preset reference constant used to adjust the impact of the exponential function. is the weight at the rth time point, is the abnormal behavior index at the rth time point, is the abnormal sound index at the rth time point, The value is between 0 and 1, where 0 indicates no abnormality and 1 indicates a high abnormality.

[0045] Preferably, in the emergency safety assurance module, the process of starting the safety assurance mechanism includes:

[0046] Use deep learning algorithm modules and anomaly detection modules to continuously analyze sound and visual data, monitor the behavior of the elderly, and trigger warning signals when abnormal behavior is identified;

[0047] When the warning signal is triggered, it automatically sends a help message to the preset emergency contact, including text messages, phone calls or app push notifications, and uses the built-in GPS of the mobile robot module to obtain the precise location of the elderly;

[0048] After sending the help message, a pre-recorded help voice is played in the environment around the elderly to attract the attention of nearby people and seek help, monitor the elderly's condition, and provide real-time updates to emergency contacts. After detecting that the abnormal condition is lifted, a message to lift the alarm is sent.

[0049] The present invention provides a remote monitoring robot for the health monitoring of the elderly. It has the following beneficial effects:

[0050] First, the remote monitoring robot for the health monitoring of the elderly integrates deep learning algorithms and mobile robot technologies to provide a safe living space for the elderly, monitors the activities of the elderly in real time, and ensures that help can be obtained in a timely manner when the elderly encounter a health crisis through a fall detection and emergency feedback mechanism. When the elderly are detected to have fallen by the YOLOv5 algorithm, a help message is immediately sent to the preset guardian. In addition, the mobile following function of the robot provides continuous guardianship when the elderly go out. Whether at home or outdoors, the robot can closely follow to ensure that there are no blind spots in guardianship, significantly improving the safety of the elderly.

[0051] Second, the remote monitoring robot for the health monitoring of the elderly analyzes video information and sound data to evaluate the behavior patterns of the elderly and issues a warning when abnormal behavior is detected. By comprehensively considering sound and visual information, a more accurate assessment of abnormal behavior is provided. If the elderly make abnormal sounds or their behavior patterns are significantly different from the normal patterns, the warning mechanism is triggered and the guardian is notified through the preset communication channels, enabling the guardian to understand the situation of the elderly in a timely manner and take necessary intervention measures to prevent potential risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is the overall flowchart of a remote monitoring robot for the health monitoring of the elderly according to the present invention;

[0053] Figure 2 is the block diagram of the present invention;

[0054] Figure 3 is the flowchart of the present invention for monitoring the behavior of the elderly. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments. The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the present invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are selected and described to better illustrate the principles and practical applications of the present invention, and enable those of ordinary skill in the art to understand the present invention and thus design various embodiments with various modifications suitable for specific purposes.

[0056] The first embodiment is as Figure 1 、 Figure 2As shown in the figure, the present invention provides a technical solution: a remote monitoring robot for the health monitoring of the elderly, including a mobile robot module, which is communicatively connected to a deep learning algorithm module, a voice module, an anomaly detection module, a positioning module, and an emergency safety guarantee module. Among them, the electrical signals are connected between the modules;

[0057] The deep learning algorithm module collects video streams in real time through a USB camera, uses the YOLOv5 algorithm for fall detection, intelligently analyzes the video content, and identifies the body postures of the elderly. Once a fall behavior is detected, subsequent safety guarantee measures will be immediately triggered, improving the detection accuracy and response speed of elderly fall incidents, effectively preventing serious consequences that may be caused by falls. Install a USB camera on the mobile robot module to collect video streams of the elderly's activities in real time, continuously monitor the activity area of the elderly, capture dynamic images, pre-collect a batch of action samples of falls and non-falls, and construct a fall detection model according to the ratio of training set: test set = 9:1, learn the unique body posture change patterns during falls, and deploy the fall detection model in the deep learning algorithm module to segment the collected video streams, divide them into individual video frames, each frame being a static image, and preprocess the segmented video frame images, including resizing to match the model input size, normalizing pixel values, etc., to improve the model processing efficiency and accuracy. Send the preprocessed video frames into the YOLOv5 algorithm for object detection, detect the human contour of the elderly in the video frame image, distinguish the elderly from other objects or the background. The deep learning algorithm module further analyzes the human posture in the human contour, detects the positions and relative relationships of key points (such as the head, shoulders, knees, etc.) of the elderly's body, and combines the preset fall detection model to determine whether the elderly has fallen. Run the YOLOv5 algorithm on the video frame, and according to the detection of the human body in the video frame by the deep learning algorithm, output the bounding box of the human body. In the output bounding box area of the human body, apply a pose estimation algorithm (such as OpenPose, HRNet, etc.) to identify the key points of the human body, such as the head, shoulders, knees, ankles, etc., and represent the key points in the image in the form of two-dimensional coordinates Let the coordinates of the th key point be , where represents the index of the key point (such as the head is 1, the left shoulder is 2,..., the right ankle is N), extract fall features from the detected key points. Among them, the fall features include the distance, angle, and relative position between the key points. According to the preset fall detection model, analyze whether the fall features conform to the behavior pattern during falls. Among them, the distance between the key points is calculated using the Euclidean distance formula, and its expression is: , is the key point and The distance between is the coordinate of the key point . is the coordinate of the key point . The angle between key points is calculated using the vector dot product formula, and its expression is: , is the key point , and form an angle is the coordinate of the key point . is the distance between the key point and .

[0058] are preset thresholds for different fall characteristics. When the fall detection node detects that the characteristic value is lower than the preset threshold for 5 consecutive times, combined with the preset fall detection model, it is judged that the fall characteristic conforms to the behavior pattern during a fall;

[0059] Mobile robot module, integrated with CSRT_Tracker tracking technology, autonomously adjusts the moving direction and speed according to the target position information provided by the deep learning algorithm module, maintains a safe distance from the elderly and continuously monitors, realizing real-time monitoring of the elderly when going out. Whether the elderly are at home or out for activities, the robot can closely follow to ensure no dead corners in guardianship. The mobile robot module includes a mobile robot, a development board, a USB camera and a microphone. Among them, the development board: uses the Fibocom sc171 development board as the hardware platform. The SC171 series is a multi-network mode 5G intelligent module launched by Fibocom, designed based on the Qualcomm SM6490 platform, adopting an 8-core high-performance processor (1*Kryo Gold plus 2.7GHz + 3*Kryo Gold 2.4GHz + 4*Kryo Sliver 1.95GHz), built-in VDSP, integrated with a high-performance graphics engine, and can smoothly play 4K videos. The SC171 series supports 5G NSA and SA modes, is downward compatible with 4G / 3G networks, supports 5GNR sub-6Ghz, DL4x4 MIMO, UL 2x2 MIMO, and supports 2.4G + 5G WLAN wireless communication, supports GPS (L1 + L5) / Beidou / GLONASS wireless positioning technology. SC171 is equipped with the Android 12 operating system and has various expansion interfaces such as MIPI / USB / UART / SPI / I2C. Deploy aidlux as a deep learning platform in the environment, and the native environment drives the camera and positioning module for aidlux to call and read data. CSTR_Tracker sends drive signals to the chassis stm32 to drive the motor through serial communication. Before going out, the elderly turn on the robot and start the mobile robot through voice control or mobile APP operation. Use the deep learning algorithm module to process the video frame data, output the human key points of the elderly, use the pose estimation algorithm to output the key point information, initialize the CSRT_Tracker tracking algorithm, use CSRT_Tracker to track the position of the elderly in consecutive video frames, and update its bounding box coordinates in real time. Obtain the real-time position information of the elderly from CSRT_Tracker, including the center point coordinates of the bounding box. The mobile robot uses its own sensors (such as wheel encoders, IMU, etc.) to determine its position. After identifying the target, calculate the pixel center coordinates of the target. If the coordinates appear on the right side of the USB camera, control the mobile robot to rotate to the right, and the same applies to the left side. By continuously fine-tuning the rotation direction and speed, ensure that the target is always in the center of the screen. According to the real-time position of the elderly and the current position of the mobile robot, use the path planning algorithm to calculate the best path for the mobile robot to reach the target position, and based on the path planning result, calculate the speed and direction that the mobile robot needs to adjust to maintain a safe distance from the elderly. The drive system of the mobile robot performs corresponding movement operations according to the calculation results, including adjusting speed and steering, and continuously uses the USB camera and CSRT_Tracker to track the position of the elderly during the movement of the mobile robot. The mobile robot combines the movement speed and direction of the elderly, dynamically adjusts the speed and position to ensure the smoothness and continuity during the following process and avoid losing the target;.

[0060] The voice module is used to deploy a large language model on the development board of the mobile robot module, extract and recognize the voice of the elderly from the ambient sound, including dialect recognition, enabling the elderly to send instructions to the robot in the form of voice conversations, such as "start following", "stop following", etc. In addition, a corresponding mobile phone APP is provided to facilitate the elderly to issue instructions from the mobile phone APP, simplifying the operation process and improving the usage experience of the elderly. The elderly do not need to learn complex operation interfaces or buttons and can control the robot only by voice, realizing convenient human-computer interaction. Preset command vocabulary is pre-entered by the elderly and stored to match the voice commands of the elderly. Voice commands such as "start following" or "stop following" are used. The microphone equipped with the mobile robot module continuously collects ambient sound and preprocesses the collected ambient sound, including preprocessing operations such as noise reduction, echo cancellation, and gain control, to improve the accuracy of voice recognition. Through voice activity detection technology, the effective voice part in the ambient sound is recognized, and the large language model deployed on the development board is used to recognize the recognized effective voice part, handle different dialects and accents, recognize the voice text, improve the recognition accuracy, and ensure that the elderly can communicate with the robot in a natural and habitual language. Natural language processing is performed on the voice text, including semantic understanding and command classification, to ensure correct understanding of the elderly's instructions, match the recognized voice text with the preset command vocabulary, parse the intention of the elderly, and send the parsed instructions to the development board of the mobile robot module. The development board controls the mobile robot to perform corresponding actions. After the robot executes the instructions, feedback is provided to the elderly synchronously through voice and the display screen to confirm that the instructions have been executed;

[0061] The anomaly detection module is used to monitor the behavior of the elderly by combining the results of sound recognition and visual recognition, and detect and give early warnings of the elderly falling and the occurrence of emergencies;

[0062] The positioning module enables instant communication between the robot and the emergency contact person, as well as precise positioning of the robot, ensuring the rapid transmission of help information and the accurate arrival of rescue personnel, improving the rescue efficiency, reducing the risks faced by the elderly, and ensuring that help can be obtained in a timely manner even if an accident occurs when the elderly are out;

[0063] The emergency safety guarantee module is used to combine the abnormal behavior detection results of the anomaly detection module, immediately send help and positioning information to the emergency contact person, and play a help voice around. In case of an emergency, the safety guarantee mechanism is activated to provide timely and effective assistance to the elderly. At the same time, through sound recognition technology, the intelligence level of the robot is enhanced, making it more adaptable to the actual needs of the elderly.

[0064] The second embodiment is based on the first embodiment. Please refer to Figure 3As shown, in the anomaly detection module, the process of monitoring the behavior of the elderly includes:

[0065] Collect audio and video data of the normal and abnormal behaviors (such as falling, calling for help, etc.) of the elderly under different environments, lighting conditions, and background noises, preprocess the collected audio and video data, mark the abnormal behavior data, analyze the preprocessed audio and video data, extract sound features and visual features from it, and respectively mark the abnormal sound features and abnormal action features in the sound features and visual features to obtain a sound feature set and a visual feature set. Use the abnormal sound features in the sound feature set to perform sound intensity and frequency analysis on the sound frequency data, calculate the decibel value and frequency deviation of the abnormal sound features to obtain a sound feature weighted summation function, calculate the weights of all sampling points, find the reference value in the long-term trend to obtain a reference adjustment exponential function, and perform weighted calculation on the abnormal sound to obtain an abnormal sound index to evaluate the degree of the overall abnormal sound. Use the abnormal action features in the visual feature set to preset the normal behavior reference value, calculate the deviation between a single abnormal action feature and the normal behavior reference value to obtain a behavior deviation weighted summation function, calculate the abnormal scores of all video frames at all time points, find the maximum value in the long-term trend to obtain a time series maximum weighted abnormal score function, perform weighted calculation on multiple abnormal action features to obtain an abnormal behavior index to evaluate the degree of the overall abnormal behavior. Use the associated data of the sound feature set and the visual feature set to construct a training data set, train an anomaly detection model through a neural network model to identify the features of normal and abnormal behaviors, based on the anomaly detection model and the abnormal behavior data in the sound feature set and the visual feature set, perform comprehensive analysis on the abnormal sound index and the abnormal behavior index to obtain an anomaly warning coefficient, analyze the detected abnormal behaviors, set a warning threshold, determine the nodes that trigger the warning mechanism, use the microphone equipped in the mobile robot module to continuously collect environmental sounds, and use the USB camera installed on the robot to collect video streams in real time, extract the sound features and visual features from them, input them into the trained anomaly detection model, and judge whether to trigger the warning mechanism according to the output of the anomaly detection model and the anomaly warning coefficient;

[0066] Furthermore, the abnormal sound index is obtained based on the sound feature weighted summation function and the reference adjustment exponential function, and its expression is:

[0067]

[0068] Where, is the abnormal sound index, is the sound feature weighted summation function, is the reference adjustment exponential function, is the decibel value of the th sampling point, representing the intensity of the sound, The frequency deviation of each sampling point, representing the deviation between the sound frequency and normal speech, is the total number of sampling points, is a preset reference constant used to adjust the influence of the exponential function, is the weight of the th sample, with a value ranging from 0 to 1, where 0 indicates no abnormality and 1 indicates a high degree of abnormality. As and increase (i.e., the sound intensity and frequency deviation increase), increases, indicating an increased likelihood of abnormality, represents calculating the decibel value and frequency deviation of each sound sampling point and adjusting them in combination with the logarithmic function to evaluate the abnormality degree of the sound. By summation and radical operation, the abnormality degrees of all sampling points are integrated to obtain an overall sound abnormality score, represents calculating the weights of all sampling points and finding the reference value in the long-term trend. Through the exponential function and limit function, the reference value of the sound characteristics in the long-term trend is focused on to adjust the sensitivity of the overall abnormal sound index;

[0069] The abnormal behavior index is obtained based on the behavior deviation weighted summation function and the time series maximum weighted abnormal score function, and its expression is:

[0070]

[0071] where, is the abnormal behavior index, is the behavior deviation weighted summation function, is the time series maximum weighted abnormal score function, is the measurement value of the th action feature in the th video frame, is the normal behavior reference value of the th action feature, is the total number of frames of a single action feature, is the standard deviation of the th action feature, is the normal behavior standard deviation of the th action feature, is the weight of the th video frame, is the abnormal score of the th video frame in the first evaluation, with a value ranging from 0 to 1, where 0 indicates no abnormality and 1 indicates a high degree of abnormality. As the degree of deviation of the action feature from normal behavior increases, An increase indicates an increased likelihood of an anomaly. It means calculating the deviation of each action feature in each video frame and adjusting it in combination with the standard deviation to evaluate the degree of anomaly of the action. Through summation and exponential functions, the degrees of anomaly of all action features are combined to obtain an overall abnormal behavior score. It means calculating the anomaly scores of all video frames at all time points and finding the maximum value in the long-term trend. Through limit functions and exponential functions, focusing on the maximum anomaly score in the long-term trend to evaluate the overall abnormal behavior risk.

[0072] Furthermore, the expression of the anomaly warning coefficient is:

[0073]

[0074] where The anomaly warning coefficient is the abnormal behavior index at time ; is the abnormal sound index at time t, and T is the total number of time points. and are weight coefficients used to adjust the influence of the abnormal behavior index and the abnormal sound index on the anomaly warning coefficient. is a preset reference constant used to adjust the influence of the exponential function. is the weight at the r-th time point. is the abnormal behavior index at the r-th time point. is the abnormal sound index at the r-th time point. The value of and is between 0 and 1, where 0 indicates no anomaly and 1 indicates a high degree of anomaly. As and increase, increases, indicating an increased likelihood of an anomaly, which may trigger the warning mechanism. For each time point, the abnormal behavior index and the abnormal sound index are combined. Using radical and logarithmic functions is to reduce the amplification effect of outliers while remaining sensitive to anomalies. Using and adjusts the influence of the behavior and sound indices to reflect their importance in the warning. The adjusted scores of all time points are weighted and summed to evaluate the overall anomaly warning degree.

[0075] In the emergency safety guarantee module, the process of activating the safety guarantee mechanism includes:

[0076] The deep learning algorithm module and the anomaly detection module are used to continuously analyze the sound and visual data and monitor the behavior of the elderly. When abnormal behavior is identified, an early warning signal is triggered. When the early warning signal is triggered, a help message is automatically sent to the preset emergency contact, including text messages, phone calls or application push notifications, and the built-in GPS of the mobile robot module is used to obtain the precise location of the elderly. After sending the help message, the pre-recorded help voice is played in the environment around the elderly to attract the attention of nearby people and seek help. The elderly's status is monitored and real-time updates are provided to the emergency contact. After detecting that the abnormal status is lifted, a message to lift the alarm is sent.

[0077] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without creative work should fall within the scope of protection of the present invention. The structures, devices and operating methods not specifically described and explained in the present invention are implemented according to the conventional means in the field unless otherwise specified and limited.

Claims

1. A remote monitoring robot for health monitoring of the elderly, characterized in that: It includes a mobile robot module, which is communicatively connected to a deep learning algorithm module, a voice module, an anomaly detection module, a positioning module, and an emergency safety assurance module, wherein electrical signals are connected between the modules; The deep learning algorithm module collects video streams in real time through a USB camera, uses the YOLOv5 algorithm to detect falls, intelligently analyzes video content, and recognizes the body posture of the elderly. In the deep learning algorithm module, the process of recognizing the body posture of the elderly includes: A USB camera is installed on the mobile robot module to collect real-time video streams of the elderly's activities, continuously monitor the elderly's activity area, and capture dynamic images; Collect a batch of fall and non-fall action samples in advance, build a fall detection model with a ratio of training set: test set = 9:1, learn the unique body posture change pattern when falling, and deploy the fall detection model in the deep learning algorithm module; Segment the captured video stream into separate video frames, each frame being a static image, and preprocess the segmented video frame images; The preprocessed video frames are sent to the YOLOv5 algorithm for target detection, detecting the human body contour of the elderly in the video frame image and distinguishing the elderly from other objects or backgrounds; The deep learning algorithm module further analyzes the human body posture in the human body contour, detects the position and relative relationship of the key points of the elderly's body, and combines the preset fall detection model to determine whether the elderly has fallen. The process of determining whether the elderly person has fallen down includes: Run the YOLOv5 algorithm on the video frame, detect the human body in the video frame based on the deep learning algorithm, and output the bounding box of the human body; In the output human body bounding box area, the posture estimation algorithm is applied to identify the key points of the human body and calculate them in two-dimensional coordinates. The key points are represented in the image in the form of The coordinates of the key points are ,in Indicates the index of the key point; The fall features are extracted from the detected key points, wherein the fall features include the distance, angle, and relative position between the key points. According to the preset fall detection model, whether the fall features conform to the behavior pattern when falling is analyzed. The distance between the key points is calculated using the Euclidean distance formula, and its expression is: , For key points and The distance between For key points The coordinates of For key points The angles between key points are calculated using the vector dot product formula, which is expressed as: , For key points , and The angle formed, For key points The coordinates of For key points and The distance between them; preset thresholds for different fall features. When the fall detection node detects that the feature value is lower than the preset threshold for 5 consecutive times, the fall detection model is combined to determine that the fall feature meets the behavior pattern when falling; The mobile robot module integrates CSRT_Tracker tracking technology, and autonomously adjusts the moving direction and speed according to the target position information provided by the deep learning algorithm module, so as to maintain a safe distance from the elderly and continuously monitor them, thereby realizing real-time monitoring of the elderly when they go out. The mobile robot module includes a mobile robot, a development board, a USB camera and a microphone. In the mobile robot module, the process of real-time monitoring of the elderly when they go out includes: The elderly person turns on the robot before going out and starts the mobile robot through voice control or mobile phone APP operation. The deep learning algorithm module is used to process the video frame data and output the key points of the elderly person's body. Use the posture estimation algorithm to output key point information, initialize the CSRT_Tracker tracking algorithm, use CSRT_Tracker to track the position of the elderly in consecutive video frames, and update its bounding box coordinates in real time; The real-time location information of the elderly is obtained from CSRT_Tracker, including the coordinates of the center point of the bounding box. The mobile robot uses its own sensors to determine its location. After identifying the target, calculate the pixel center coordinates of the target. If the coordinates appear on the right side of the USB camera, control the mobile robot to rotate right. The same goes for the left side. By constantly fine-tuning the rotation direction and speed, the target is always in the center of the picture. According to the real-time position of the elderly and the current position of the mobile robot, the path planning algorithm is used to calculate the best path for the mobile robot to reach the target location. Based on the path planning results, the speed and direction that the mobile robot needs to adjust are calculated to maintain a safe distance from the elderly. The driving system of the mobile robot performs corresponding movement operations according to the calculation results, including adjusting the speed and steering. During the movement of the mobile robot, the USB camera and CSRT_Tracker are continuously used to track the position of the elderly. The mobile robot dynamically adjusts the speed and position according to the movement speed and direction of the elderly. The voice module is used to deploy a large language model on the development board of the mobile robot module, extract and recognize the voice of the elderly in the ambient sound, including dialect recognition, so that the elderly can send instructions to the robot in the form of voice dialogue; The anomaly detection module is used to monitor the behavior of the elderly by combining the results of voice recognition and visual recognition, and to detect and warn the occurrence of elderly falls and emergencies; The positioning module enables instant communication between the robot and the emergency contact, as well as precise positioning of the robot; The emergency safety guarantee module is used to combine the abnormal behavior detection result of the abnormal detection module, immediately send help and positioning information to the emergency contact, and play help voice in the surrounding area.

2. A remote monitoring robot for health monitoring of the elderly according to claim 1, characterized in that: In the voice module, the process of extracting and identifying the elderly person's voice from the ambient sound includes: Pre-enter the command vocabulary preset by the elderly and store the preset command vocabulary to match the elderly's voice commands; The microphone equipped with the mobile robot module is used to continuously collect environmental sounds and pre-process the collected environmental sounds; Through voice activity detection technology, the effective speech part in the environmental sound is identified, and the large language model deployed on the development board is used to recognize the effective speech part, process different dialects and accents, and recognize the speech text; Perform natural language processing on the voice text, including semantic understanding and command classification, and match the recognized voice text with the preset command vocabulary to analyze the elderly's intentions; The parsed instructions are sent to the development board of the mobile robot module, which controls the mobile robot to perform corresponding actions. After the robot executes the instructions, it provides feedback to the elderly through voice and display screen to confirm that the instructions have been executed.

3. A remote monitoring robot for health monitoring of the elderly according to claim 2, characterized in that: In the anomaly detection module, the process of monitoring the behavior of the elderly includes: Collect audio and video data of normal and abnormal behaviors of the elderly under different environments, lighting conditions, and background noises, pre-process the collected audio and video data, and mark abnormal behavior data; Analyze the preprocessed audio and video data, extract sound features and visual features therefrom, and mark abnormal sound features and abnormal action features in the sound features and visual features respectively, to obtain a sound feature set and a visual feature set; Using the abnormal sound features in the sound feature set, the sound intensity and frequency analysis of the audio data are performed, the decibel value and frequency deviation of the abnormal sound features are calculated, and the sound feature weighted sum function is obtained. The weights of all sampling points are calculated, and the benchmark value in the long-term trend is found to obtain the benchmark adjustment index function. The abnormal sound is weighted to obtain the abnormal sound index, and the degree of overall abnormal sound is evaluated. Using the abnormal action features in the visual feature set, the normal behavior benchmark value is preset, and the deviation between a single abnormal action feature and the normal behavior benchmark value is calculated to obtain the behavior deviation weighted sum function. The abnormal scores of all video frames at all time points are calculated, and the maximum value in the long-term trend is found to obtain the time series maximum weighted abnormal score function. Multiple abnormal action features are weighted to obtain the abnormal behavior index, and the overall degree of abnormal behavior is evaluated. The training data set is constructed by using the associated data of the sound feature set and the visual feature set, and the anomaly detection model is trained through the neural network model to identify the characteristics of normal and abnormal behaviors; Based on the anomaly detection model, combined with the abnormal behavior data in the sound feature set and the visual feature set, the abnormal sound index and the abnormal behavior index are comprehensively analyzed to obtain the abnormal warning coefficient, analyze the detected abnormal behavior, set the warning threshold, and determine the node that triggers the warning mechanism; The microphone equipped with the mobile robot module is used to continuously collect environmental sounds, and the USB camera installed on the robot is used to collect video streams in real time. The sound and visual features are extracted and input into the trained anomaly detection model. According to the output of the anomaly detection model and the anomaly warning coefficient, it is determined whether the warning mechanism should be triggered.

4. A remote monitoring robot for elderly health monitoring according to claim 3, characterized in that: The abnormal sound index is obtained based on the sound feature weighted sum function and the reference adjustment index function, and its expression is: ; ; ; in, is the abnormal sound index, is the weighted sum function of sound features, The exponential function is adjusted for the benchmark, For the The decibel value of the sampling point represents the intensity of the sound. For the The frequency deviation of the sampling points indicates the deviation of the sound frequency from normal speech. is the total number of sampling points, is a preset reference constant used to adjust the impact of the exponential function. For the The weight of the samples, The value is between 0 and 1, where 0 indicates no abnormality and 1 indicates a high degree of abnormality.

5. A remote monitoring robot for elderly health monitoring according to claim 4, characterized in that: The abnormal behavior index is obtained based on the weighted sum function of the behavior deviation and the maximum weighted abnormal scoring function of the time series, and its expression is: ; ; ;in, is the abnormal behavior index, is the weighted sum function of behavioral deviation, is the maximum weighted anomaly scoring function for time series, For the The action feature is The measurements in the video frames, For the The normal behavior baseline value of the action feature, is the total number of frames of a single action feature, For the The standard deviation of the action features, For the The standard deviation of normal behavior of action features, is the number of action features, For the The weight of the video frame, For the The video frame is Abnormal scores in the evaluation, The value is between 0 and 1, where 0 indicates no abnormality and 1 indicates a high abnormality.

6. A remote monitoring robot for elderly health monitoring according to claim 5, characterized in that: The expression of the abnormal warning coefficient is: ;in, is the abnormal warning coefficient, For in time The abnormal behavior index, For in time The abnormal sound index, is the total time points, and is the weight coefficient, which is used to adjust the influence of abnormal behavior index and abnormal sound index in the abnormal warning coefficient. is a preset reference constant used to adjust the impact of the exponential function. For the The weight of each time point, For the Abnormal behavior index at each time point, For the Abnormal sound index at a time point.

7. A remote monitoring robot for elderly health monitoring according to claim 6, characterized in that: In the emergency safety assurance module, the process of starting the safety assurance mechanism includes: Use deep learning algorithm modules and anomaly detection modules to continuously analyze sound and visual data, monitor the behavior of the elderly, and trigger warning signals when abnormal behavior is identified; When the warning signal is triggered, it automatically sends a help message to the preset emergency contact, including text messages, phone calls or app push notifications, and uses the built-in GPS of the mobile robot module to obtain the precise location of the elderly; After sending the help message, a pre-recorded help voice is played in the environment around the elderly to attract the attention of nearby people and seek help, monitor the elderly's condition, and provide real-time updates to emergency contacts. After detecting that the abnormal condition is lifted, a message to lift the alarm is sent.

Citation Information

Patent Citations

  • Method for analyzing and monitoring health condition of patient

    CN113688736A

  • Telehealthcare system and service using an intelligentmobile robot

    KR1020060084916A