A human presence detection system and method based on multi-sensor collaboration of robots
By using a multi-sensor collaborative robot system that combines vision, lidar, and millimeter-wave radar sensors, comprehensive monitoring of the human body's state is achieved. This solves the problems of false alarms and missed alarms in existing technologies, improves the accuracy and efficiency of monitoring, and is suitable for rapid response to emergencies and long-term health management.
Patent Information
- Application Number
- CN202411618003.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-13
AI Technical Summary
In home health monitoring, existing technologies such as millimeter-wave radar and lidar have false alarms and missed alarms when used to detect the presence of the human body, making it difficult to achieve comprehensive and accurate monitoring.
A multi-sensor collaborative robot system is adopted, which combines visual sensors, lidar sensors and millimeter-wave radar sensors to collect image data, point cloud data and human physiological data in real time. Decisions are made through fusion analysis to achieve accurate monitoring of the human body's state.
It enables comprehensive monitoring of the human body's condition, improving the accuracy and efficiency of monitoring, allowing for rapid response in emergencies and long-term health management, and providing personalized health management and early warning services.
Smart Images

Figure CN119716836B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human body monitoring technology, specifically to a human presence detection system and method based on multi-sensor collaboration of robots. Background Technology
[0002] In current home health monitoring solutions, millimeter-wave radar technology is used to detect the presence of the human body. For example, existing AI-powered personal health management robots use millimeter-wave radar to detect human presence, including identifying standing, walking, sleeping, falls, heartbeat, and breathing; and use lidar waves for obstacle avoidance detection, including identifying pedestrians, furniture, and obstacles. However, existing millimeter-wave radar requires a height difference to achieve human presence detection. The detection method involves placing the millimeter-wave device at a high position, and judging whether the user has fallen by assessing the height difference after a fall. Its applicability is relatively limited. Lidar waves are mostly used in vehicles and have not been widely adopted in robots.
[0003] However, the limitations of existing technologies have led to false alarms and missed alarms in the monitoring of human body conditions, and it is difficult to achieve accurate identification and comprehensive monitoring of each individual. Summary of the Invention
[0004] Based on the above-mentioned technical problems, this application aims to provide a human presence detection system and method based on robot multi-sensor collaboration, so as to at least solve one of the above-mentioned technical problems.
[0005] The first aspect of this application provides a human presence detection system based on robot multi-sensor collaboration. The human presence detection system includes a robot, a data acquisition module, and a robot operating system. The robot operating system is built into the robot. The data acquisition module includes a vision sensor, a lidar sensor, and a millimeter-wave radar sensor. The vision sensor is disposed in the head of the robot, the millimeter-wave radar sensor is disposed in the body of the robot, and the lidar sensor is disposed in the bottom of the robot. The vision sensor, the lidar sensor, and the millimeter-wave radar sensor are all connected to the robot operating system.
[0006] The vision sensor is used to collect image data of the target area in real time while the robot is walking and send the image data to the robot operating system;
[0007] The lidar sensor is used to collect point cloud data of the target area in real time while the robot is walking and send the point cloud data to the robot operating system;
[0008] The millimeter-wave radar sensor is used to collect human physiological data of the target area in real time while the robot is walking and send the human physiological data to the robot operating system.
[0009] The robot operating system is used to fuse and analyze the image data, the point cloud data, and the human physiological data, and make decisions based on the analysis results to control the robot's response.
[0010] In some embodiments of this application, the fusion analysis of the image data, the point cloud data, and the human physiological data includes:
[0011] The image data, the point cloud data, and the human physiological data are filtered respectively to obtain filtered image data, filtered point cloud data, and filtered human physiological data.
[0012] Based on the filtered point cloud data, a 3D model is constructed to obtain an environment model and a user model.
[0013] Based on the aforementioned environment model and user model, and the filtered image data, a fusion analysis is performed to obtain the first analysis result for the target user.
[0014] Based on the primary analysis results and the filtered human physiological data, a fusion analysis is performed to obtain the second analysis results for the target user.
[0015] In some embodiments of this application, the step of constructing a 3D model based on the filtered point cloud data to obtain an environment model and a user model includes:
[0016] A three-dimensional model is constructed from the filtered point cloud data using triangulation or voxelization methods to determine the model representing the terrain of the target area, and the model representing the terrain of the target area is used as the environment model.
[0017] A model representing the active state of the human body is determined, and this model is used as the user model.
[0018] In some embodiments of this application, the step of performing fusion analysis based on the environment model and user model, and the filtered image data, to obtain a first analysis result for the target user, includes:
[0019] Based on the aforementioned environment model and user model, potential problem users are identified;
[0020] Feature extraction is performed on the filtered image data, and the target user is determined based on the extracted features and the potential problem users.
[0021] The target user's physical activity state is analyzed to obtain the first analysis result of the target user, wherein the target user's physical activity state includes the target user's location, the target user's posture, and the target user's interaction state with the environment.
[0022] In some embodiments of this application, the step of performing fusion analysis based on the environment model and user model, and the filtered image data, to obtain a first analysis result for the target user, includes:
[0023] The processed image data is used to perform pixel mapping on the environment model and user model in order to identify the target user;
[0024] The target user's physical activity state is analyzed to obtain the first analysis result of the target user, wherein the target user's physical activity state includes the target user's location, the target user's posture, and the target user's interaction state with the environment.
[0025] In some embodiments of this application, the step of performing a fusion analysis based on the primary analysis results and the filtered human physiological data to obtain a second analysis result for the target user includes:
[0026] Features are extracted from the filtered human physiological data, and the health information of the target user is determined based on the extracted features.
[0027] A second analysis result of the target user is obtained by performing a deep fusion analysis on the first analysis result of the target user, the target user's health information, and the target user's historical data.
[0028] In some embodiments of this application, the step of performing deep fusion analysis on the first analysis result of the target user, the health information of the target user, and the historical data of the target user to obtain the second analysis result of the target user includes:
[0029] A long short-term memory network is constructed and trained using the target user's historical data, wherein the target user's historical data includes the target user's historical behavior, activity records, historical medical data, and historical physiological data;
[0030] The health information of the target user is predicted by a trained long short-term memory network, and the prediction result is obtained.
[0031] Using the Framingham model, a multiple regression analysis is performed on the prediction results and the first analysis results of the target user to obtain the second analysis results of the target user.
[0032] A second aspect of this application provides a human presence detection method based on multi-sensor collaboration of a robot, the method comprising:
[0033] The robot collects image data, point cloud data, and human physiological data of the target area in real time while it is walking.
[0034] The image data, point cloud data, and human physiological data are fused and analyzed, and decisions are made based on the analysis results to control the robot's response.
[0035] A third aspect of this application provides an electronic device, including a memory and a processor. The memory stores computer-readable instructions, which, when executed by the processor, cause the processor to perform the following steps:
[0036] The robot collects image data, point cloud data, and human physiological data of the target area in real time while it is walking.
[0037] The image data, point cloud data, and human physiological data are fused and analyzed, and decisions are made based on the analysis results to control the robot's response.
[0038] The fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps: real-time acquisition of image data, point cloud data, and human physiological data of a target area while the robot is walking.
[0039] The image data, point cloud data, and human physiological data are fused and analyzed, and decisions are made based on the analysis results to control the robot's response.
[0040] The technical solutions provided in this application embodiment have at least the following technical effects or advantages:
[0041] The human presence detection system based on robot multi-sensor collaboration described in the various embodiments of this application realizes the innovative integration of visual sensors, lidar sensors and millimeter-wave radar sensors into the application of robots. It achieves low-position human presence detection, facial recognition and deep monitoring of biometrics (such as heartbeat, blood pressure and sleep) from surface to point. This integrated application is the first of its kind at home and abroad, and breaks through the limitations of single sensors in human monitoring.
[0042] Utilizing lidar and millimeter-wave radar sensors, this system can operate stably under varying lighting and environmental conditions, enhancing its adaptability and stability in complex environments. Through real-time data acquisition and analysis, the system can monitor the physiological and behavioral status of target users in real time, making it suitable for rapid response to emergencies and long-term health management. Through multi-sensor data fusion analysis, the system can more accurately and comprehensively monitor and understand human activity and physiological conditions, improving the accuracy and efficiency of monitoring.
[0043] Furthermore, by combining users' historical data for in-depth analysis and health prediction, the system can provide users with personalized health management and early warning services. In particular, when a user accidentally falls and loses consciousness, the system can act as the first responder to provide medical emergency assistance, significantly improving the safety and timeliness of emergency care for home users and effectively saving lives.
[0044] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0045] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0046] Figure 1 This is a schematic diagram of the structure of a human presence detection system based on robot multi-sensor collaboration in an exemplary embodiment of this application;
[0047] Figure 2 This is a schematic diagram of another human presence detection system based on robot multi-sensor collaboration in an exemplary embodiment of this application;
[0048] Figure 3 This is a schematic diagram comparing the properties of a camera and a lidar wave sensor in an exemplary embodiment of this application;
[0049] Figure 4 This is a schematic diagram of transmitting a linear frequency modulated pulse to a patient according to an exemplary embodiment of this application;
[0050] Figure 5 This is a schematic diagram illustrating the steps of a human presence detection method based on multi-sensor collaboration of a robot in an exemplary embodiment of this application;
[0051] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this application.
[0052] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Detailed Implementation
[0053] Current home health monitoring solutions heavily rely on millimeter-wave radar technology for detecting human condition (e.g., determining if a fall has occurred). A primary application of this technology is placing the detection device at a high position, such as on a room wall or near a light source, to identify different postures by measuring changes in body height. On the other hand, while lidar technology is widely used in automotive systems, its application in health monitoring robots is relatively limited. These current technological limitations lead to potential false alarms and missed detections when monitoring human condition, and make accurate and comprehensive monitoring of each individual difficult.
[0054] Therefore, this application provides a human presence detection system and method based on robot multi-sensor collaboration. By fusing and analyzing image data, point cloud data, and physiological data collected by multiple sensors, it achieves comprehensive monitoring of the human presence and health status of each target user in the target area.
[0055] The present application will now be described in further detail with reference to the accompanying drawings and several embodiments. It should be understood that the embodiments depicted herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0056] Example 1
[0057] This embodiment provides a human presence detection system based on multi-sensor collaboration of robots. Figure 1 This is a schematic diagram of a human presence detection system based on multi-sensor collaboration of robots, such as... Figure 1 As shown, the human detection system includes a robot, a data acquisition module, and a robot operating system (ROS). The robot operating system is built into the robot, and the data acquisition module includes a vision sensor, a lidar sensor, and a millimeter-wave radar sensor, such as... Figure 2 As shown, a vision sensor is located in the robot's head, a millimeter-wave radar sensor is located in the robot's body, and a lidar sensor is located in the robot's bottom. All three sensors are connected to the robot's operating system.
[0058] Specifically, the visual sensor is a camera, which is equipped with a video sensor, specifically a human scene sensor. The millimeter-wave radar sensor is a millimeter-wave radar chip, which is uniformly distributed circumferentially on the robot's body. The lidar sensor is a lidar chip used to collect point cloud data of a preset environment. Preferably, one camera, three millimeter-wave radar chips, and one lidar chip are installed on the robot. The human presence detection system also includes an alarm module, which issues an alarm message when the human presence detection system detects abnormal data.
[0059] The camera captures image and video data for facial, posture, and other visual feature recognition. Located at the front of the robot, it acquires real-time photos of a preset environment (such as a family living room) to obtain image data within the detection environment. The camera is equipped with a video sensor, specifically a human scene sensor, which uses FP2 for non-visual and non-intrusive security detection. When individual identification is required, the camera is invoked for facial recognition. The system conforms to standard GB 4943.1-2022 and uses Wi-Fi IEEE 802.11 and Bluetooth 4.2 for wireless connectivity.
[0060] The millimeter radar wave chip is used to monitor human activities such as heartbeat and sleep in real time. Each millimeter radar wave chip has a detection range of 120 degrees, a radial distance of about 8 meters, a data refresh rate of 0.5 seconds, and a grid of 320.
[0061] The aforementioned lidar wave chip is used to collect discrete point cloud data, accurately acquire 3D structural data within the human detection environment, thereby achieving environmental modeling and outputting high-precision distance information of personnel within the human detection environment to identify their activity paths. Combined with data acquired by a camera, the lidar wave chip enables perception of complex and changing environments, enhancing the intelligence and data accuracy of the system's human detection, making motion planning and decision-making safer and more reliable.
[0062] The human presence detection system based on robot multi-sensor collaboration can set the data acquisition frequency according to user needs, with a default acquisition frequency of once per hour. Furthermore, the system also sets detection parameter thresholds; when the monitored data reaches the threshold, a new round of detection is automatically initiated. These thresholds include: detecting noise levels exceeding 50 decibels from actions within the monitoring environment; stopping for 30 seconds upon acquiring user heartbeat information; and detecting user behavior such as falling and remaining motionless for an extended period. After repeated detections, if the data remains abnormal, the system automatically triggers an alarm, sending a distress signal to the associated contact person, community hospital, or clinic. It should be noted that the lidar sensor is used for distance measurement and map building. By emitting a laser beam and receiving the reflected beam, it measures distance and can generate high-precision 3D environment and object models, suitable for accurately monitoring objects and human postures in space. In this application, the lidar sensor is used to collect point cloud data of the target area in real time while the robot is moving and sends the point cloud data to the robot operating system.
[0063] Millimeter-wave radar sensors utilize radar waves in the millimeter-wave frequency band to measure distance, velocity, and angle. Millimeter-wave radar is particularly suitable for monitoring human physiological signals (such as heartbeat and respiration) and movement status in harsh visual environments (such as smoke, nighttime, or bright light) because it can penetrate some visual obstructions. In this application, at least three millimeter-wave radar sensors are used, arranged around the robot's body to collect real-time human physiological data of the target area while the robot is walking, and then transmit this data to the robot's operating system. This allows for the collection of more comprehensive and detailed human physiological data.
[0064] The target area mentioned above can be different rooms within a home, such as the living room. When navigating indoors, the Robot Operating System (ROS) makes decisions to control the robot to plan the optimal path based on the environment. This includes moving in a straight line, bypassing stationary or dynamic obstacles, carefully navigating narrow passages, and performing turning maneuvers when necessary, whether it's an acute-angle turn, a curved path, or other complex walking patterns. This enables the robot to recognize and adapt to various geometries and obstacles in the environment, achieving efficient and safe navigation.
[0065] Visual sensors are positioned in the robot's head, such as cameras connected to the robot's operating system via USB or similar high-speed data interfaces, enabling efficient image data transmission and processing. LiDAR sensors are located on the robot's bottom, primarily using a network interface and communicating via the UDP protocol to handle the large data volumes generated by the LiDAR. While network configuration for the LiDAR does not require dedicated drivers, users need to configure the network information according to the product manual to acquire data correctly. Millimeter-wave radar sensors are positioned in the robot's body. Connection to the robot's operating system can also be via a network interface or other interfaces suitable for high-speed data transmission to support real-time acquisition of physiological signals and motion data.
[0066] The vision sensors, lidar sensors, and millimeter-wave radar sensors in the detection system are all connected to the robot operating system through corresponding interfaces. Furthermore, such as... Figure 2 As shown, the entire detection system also works in conjunction with a Zigbee gateway via a Wi-Fi local area network to achieve wireless communication and data exchange. The robot's operating system supports the GB4943.1-2022 standard and utilizes Wi-Fi's IEEE 802.11 and Bluetooth 4.2 for wireless connectivity. For applications requiring non-visual and non-intrusive security detection, the millimeter-wave radar sensor employs FP2 technology. FP2 technology incorporates data encryption, non-intrusive data transmission, and specific communication protocols to enhance security and privacy, especially during sensitive operations such as facial recognition.
[0067] It is understood that this embodiment can plan and utilize cameras, lidar waves, and millimeter-scale radar waves, effectively integrating the functions of each component to achieve overall decision-making and planning; it establishes a terrain surface model through lidar waves and monitors personnel activities in the environment in real time; it uses cameras to identify personnel, while simultaneously deploying millimeter-scale radar waves to comprehensively monitor personnel activities in the environment, effectively enhancing the adaptability and stability of the health robot system in various complex environments and reducing the possibility of data distortion and loss during monitoring; in addition, this system can more accurately and comprehensively achieve the goal of real-time monitoring of the physiological and behavioral status of target users, making it suitable for rapid response to emergency events and long-term health management.
[0068] In summary, combining these three sensors in a health monitoring robot can significantly improve the accuracy and comprehensiveness of human body monitoring. The visual sensor acquires detailed visual information, the lidar sensor provides precise spatial positioning and environmental modeling, and the millimeter-wave radar sensor enhances monitoring capabilities in visually limited environments. Together, they achieve accurate individual identification and comprehensive health monitoring, effectively reducing false alarms and missed alarms. The detection system also innovatively integrates visual, lidar, and millimeter-wave radar sensors into the robot, enabling low-position human presence detection, facial recognition, and in-depth monitoring of biometrics (such as heart rate, blood pressure, and sleep) from surface to point. This integrated application is the first of its kind both domestically and internationally, breaking through the limitations of single sensors in human body monitoring.
[0069] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application.
[0070] Example 2
[0071] This embodiment provides a human presence detection system based on multi-sensor collaboration of robots, referencing... Figure 1 and Figure 2 Based on the specific structure of the human presence detection system based on multi-sensor collaboration of robots shown in Example 1, this example provides the working method of the detection system. The robot mentioned above refers to a health robot.
[0072] It's understandable that the lidar waves on the health robot are used to perceive and model the user's indoor environment. Based on the input environmental information and monitoring mode, the lidar sets the scanning frequency, laser beam angle, and scanning speed. By frequently emitting laser beams and receiving reflected signals, it acquires sufficiently accurate point cloud data in a short time. During point cloud data acquisition, it's crucial to ensure the stability of the health robot's position and angle to avoid data acquisition errors or repeated scanning.
[0073] A continuous terrain surface model is obtained by processing discrete point cloud data collected by lidar sensors. First, the raw data needs to be cleaned to remove some noise points that do not meet the requirements. Second, Gaussian filtering and median filtering are applied to the point cloud data to eliminate irregular fluctuations and unevenness in the data. Then, data fitting and surface fitting are performed to obtain a smoother and more continuous terrain surface.
[0074] The terrain surface model constructed by the system depends on the division of the collected point cloud data, which can be set as a triangular mesh or a voxel mesh. Specifically, the point cloud data is divided into different regions through clustering methods, and then each region is further processed and modeled. The triangular mesh method connects the point cloud data into a triangular mesh to form a continuous terrain surface model. The voxel mesh method divides the point cloud data into regular voxel units and interpolates each voxel to obtain a smoother and more continuous terrain surface model.
[0075] After completing the terrain surface modeling, the accuracy of the modeling results needs to be evaluated and checked. This can be done by comparing the modeling results with actual terrain data to calculate the error between the modeling results and the actual data; at the same time, visual inspection can also be performed to analyze the modeling results in order to determine the accuracy and reliability of the modeling results.
[0076] After confirming the terrain surface modeling, the next step is user modeling. User modeling utilizes camera vision for judgment. Based on the first model (ground model) data, which includes terrain information and user information, a pre-trained model of LiDAR waves is applied to the robot's vision to determine whether the terrain model and user model created by the LiDAR waves contain errors due to curved ground, and to identify which user is identified by facial recognition for each user. This yields the second model (user model) data, which includes location information and user information.
[0077] Figure 3 This is a schematic diagram comparing the properties of a camera and a LiDAR wave sensor, where the brown line represents the camera and the red line represents the LiDAR wave sensor. Figure 3 It can be seen that cameras can output high-density image information, but their distance estimation capabilities and robustness are poor; lidar sensors can directly output high-precision distance information and clear 3D structural data, but their point clouds are sparse. The radar chart also shows that the two sensors are highly complementary. By combining fine-grained camera images and long-range detection capabilities, high-resolution depth maps and 3D representations can be generated, providing far more information than any single sensor can offer. Compared to the rich readings of a single sensor, the rational fusion of vision-lidar data will greatly enhance the perception of complex and changing environments, improve the intelligence level of the perception system, reduce uncertainty, and make motion planning and decision-making safer and more reliable.
[0078] The principle of lidar wave sensor and camera fusion is as follows: First, the relative positional relationship parameters of the lidar and camera are determined by calibration. Then, the three-dimensional space of the lidar point cloud is mapped to the two-dimensional space of the camera image, so that the detected object in the image obtains its depth and position information in three-dimensional space, as well as the composition information of three-dimensional space elements. Next, the ground part in the point cloud is extracted and separated from the lidar point cloud. Then, the three-dimensional information of the detected target is used to process the detected object in the lidar point cloud, filter out point cloud noise, further optimize the point cloud matching relationship, and densify the point cloud of the detected object. After that, the point cloud parts of the detected object are separated one by one, and the three-dimensional bounding box of its point cloud is calculated. Finally, the vertices of the three-dimensional bounding box of the detected object are projected onto the camera image space to obtain its pseudo three-dimensional bounding box information.
[0079] The specific process of data fusion processing based on LiDAR wave sensors and cameras includes the following steps: First, before using the LiDAR wave sensors and cameras, the relative positional relationship between them needs to be pre-calculated; this process of determining these positional relationships is usually called calibration. After calibration, the data collected by the LiDAR wave sensors and cameras are fused together. Next, data fusion is performed. The data processing here refers to the image and LiDAR point cloud data collected using the LiDAR wave sensors and cameras. Based on the calibrated parameters, the correspondence between spatial points in the LiDAR point cloud and pixels in the image is calculated, thus fusing the data collected by different sensors. Fusion processing can fully utilize the features of data collected by different sensors. For example, color curvature information collected by the camera, and spatial coordinates and reflectivity information collected by the LiDAR wave sensors. Then, clustering processing is performed on the initially fused data because the initially fused data may contain some errors, usually caused by matching errors due to differences in sensor positions. Therefore, clustering based on feature information can correct some of the erroneous fusion processing.
[0080] Next, data optimization is performed. When generating the 3D spatial detection box, the size and position of the bounding box are optimized based on object features, spatial relationships between objects, and environmental information to obtain a more accurate 3D spatial detection box. This application uses a LiDAR sensor to perceive the detected object, enabling the acquisition of centimeter-level depth information and effectively reducing depth estimation errors. Simultaneously, a complete 3D structure is established: due to the positional difference between the LiDAR and the camera, and the fact that the LiDAR point cloud possesses 3D information, the 3D structural information of the detected object can be constructed.
[0081] Then, the second model data output from the previous step is combined with millimeter-wave radar waves from the robot to obtain the third model data (i.e., the physiological model of the target user). The third model data includes monitoring information and user information. This step utilizes the millimeter-wave radar's capabilities, achieving centimeter-level positioning indoors, and employing MIMO virtual aperture super-resolution point cloud imaging to intelligently identify the location, angle, distance, and amplitude of personnel. It accurately extracts and intelligently identifies micro-motion information such as breathing, heartbeat, sleep, and complex physiological characteristic signals. This enables precise monitoring of each user's presence indoors. The fourth model data is then output, including user ID information (UserNumber) and physiological status information (PhysioligicalStatus).
[0082] The advantage of short wavelengths in millimeter radar is its high accuracy. Millimeter-wave radar with frequencies of 60 or 77 GHz (corresponding to wavelengths within the 4 mm range) can detect movements as short as less than 1 mm. Figure 3 This is a schematic diagram illustrating the use of millimeter-wave radar to project linear frequency modulated pulses onto a patient's chest region. Due to chest movement, the reflected signal is phase-modulated. The modulation encompasses all components of motion, including those caused by heartbeat and respiration, such as... Figure 4 As shown, the modulation and analysis of the reflected signal yielded a respiratory rate of 26 and a heart rate of 80. Millimeter radar waves transmit multiple linearly frequency-modulated pulses at predetermined time intervals. Each pulse undergoes a range Fast Fourier Transform (FFT), selecting a range corresponding to the person's chest position. Each linearly frequency-modulated pulse records the signal phase within that selected range. The phase change is calculated, and thus the velocity is derived. The obtained velocity still includes all motion components. A Doppler FFT is performed to analyze the spectrum of the obtained velocity, thereby resolving the various component data.
[0083] Finally, a multivariate regression analysis is performed on the target user's physiological data based on the Framingham model. First, the Framingham model requires various health indicators as input, which can be obtained from the fourth model data. These indicators include risk factors required by the Framingham model, such as the target user's age, gender, blood pressure, cholesterol levels, etc. Then, the regression equation of the Framingham model is applied to the target user's data, i.e., the Framingham model is used to calculate the target user's health risk score, which reflects the probability of the target user developing cardiovascular disease in the future. Based on the output of the Framingham model, a risk threshold is set. When the target user's risk score exceeds this threshold, the system automatically issues a health warning. This warning can help the target user take timely preventative measures, such as changing lifestyle or seeking medical help. Optionally, the risk score and warning information from the Framingham model can be integrated into the target user's health report and interacted with through the system, such as sending the report to the target user's mobile phone. It can be understood that after modeling a health model for each individual user, not only is a daily health data model generated, but this data is also combined with other health data collected by the robot to complete a more comprehensive and user-tailored health model. Meanwhile, the Framingham model can analyze each user's health trend and risk factors related to the occurrence of related diseases, and issue early warnings when a danger threshold is reached, thus proactively establishing a disease risk prediction model.
[0084] As can be seen, the human presence detection system based on robot multi-sensor collaboration has achieved the innovative integration of visual sensors, lidar sensors and millimeter-wave radar sensors into the application of robots. It has realized the detection of human presence at low position, facial recognition and deep monitoring of biometrics (such as heartbeat, blood pressure and sleep) from surface to point. This integrated application is the first of its kind at home and abroad, and has broken through the limitations of single sensors in human monitoring.
[0085] In one specific implementation, the fusion analysis of image data, point cloud data, and human physiological data first involves filtering each data separately to obtain filtered image data, filtered point cloud data, and filtered human physiological data. Next, a three-dimensional model is constructed based on the filtered point cloud data to obtain an environment model and a user model. Then, based on the environment model, user model, and filtered image data, a fusion analysis is performed to obtain the first analysis result for the target user. Finally, based on the initial analysis result and the filtered human physiological data, a fusion analysis is performed to obtain the second analysis result for the target user.
[0086] Filtering aims to remove noise and outliers from data, improving data quality. Different filtering techniques can be used for point cloud data, image data, and human physiological data. For example, statistical analysis filtering (such as moving average or Gaussian filtering) can be used for point cloud data, edge-preserving filtering algorithms (such as bilateral filtering or nonlocal mean filtering) can be used for image data, and bandpass filters can be applied to human physiological data to remove non-target frequency components.
[0087] Specifically, when constructing a 3D model based on filtered point cloud data to obtain an environment model and a user model, the filtered point cloud data is first used to construct a 3D model through triangulation or voxelization to determine the model representing the terrain of the target area, and the model representing the terrain of the target area is used as the environment model; then the model representing the human body's activity state is determined, and the model representing the human body's activity state is used as the user model.
[0088] In one possible implementation, a fusion analysis is performed based on an environment model and a user model, as well as filtered image data, to obtain the first analysis result for the target user, including the following steps.
[0089] S100. Identify potential problem users based on the environment model and user model.
[0090] Specifically, LiDAR sensors efficiently collect accurate point cloud data by determining the scanning frequency, laser beam angle, and scanning speed. The Robot Operating System (ROS) then performs clutter removal and filtering on this point cloud data, and constructs a continuous terrain surface model using triangulation or voxelization methods. This model includes both environmental and user models. This process involves not only accurate environmental perception but also the initial identification of potential problem users, such as identifying those who may be located on unstable ground or near potential obstacles by analyzing the terrain model.
[0091] For example, the robot uses a lidar sensor on its bottom to collect point cloud data of the room. This data includes the position and shape of the floor, furniture, walls, etc. The Robotic Operating System (ROS) uses the point cloud data to construct a 3D model of the environment through triangulation or voxelization. This environment model represents the static world around the robot, including curved floors, furniture layout, and so on. The ROS uses the environment model to identify the precise location and shape of curved floors, and then calculates an optimal path to control the robot to walk at the appropriate angle and speed to avoid slipping or deviating from the path.
[0092] For example, if the robot encounters an obstacle during autonomous navigation (e.g., performing autonomous navigation once every hour), and the user model determines that the obstacle is a user, ROS controls the robot to adjust its path to bypass the user. Simultaneously, it utilizes human physiological data for fusion analysis; for instance, if it happens to encounter a family member who has fallen, that family member can be identified as a potential problem user. If it encounters dynamic obstacles such as pets, ROS controls the robot to adjust its path in real time, ensuring both efficient navigation and collision avoidance.
[0093] S200. Extract features from the filtered image data, and determine the target user based on the extracted features and potential problem users.
[0094] The potential problem users identified in step S100 may not be clearly identified. That is, ROS analyzes the point cloud data collected by the LiDAR sensor to determine the approximate geographical range (environment model) and identify users who may pose health risks (user model). For example, the area where a fall occurred needs to be further confirmed by combining filtered image data.
[0095] The purpose of step S200 is to identify and analyze image content, such as object detection, scene understanding, and face recognition. Optionally, convolutional neural networks (CNNs) can be used to extract image features and perform classification or regression analysis. By combining the potential problem users identified by the environment model and user model, face recognition can be performed to determine the target user, i.e., to clearly see which specific family member.
[0096] S300. Analyze the human body activity state of the target user to obtain the first analysis result of the target user. The human body activity state of the target user includes the target user's location, the target user's posture, and the target user's interaction state with the environment.
[0097] In one specific implementation, suppose ROS identifies a family member approaching a slippery area in the living room, such as a freshly cleaned and slippery floor, or someone has already fallen. It then performs a detailed analysis of the target user's physical activity, including their specific location, posture, and interaction with their surroundings. For example, the first analysis might reveal that the target user is located on the floor near the center of the living room, specifically 2 meters from the north wall and 3 meters from the east wall; their posture is lying flat on the floor, slightly tilted to the left, with arms outstretched; and they are near a slippery area with no obstacles nearby. Based on this analysis, ROS can further process the physiological data collected by the millimeter-wave radar sensor and determine whether action is necessary based on the fusion analysis results.
[0098] As a variable implementation method, based on the environment model and user model, and the filtered image data, a fusion analysis is performed to obtain the first analysis result of the target user, and the following steps can also be performed.
[0099] S101. Utilize the processed image data to perform pixel mapping on the environment model and user model in order to identify the target user.
[0100] The pixel mapping described above establishes a pixel-level correspondence between the feature information (such as edges and shapes) of the environment model and user model and the processed image data. The Robot Operating System (ROS) maps the processed image data onto the environment model, enhancing the realism of the model and making it closer to the appearance of the real world. Mapping image data onto the user model helps identify human movement patterns within the user model, thereby accurately locating and identifying the target user. For example, by performing pixel mapping on the user model, the robot can more accurately identify the characteristics of the target user, such as clothing color, possible postures, and relationship with the surrounding environment. This method significantly improves the accuracy of target recognition and the depth of environmental understanding.
[0101] S201. Analyze the human body activity state of the target user to obtain the first analysis result of the target user. The human body activity state of the target user includes the target user's location, the target user's posture, and the target user's interaction state with the environment.
[0102] If the robot encounters a user falling during autonomous navigation, it can determine the target user and their interaction with the surrounding environment through image data mapping, i.e., a visually enhanced environment model and user model. For example, the first analysis result is: the target user is located on the floor near the center of the living room, specifically 1 meter from the north wall and 1 meter from the east wall; the target user is lying flat on the ground, with their body slightly tilted to the right and their arms spread apart; the target user is near a slippery floor area, next to a standing clothes rack.
[0103] The above implementation enhances the accuracy and realism of the environment and user models through pixel mapping, making the identification of target users more precise. Simultaneously, by analyzing the target user's physical activity state, the robot can gain a deeper understanding of the user's specific situation and needs, providing more effective support for subsequent fusion analysis.
[0104] In one specific implementation, a fusion analysis is performed based on the primary analysis results and the filtered human physiological data to obtain the second analysis result of the target user. This includes: extracting features from the filtered human physiological data and determining the target user's health information based on the extracted features; and performing a deep fusion analysis on the target user's primary analysis result, the target user's health information, and the target user's historical data to obtain the second analysis result of the target user.
[0105] Specifically, a deep fusion analysis is performed on the target user's first analysis results, the target user's health information, and the target user's historical data to obtain the target user's second analysis results, including:
[0106] A long short-term memory network is constructed and trained using the target user's historical data, which includes the target user's historical behavior, activity records, historical medical data, and historical physiological data. The trained long short-term memory network is used to predict the target user's health information to obtain the prediction results. Using the Framingham model, a multiple regression analysis is performed on the prediction results and the target user's first analysis results to obtain the target user's second analysis results.
[0107] Implementation requires building and training a Long Short-Term Memory (LSTM) network. First, extensive historical data on the target user is collected, including but not limited to historical behavioral patterns, activity records, medical records, and physiological data (such as heart rate and blood pressure). This data provides a rich foundation for building a personalized health prediction model. Then, the collected historical data is used to construct the LSTM network. LSTM is a special type of recurrent neural network (RNN), particularly suitable for handling and predicting long-term dependencies in time series data. By training the LSTM, the model learns to identify health trends and potential risks from the user's historical behavioral and physiological data. The trained LSTM network is then used to predict the target user's future health status based on current health information. Predictions include potential health risks, the development trends of potential diseases, and changes in the user's health status, among other things.
[0108] Furthermore, using the Framingham model, a multiple regression analysis was performed on the prediction results of the trained LSTM network and the first analysis results of the target user to obtain the second analysis results of the target user, achieving in-depth monitoring of the human body's state from a broad perspective to a specific level. The Framingham model, also known as the cardiovascular disease risk model, is a tool used to assess an individual's risk of developing cardiovascular disease over a future period. The Framingham model considers multiple risk factors based on historical medical and physiological data, such as the target user's age, gender, blood pressure, cholesterol levels, etc., and finally combines the user's first analysis results with multiple regression analysis.
[0109] By combining LSTM-predicted future health information, the user's first analysis results, and Framingham model scores, a multiple regression analysis is performed. This considers not only the user's physiological parameters and environmental interactions but also their long-term health trends and risk factors, resulting in a comprehensive future health prediction and risk assessment. The second analysis, derived from the fusion analysis, provides the target user with a comprehensive health risk assessment, including potential health challenges, risk factors requiring attention, and personalized health recommendations or preventative measures based on the prediction results. For example, if the prediction results indicate a high risk of cardiovascular disease, the second analysis results may recommend lifestyle changes, regular medical checkups, or seeking medical consultation and intervention when necessary.
[0110] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application.
[0111] Example 3
[0112] This embodiment provides a human presence detection method based on multi-sensor collaboration of robots, such as Figure 5 As shown, the method includes the following steps.
[0113] S1. Real-time acquisition of image data, point cloud data, and human physiological data of the target area while the robot is walking.
[0114] Specifically, a camera mounted on the robot's head can collect real-time images of a target area (such as a room indoors, or a basement activity hall in a villa), including environmental features and visual information of people present. The image data provides intuitive visual information about the scene, such as human position, posture, and environmental layout. A LiDAR sensor mounted on the robot's bottom collects point cloud data through scanning, capturing the three-dimensional structural information of the target area. Point cloud data enhances the understanding of the three-dimensional shape and spatial position of the environment and objects (including the human body). Millimeter-wave radar sensors surrounding the robot's body collect human physiological data, i.e., real-time monitoring and recording of physiological signals such as heart rate and respiratory rate. This data provides important clues about the human body's current health status and possible physiological stress responses.
[0115] S2. The image data, point cloud data, and human physiological data are fused and analyzed, and decisions are made based on the analysis results to control the robot's response.
[0116] Specifically, all data is first processed by removing clutter and filtering. Then, image data and point cloud data are fused and analyzed to construct a detailed environment and human body model. Through this analysis, the target user's location, posture, and interaction with the environment can be identified, yielding preliminary analysis results. Then, further fusion analysis is performed by combining the clutter-removed and filtered human physiological data with the preliminary analysis results. This process involves using advanced algorithms, such as deep learning and pattern recognition techniques, to understand and predict human health status and potential safety risks. The analysis process is similar to that in Example 2 and will not be elaborated upon here.
[0117] Based on the results of the fusion analysis, the system will make decisions to guide the robot's response behavior. These decisions may include providing necessary support and assistance to the user, warning the user to avoid potential dangers, or adjusting the robot's action plan to better serve the user. For example, if the analysis indicates that a user is at risk of falling, the robot may be instructed to provide an immediate audible warning or move to the user's location to provide physical support.
[0118] The aforementioned human presence detection method based on multi-sensor collaboration of robots can provide users with personalized health management and early warning services by fusing and analyzing the image data, point cloud data, and human physiological data. In particular, this system can act as the first responder to provide medical alarm assistance when a user accidentally falls and loses consciousness, significantly improving the safety and timeliness of emergency care for home users and effectively saving lives.
[0119] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application.
[0120] Example 4
[0121] Please refer to the following. Figure 6 This embodiment provides a schematic diagram of an electronic device. For example... Figure 6 As shown, the electronic device 2 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected via the bus 202. The memory 201 stores a computer program that can run on the processor 200. When the processor 200 runs the computer program, it executes any of the human presence detection methods based on robot multi-sensor collaboration in the embodiments of this application.
[0122] The memory 201 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 203 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network.
[0123] Bus 202 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory 201 is used to store programs. After receiving an execution instruction, the processor 200 executes the program. The human presence detection method based on multi-sensor collaboration of robots disclosed in any of the foregoing embodiments of this application can be applied to the processor 200, or implemented by the processor 200.
[0124] The processor 200 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 200 or by instructions in software form. The processor 200 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 201. The processor 200 reads the information in memory 201 and, in conjunction with its hardware, completes the steps of the human presence detection method based on robot multi-sensor collaboration.
[0125] This embodiment also provides a computer-readable storage medium corresponding to the human presence detection method based on robot multi-sensor collaboration provided in the foregoing embodiments. The medium stores a computer program, which, when executed by a processor, performs the human presence detection method based on robot multi-sensor collaboration provided in any of the foregoing embodiments. Furthermore, examples of the computer-readable storage medium may include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical or magnetic storage media, which will not be elaborated upon here.
[0126] In addition, this application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the aforementioned human presence detection methods based on robot multi-sensor collaboration.
[0127] Those skilled in the art will understand that the various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art should understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation apparatus according to embodiments of this application.
[0128] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A human presence detection system based on multi-sensor collaboration of robots, characterized in that, The human presence detection system includes a robot, a data acquisition module, and a robot operating system. The robot operating system is built into the robot. The data acquisition module includes a vision sensor, a lidar sensor, and a millimeter-wave radar sensor. The vision sensor is located in the head of the robot, the millimeter-wave radar sensor is located in the body of the robot, and the lidar sensor is located in the bottom of the robot. The vision sensor, the lidar sensor, and the millimeter-wave radar sensor are all connected to the robot operating system. The vision sensor is used to collect image data of the target area in real time while the robot is walking and send the image data to the robot operating system; The lidar sensor is used to collect point cloud data of the target area in real time while the robot is walking and send the point cloud data to the robot operating system; The millimeter-wave radar sensor is used to collect human physiological data of the target area in real time while the robot is walking and send the human physiological data to the robot operating system. The robot operating system is used to fuse and analyze the image data, the point cloud data, and the human physiological data, and make decisions based on the analysis results to control the robot's response; The fusion analysis of the image data, the point cloud data, and the human physiological data includes: The point cloud data is filtered to obtain filtered point cloud data; and Based on the filtered point cloud data, a 3D model is constructed to obtain an environment model and a user model. The construction of a 3D model based on the filtered point cloud data, resulting in an environment model and a user model, includes: A three-dimensional model is constructed from the filtered point cloud data using triangulation or voxelization methods to determine the model representing the terrain of the target area, and the model representing the terrain of the target area is used as the environment model. A model representing the active state of the human body is determined, and this model is used as the user model.
2. The human presence detection system based on multi-sensor collaboration of robots according to claim 1, characterized in that, The fusion analysis of the image data, the point cloud data, and the human physiological data further includes: The image data and the human physiological data are filtered respectively to obtain filtered image data and filtered human physiological data. Based on the aforementioned environment model and user model, and the filtered image data, a fusion analysis is performed to obtain the first analysis result for the target user. Based on the first analysis result and the filtered human physiological data, a fusion analysis is performed to obtain the second analysis result for the target user.
3. The human presence detection system based on multi-sensor collaboration of robots according to claim 1, characterized in that, The first analysis result for the target user is obtained by fusing and analyzing the environment model, the user model, and the filtered image data, including: Based on the aforementioned environment model and user model, potential problem users are identified; Feature extraction is performed on the filtered image data, and the target user is determined based on the extracted features and the potential problem users. The target user's physical activity state is analyzed to obtain the first analysis result of the target user, wherein the target user's physical activity state includes the target user's location, the target user's posture, and the target user's interaction state with the environment.
4. The human presence detection system based on multi-sensor collaboration of robots according to claim 1, characterized in that, The first analysis result for the target user is obtained by fusing and analyzing the environment model, the user model, and the filtered image data, including: The processed image data is used to perform pixel mapping on the environment model and user model in order to identify the target user; The target user's physical activity state is analyzed to obtain the first analysis result of the target user, wherein the target user's physical activity state includes the target user's location, the target user's posture, and the target user's interaction state with the environment.
5. The human presence detection system based on robot multi-sensor collaboration according to claim 3 or 4, characterized in that, The step of performing a fusion analysis based on the first analysis result and the filtered human physiological data to obtain the second analysis result for the target user includes: Features are extracted from the filtered human physiological data, and the health information of the target user is determined based on the extracted features. A second analysis result of the target user is obtained by performing a deep fusion analysis on the first analysis result of the target user, the target user's health information, and the target user's historical data.
6. The human presence detection system based on multi-sensor collaboration of robots according to claim 5, characterized in that, The second analysis result of the target user is obtained by deep fusion analysis of the first analysis result of the target user, the target user's health information, and the target user's historical data, including: A long short-term memory network is constructed and trained using the target user's historical data, wherein the target user's historical data includes the target user's historical behavior, activity records, historical medical data, and historical physiological data; The health information of the target user is predicted by a trained long short-term memory network, and the prediction result is obtained. Using the Framingham model, a multiple regression analysis is performed on the prediction results and the first analysis results of the target user to obtain the second analysis results of the target user.
7. A human presence detection method based on multi-sensor collaboration in robots, characterized in that, The method is implemented based on the human presence detection system based on robot multi-sensor collaboration according to any one of claims 1 to 6, and the method includes: The robot collects image data, point cloud data, and human physiological data of the target area in real time while it is walking. The image data, point cloud data, and human physiological data are fused and analyzed, and decisions are made based on the analysis results to control the robot's response.
8. An electronic device comprising a memory and a processor, characterized in that, The memory stores computer-readable instructions, which, when executed by the processor, cause the processor to perform the human presence detection method based on robot multi-sensor collaboration as described in claim 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the human presence detection method based on multi-sensor collaboration of robots as described in claim 7.
Citation Information
Patent Citations
Pedestrian identification method and system
CN112464782A
Life body detection system and method for fire scene
CN117214965A