Visual attitude fusion-based health care robot ethical risk monitoring method
Through visual posture fusion technology, using RGB cameras, depth sensors and sensors to collect data, combined with OpenPose and MediaPipe models for posture analysis, the data bias problem in the ethical risk monitoring of health care robots is solved, and real-time, multi-target detection and accurate assessment of ethical risks are achieved, improving the stability and accuracy of monitoring.
Patent Information
- Application Number
- CN202510574411.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing health care robots have 3D accuracy and occlusion problems in data calculation, resulting in the loss of key points, large data calculation deviations, and inaccurate ethical risk monitoring.
A method based on visual posture fusion is adopted. Data is collected through RGB cameras and depth sensors, and the posture data is supplemented by inertial measurement units and pressure sensors. OpenPose and MediaPipe real-time posture estimation models are used for posture analysis. Kalman filtering and neural network are combined to fuse data, quantify ethical risks and set response mechanisms.
It realizes real-time, multi-target detection and multi-modal output, improves the real-time and robustness of health care robots, enhances the stability and accuracy of ethical risk monitoring, and protects user privacy and dignity.
Smart Images

Figure FDA0005388096830000021 
Figure FDA0005388096830000023 
Figure FDA0005388096830000032
Abstract
Description
Technical Field
[0001] The present invention relates to a method for monitoring ethical risks of a health care robot based on visual posture fusion. Background Art
[0002] Ethical risk monitoring for healthcare robots is a multidisciplinary research area that requires the integration of computer vision, robotics ethics, and human-computer interaction technologies. This involves extensive data computation, and in actual deployment, key points can be lost due to issues like 3D precision and occlusion, leading to significant deviations in data computation. Summary of the Invention
[0003] The technical problem to be solved by the present invention is: In order to overcome the above technical problems, the present invention provides a health care robot ethical risk monitoring method based on visual posture fusion.
[0004] The technical solution adopted by the present invention to solve the technical problem is: a method for monitoring ethical risks of a health care robot based on visual posture fusion, comprising the following steps:
[0005] Step 1: Multimodal data collection:
[0006] a1. Visual data acquisition: Capture user posture through RBG camera and depth sensor;
[0007] a2. Sensor assistance: supplement attitude data through inertial measurement units and pressure sensors to reduce visual blind spots;
[0008] a3. Data anonymization: blurring background or facial information in real time;
[0009] Step 2: Posture analysis and intention recognition:
[0010] b1. Visual end algorithm selection: based on OpenPose and MediaPipe real-time pose estimation models;
[0011] b2. Fusion algorithm selection: Use Kalman filtering or neural networks to fuse vision and IMU data to improve posture tracking robustness;
[0012] b3.,Intention inference: Combining gesture temporal features with context,infer gesture intention;
[0013] Step 3: Quantify ethical risks:
[0014] c1. Privacy risk score: calculated based on the camera coverage area and data storage duration;
[0015] c2. Dignity risk marker: detects whether the robot's behavior violates user habits;
[0016] c3. Autonomy assessment: Count the ratio of user-initiated actions to robot interventions;
[0017] Step 4: Risk response: The system presets a risk response mechanism and responds to risks based on the different risk behaviors generated by the robot.
[0018] Preferably, in step c1 of step 3, the data storage time is 30 minutes.
[0019] Preferably, in step 4, the risk response mechanism includes three types: low risk, medium risk and high risk. If the robot touches the area briefly by mistake, it is defined as low risk, and the specific response method is: record the log and notify the administrator; if the robot continuously monitors the sensitive area for more than 30 minutes, it is defined as medium risk, and the specific response method is: suspend the service and request user confirmation; if the robot engages in abusive behavior, it is defined as high risk, and the specific response method is: emergency interruption of service and alarm.
[0020] Preferably, in step 2, OpenPose adopts a two-stage bottom-up detection strategy, including key point detection and key point association. Key point detection is to predict the confidence map of human key points through CNN, and key point association is to combine key points into a complete human posture through component affinity fields. OpenPose uses the VGG-19 network to extract basic features and realize multi-branch prediction, including generating key point confidence maps and generating component affinity fields to describe the connection relationship between key points. Each key point corresponds to a heat map, and the peak position indicates the possible location of the key point.
[0021] Preferably, the confidence map is calculated as follows: Where P is the coordinate of any pixel in the image, x j,k is the true position of the kth key point of the jth person, σ is the hyperparameter controlling the width of the Gaussian kernel, usually set to 7px. If P is close to someone's key point x j,k , then the Sk(p) value is close to 1, otherwise it is close to 0.
[0022] As a preference, when P is within the band region of the limb, the vector field formula of the component affinity field is: Lc When P is outside the strip area of the limb, Lc(p) = 0, where C is the limb type, v is the limb direction vector, and the component affinity field assigns a unit direction vector to each pixel in the limb area for subsequent key point matching.
[0023] Preferably, in step 2, the architecture of the MediaPipe pose estimation model includes a lightweight backbone network, stage detection, and 3D pose output, wherein the stage detection includes using a practical lightweight CNN to detect the human body bounding box and predict 33 3D human key points within the ROI, and the 3D pose output supports converting 2D key points into 3D coordinates.
[0024] As a preference, the loss function of the MediaPipe pose estimation model is: Among them, P i * is the true keypoint coordinate, λ balances the 2D / 3D error weight, P i To predict the key point coordinates, P i =(x i ,y i ,z i ), where z i is the relative depth.
[0025] The beneficial effect of the present invention is that the ethical risk monitoring method of the health care robot based on visual posture fusion performs real-time data processing based on the OpenPose and MediaPipe real-time posture estimation models, realizes the functions of multi-target detection and multi-modal output, improves real-time and robustness, supports multi-tasking processing, can realize real-time multimedia processing, and improves the stability and accuracy of monitoring. DETAILED DESCRIPTION
[0026] The present invention provides a method for monitoring ethical risks of a health care robot based on visual posture fusion.
[0027] The technical solution adopted by the present invention to solve the technical problem is: a method for monitoring ethical risks of a health care robot based on visual posture fusion, comprising the following steps:
[0028] Step 1: Multimodal data collection:
[0029] a1. Visual data acquisition: Capture user posture through RBG camera and depth sensor;
[0030] a2. Sensor assistance: supplement attitude data through inertial measurement units and pressure sensors to reduce visual blind spots;
[0031] a3. Data anonymization: blurring background or facial information in real time;
[0032] Step 2: Posture analysis and intention recognition:
[0033] b1. Visual end algorithm selection: based on OpenPose and MediaPipe real-time pose estimation models;
[0034] b2. Fusion algorithm selection: Use Kalman filtering or neural networks to fuse vision and IMU data to improve posture tracking robustness;
[0035] b3.,Intention inference: Combining gesture temporal features with context,infer gesture intention;
[0036] Step 3: Quantify ethical risks:
[0037] c1. Privacy risk score: calculated based on the camera coverage area and data storage duration;
[0038] c2. Dignity risk marker: detects whether the robot's behavior violates user habits;
[0039] c3. Autonomy assessment: Count the ratio of user-initiated actions to robot interventions;
[0040] Step 4: Risk response: The system presets a risk response mechanism and responds to risks based on the different risk behaviors generated by the robot.
[0041] In health and wellness scenarios, possible ethical risks include:
[0042] 1. Privacy infringement, that is, whether visual data collection exceeds the necessary scope, such as in bathrooms or bedrooms where privacy is a priority;
[0043] 2. Dignity risk: whether the robot's behavior causes users to feel objectified, such as inappropriate physical contact or gesture recognition;
[0044] 3. Autonomy deprivation: Over-reliance on robots leads to the degradation of users’ autonomy;
[0045] 4. Algorithmic bias: posture recognition models may misjudge specific groups, such as people with disabilities and the elderly.
[0046] The pressure sensors here can be set on the floor or under the mattress to assist in supplementing posture data and reducing visual blind spots.
[0047] The data here is anonymized in compliance with GDPR privacy rules.
[0048] In ethical risk monitoring, a real-time monitoring module is adopted, which includes a rule engine and a dynamic learning module.
[0049] The rule engine is used to preset ethical thresholds, such as triggering an alarm if a camera is on for more than 30 minutes;
[0050] The dynamic learning module optimizes risk judgment strategies through reinforcement learning, such as correcting false positives based on user feedback.
[0051] In fact, the accuracy of judgment can be improved through human-machine collaborative review.
[0052] For example, a visual dashboard can display a heat map of ethical risks, making it easy to see areas with high incidence of privacy violations.
[0053] Manual review was adopted, and nursing staff were introduced to make the final judgment on events marked by the algorithm.
[0054] Example 1
[0055] The robot uses vision to detect that the elderly person is trying to use the milling machine by himself, but posture analysis shows that the risk of falling is high. The robot's ethical decision-making process can be:
[0056] 1. First, ask if you need help. This is to respect your autonomy.
[0057] 2. If there is no response, slowly approach and lightly touch the arm to provide support. This is to avoid sudden contact that may cause discomfort to the user.
[0058] 3. The facial data is hidden in the time record, and the skeleton key point information is retained. This is mainly to protect user privacy.
[0059] Here, by encoding posture recognition technology and ethical principles, the health care robot can achieve a balance between "effective care" and "ethical bottom line".
[0060] Preferably, in step c1 of step 3, the data storage time is 30 minutes.
[0061] Preferably, in step 4, the risk response mechanism includes three types: low risk, medium risk and high risk. If the robot touches the area briefly by mistake, it is defined as low risk, and the specific response method is: record the log and notify the administrator; if the robot continuously monitors the sensitive area for more than 30 minutes, it is defined as medium risk, and the specific response method is: suspend the service and request user confirmation; if the robot engages in abusive behavior, it is defined as high risk, and the specific response method is: emergency interruption of service and alarm.
[0062] Preferably, in step 2, OpenPose adopts a two-stage bottom-up detection strategy, including key point detection and key point association. Key point detection is to predict the confidence map of human key points through CNN, and key point association is to combine key points into a complete human posture through component affinity fields. OpenPose uses the VGG-19 network to extract basic features and realize multi-branch prediction, including generating key point confidence maps and generating component affinity fields to describe the connection relationship between key points. Each key point corresponds to a heat map, and the peak position indicates the possible location of the key point.
[0063] Preferably, the confidence map is calculated as follows: Where P is the coordinate of any pixel in the image, x j,k is the true position of the kth key point of the jth person, σ is the hyperparameter controlling the width of the Gaussian kernel, usually set to 7px. If P is close to someone's key point x j,k , then the Sk(p) value is close to 1, otherwise it is close to 0.
[0064] As a preference, when P is within the band region of the limb, the vector field formula of the component affinity field is: Lc When P is outside the strip area of the limb, Lc(p) = 0, where C is the limb type, v is the limb direction vector, and the component affinity field assigns a unit direction vector to each pixel in the limb area for subsequent key point matching.
[0065] OpenPose is a real-time multi-person pose estimation algorithm based on deep learning. It can simultaneously detect the key points of the body, hands, face and feet of multiple people from images or videos.
[0066] Its core features are:
[0067] Multi-target detection: Simultaneously estimate the poses of all people in an image without having to detect the number of people in advance;
[0068] Multimodal output: supports joint or separate detection of body (25 keypoints), hands (42 keypoints), and face (70 keypoints).
[0069] Real-time performance: After optimization, it can achieve real-time processing on ordinary GPUs;
[0070] Robustness: It has certain adaptability to occlusion and complex background.
[0071] Key steps include:
[0072] Input RGB image, image resolution ≥ 368*368;
[0073] Use VGG-19 network to extract basic features;
[0074] Generate confidence maps and component affinity fields;
[0075] The key points are associated into a complete human skeleton using the Hungarian algorithm.
[0076] In the confidence map here, each key point corresponds to a heat map, and the peak position indicates the possible location of the key point. For example, the heat map of the left shoulder key point will highlight all left shoulder positions in the image.
[0077] The component affinity field here defines the connection vector field between key points and solves the key point attribution problem in multi-person scenarios, such as determining which wrists and elbows belong to the same person.
[0078] Here, feature maps of different resolutions are combined to improve the detection effect of small objects. The confidence map and component affinity field are gradually optimized through multiple stages, usually 6 stages.
[0079] This allows you to:
[0080] 1. Determine the risk of falling by detecting sudden displacement of key posture points, such as the head and hips.
[0081] 2. Only use skeleton key point data for behavior analysis to protect privacy.
[0082] 3. Check whether the distance between the robot and the user and the contact points comply with ethical standards.
[0083] Example 2
[0084] In a healthcare scenario, if an elderly person falls, the algorithm will:
[0085] 1. Localize the hip and head using confidence maps;
[0086] 2. Calculate whether the direction of the hips and head suddenly becomes vertical;
[0087] 3. If || L torso (0,1)||<∈ (the dot product with the ground normal vector is too small), triggering a fall alarm.
[0088] Preferably, in step 2, the architecture of the MediaPipe pose estimation model includes a lightweight backbone network, stage detection, and 3D pose output, wherein the stage detection includes using a practical lightweight CNN to detect the human body bounding box and predict 33 3D human key points within the ROI, and the 3D pose output supports converting 2D key points into 3D coordinates.
[0089] As a preference, the loss function of the MediaPipe pose estimation model is: Among them, P i * is the true keypoint coordinate, λ balances the 2D / 3D error weight, P i To predict the key point coordinates, P i =(x i ,y i ,z i ), where z i is the relative depth.
[0090] MediaPipe is an open source cross-platform framework developed by Google, focusing on real-time multimedia processing, such as gesture recognition, pose estimation, and face tracking. It supports multi-task parallel processing, provides out-of-the-box models such as gesture recognition, iris tracking, object detection, and supports privacy protection.
[0091] Its model architecture includes a lightweight backbone network based on a cropped version of MobileNetV2 or EfficientNet, uses a lightweight CNN to detect the human body bounding box (ROI), predicts 33 3D human key points within the ROI, including face and hands, supports conversion of 2D key points into 3D coordinates, and is suitable for rehabilitation motion analysis.
[0092] The framework supports local processing without the need for cloud computing, which can better protect user privacy.
[0093] This framework quantizes the model and uses TFLite to convert the model from FP32 to INT8, which increases the speed by 3 times. It dynamically adjusts the input image resolution based on distance, such as using 256x256 for long distances, and enables Android's GPU inference or Coral TPU acceleration to achieve hardware acceleration.
[0094] MediaPipe is suitable for mobile devices, embedded devices, and real-time interactive scenarios, while OpenPose is suitable for high-precision research and large-scale multi-person scenarios. The combination of the two can meet the needs of various scenarios, expand the scope of application, balance accuracy and efficiency, greatly enhance the level of privacy protection, and thus improve the adaptability of robots and the accuracy of data.
[0095] The beneficial effect of the present invention is that the ethical risk monitoring method of a health care robot based on visual posture fusion of the present invention performs real-time data processing based on the OpenPose and MediaPipe real-time posture estimation models, realizes the functions of multi-target detection and multi-modal output, improves real-time and robustness, supports multi-tasking processing, can realize real-time multimedia processing, and improves the stability and accuracy of monitoring.
[0096] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification and must be determined according to the scope of the claims.
Claims
1. A method for monitoring ethical risks of health care robots based on visual posture fusion, characterized in that: The following steps are included: Step 1: Multimodal data collection: a1. Visual data acquisition: Capture user posture through RBG camera and depth sensor; a2. Sensor assistance: supplement attitude data through inertial measurement units and pressure sensors to reduce visual blind spots; a3. Data anonymization: blurring background or facial information in real time; Step 2: Posture analysis and intention recognition: b1. Visual end algorithm selection: based on OpenPose and MediaPipe real-time pose estimation models; b2. Fusion algorithm selection: Use Kalman filtering or neural networks to fuse vision and IMU data to improve posture tracking robustness; b3.,Intention inference: Combining gesture temporal features with context,infer gesture intention; Step 3: Quantify ethical risks: c1. Privacy risk score: calculated based on the camera coverage area and data storage duration; c2. Dignity risk marker: detects whether the robot's behavior violates user habits; c3. Autonomy assessment: Count the ratio of user-initiated actions to robot interventions; Step 4: Risk response: The system presets a risk response mechanism and responds to risks based on the different risk behaviors generated by the robot.
2. The method for monitoring ethical risks of a healthcare robot based on visual posture fusion according to claim 1, characterized in that: In step c1 of step 3, the data storage duration is 30 minutes.
3. The method for monitoring ethical risks of a healthcare robot based on visual posture fusion according to claim 1, characterized in that: In step 4, the risk response mechanism includes low risk, medium risk, and high risk. If the robot touches the area by mistake briefly, it is defined as low risk, and the specific response method is: record the log and notify the administrator; if the robot continuously monitors the sensitive area for more than 30 minutes, it is defined as medium risk, and the specific response method is: suspend the service and request user confirmation; if the robot engages in abusive behavior, it is defined as high risk, and the specific response method is: emergency interruption of service and alarm.
4. The method for monitoring ethical risks of a healthcare robot based on visual posture fusion according to claim 1, characterized in that: In step 2, OpenPose adopts a two-stage bottom-up detection strategy, including key point detection and key point association. Key point detection is to predict the confidence map of human key points through CNN. Key point association is to combine key points into a complete human pose through component affinity fields. OpenPose uses the VGG-19 network to extract basic features and implement multi-branch prediction, including generating key point confidence maps and component affinity fields to describe the connection relationship between key points. Each key point corresponds to a heat map, and the peak position indicates the possible location of the key point.
5. The method for monitoring ethical risks of a healthcare robot based on visual posture fusion as claimed in claim 4, characterized in that: The calculation formula of the confidence map is: Where P is the coordinate of any pixel in the image, x j,k is the true position of the kth key point of the jth person, σ is the hyperparameter controlling the width of the Gaussian kernel, usually set to 7px. If P is close to someone's key point x j,k , then the Sk(p) value is close to 1, otherwise it is close to 0.
6. The method for monitoring ethical risks of a healthcare robot based on visual posture fusion as claimed in claim 4, characterized in that: When P is within the strip area of the limb, the vector field formula of the component affinity field is: When P is outside the strip area of the limb, Lc(p) = 0, where C is the limb type, v is the limb direction vector, and the component affinity field assigns a unit direction vector to each pixel in the limb area for subsequent key point matching.
7. The method for monitoring ethical risks of a healthcare robot based on visual posture fusion according to claim 1, characterized in that: In step 2, the architecture of the MediaPipe pose estimation model includes a lightweight backbone network, stage detection, and 3D pose output. The stage detection includes using a lightweight CNN to detect the human bounding box and predict 33 3D human key points within the ROI. The 3D pose output supports converting 2D key points into 3D coordinates.
8. The method for monitoring ethical risks of a healthcare robot based on visual posture fusion according to claim 7, characterized in that: The loss function of the MediaPipe pose estimation model is: in, is the true keypoint coordinate, λ balances the 2D / 3D error weight, P i To predict the key point coordinates, P i =(x i ,y i ,z i ), where z i is the relative depth.