Multimodal large model elderly care robot 6d pose estimation compliant delivery system
Patent Information
- Application Number
- CN202611074688.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-25
AI Technical Summary
然而,在实际养老服务场景中,所夹持的目标物体种类多样,涉及材质形态各异的日常物件,始终输出恒定的机械夹持力,无法根据物体的软硬程度、形态结构实时调整夹持力度,易因作用力过载造成物件挤压变形、破损碎裂,降低了设备的使用可靠性与用户使用体验
[0017]与现有技术相比,本申请具有如下有益效果:本申请的多模态大模型养老机器人6D位姿估计柔顺递送系统,通过设置交互识别单元,实时采集作业场景多维环境参数数据并完成预处理,为分析定位单元提供决策基础,通过分析定位单元补全遮挡缺失的点云信息,避免居家环境杂物遮挡、物体轮廓残缺引发的目标定位失效问题,同时,抓取递送单元利用六维力传感器与柔性触觉表皮阵列传感器构建多维感知体系,结合抓取稳定性评分网络优化抓取策略,在提高抓取稳定性的同时防止易碎及柔性居家物品受挤压破损,进而提高了机器人抓取作业的柔顺性与可靠性,适配居家养老复杂作业环境与老年用户的实际使用需求。
Smart Images

Figure CN122807895A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of elderly care service technology, and in particular to a compliant delivery system for 6D pose estimation of a multimodal large-scale elderly care robot. Background Technology
[0002] With the rapid iteration and empowerment of cutting-edge technologies such as artificial intelligence, the Internet of Things, sensing technology, and biomimetic mechanical engineering, the related technological system of elderly care robots is becoming increasingly mature. Elderly care robots are intelligent service devices specifically developed for the elderly population and applied in diverse scenarios such as homes, communities, and elderly care institutions. They can provide comprehensive and intelligent elderly care services, including daily living assistance, health monitoring, rehabilitation support, emotional companionship, and safety protection. They can adapt to the personalized elderly care needs of special groups such as the disabled, empty-nest, and very old, and have been widely used in family-based and institutional elderly care scenarios, becoming an important intelligent carrier of the smart elderly care system.
[0003] Current elderly care robots employ relatively fixed grasping methods. They typically use a vision module to identify the outline of the target object, and then a mechanical gripping mechanism performs a series of actions according to preset fixed program parameters, including gripping and closing, lifting and moving, positioning and placing, and resetting. However, in actual elderly care service scenarios, the types of target objects being gripped are diverse, involving everyday objects of varying materials and shapes. Maintaining a constant mechanical gripping force, without adjusting the gripping force in real time according to the object's hardness, shape, and structure, can easily lead to overload, causing objects to be squeezed, deformed, or broken, reducing the reliability of the equipment and the user experience. Summary of the Invention
[0004] This application aims to at least partially solve one of the technical problems in the aforementioned technologies.
[0005] Therefore, one objective of this application is to propose a 6D pose estimation compliant delivery system for a multimodal large-scale elderly care robot. This system can utilize a six-dimensional force sensor and a flexible tactile epidermal array sensor to construct a multi-dimensional perception system. Combined with a grasping stability scoring network, the grasping strategy is optimized. This improves grasping stability while preventing fragile and flexible household items from being squeezed and damaged, thereby enhancing the compliance and reliability of the robot's grasping operation.
[0006] To achieve the above objectives, the first aspect of this application proposes a 6D pose estimation compliant delivery system for a multimodal large-scale model elderly care robot. The system includes: an interaction recognition unit for acquiring multidimensional environmental parameter data of the work scene when the elderly care robot receives an open-ended natural language interaction command input by a human; an analysis and localization unit for determining the spatial location of a target object in a complex occlusion scenario based on the acquired multidimensional environmental parameter data during the robot's movement; and a grasping and delivery unit for controlling the movement of the robot's robotic arm, adjusting the grasping reference position and posture, and grasping and delivering the target object according to the judgment result.
[0007] In addition, the multimodal large-model elderly care robot 6D pose estimation compliant delivery system proposed in this application may also have the following additional technical features:
[0008] Specifically, the multidimensional environmental parameter data includes visual image data, depth distance data, and human-computer interaction posture data. After acquiring the multidimensional environmental parameter data of the work scene, the interaction recognition unit is further used to: filter, calibrate, align, and preprocess the multidimensional environmental parameter data.
[0009] Specifically, the step of determining the spatial location of a target object in a complex occlusion scenario based on the acquired multidimensional environmental parameter data includes: using the Grounding DINO visual positioning model to determine the target object region in the scene corresponding to the semantic command; using the SAM segmentation model to perform pixel-level semantic segmentation of the target object, and based on the calibrated depth and distance data, performing a mapping transformation from the pixel coordinate system to the three-dimensional spatial coordinate system on the segmented target region to extract the original spatial point cloud data of the target object; using the Diffusion Policy diffusion strategy algorithm to perform dual optimization processing on the original target point cloud based on the extracted original spatial point cloud data; and using a preset spatial coordinate calculation algorithm to calculate the spatial positioning information and pose parameters of the target object based on the optimized original target point cloud data to determine the spatial location of the target object in a complex occlusion scenario.
[0010] Specifically, the Diffusion Policy algorithm is used to perform dual optimization processing on the original target point cloud, including: acquiring the original spatial point cloud data and performing point cloud denoising processing on the original spatial point cloud data; using the massive object shape prior knowledge learned by the pre-trained Diffusion Policy model to infer and restore the spatial point information of occluded or missing objects, and reconstructing a complete target object point cloud model.
[0011] Specifically, training the Diffusion Policy diffusion strategy model includes: acquiring complete 3D point cloud samples containing multiple common household objects and partially occluded point cloud samples to construct a training dataset; using the partially occluded point cloud samples as network input and the complete 3D point cloud samples as ground truth labels, iteratively training the Diffusion Policy diffusion strategy network; and obtaining the pre-trained Diffusion Policy diffusion strategy model after the network parameters converge.
[0012] Specifically, the grasping and delivery unit includes: a data receiving and acquisition module: used to collect multi-dimensional perception data of the robotic arm gripper in real time, and to fuse and summarize the collected multi-dimensional perception data with the spatial positioning information and pose parameters of the target object output by the analysis and positioning unit; a data analysis and processing module: used to calculate and select the optimal grasping strategy based on the summarized data using a pre-trained grasping stability scoring network, which includes grasping posture and grasping force parameters; an adjustment module: used to adjust the robotic arm motion trajectory and gripper holding parameters according to the optimal grasping strategy, and control the robotic arm to complete the grasping of the target object; and a delivery path optimization module: used to identify the user's human posture and hand movement data, determine the user's human-computer interaction intention, optimize the robotic arm delivery path, and achieve safety protection during the delivery process of the elderly care robot.
[0013] Specifically, the multidimensional sensing data includes real-time contact pressure and contact torque data of the gripper collected using a six-dimensional force sensor, and contact area and adhesion distribution data of the object collected using a flexible tactile epidermal array sensor.
[0014] Specifically, the step of using a pre-trained grasping stability scoring network to calculate and select the optimal grasping strategy includes: when the target object is a rigid object, controlling the robotic arm gripper to use a pinching strategy to grasp the target object; when the target object is a flexible and easily deformable object, controlling the robotic arm gripper to use a holding strategy to grasp the target object.
[0015] Specifically, after recognizing the user's human posture and hand movement data, the delivery path optimization module is further used to: use an LSTM temporal intent classification network to extract temporal features of hand movement speed, displacement amplitude, and movement trend, and compare the extracted temporal features with a preset threshold to determine the user's interaction intent in real time.
[0016] Specifically, determining the user's human-computer interaction intent and optimizing the robotic arm delivery path includes: when an external collision force is detected to be greater than a preset safety threshold, controlling the robotic arm to reverse and stop abruptly, releasing the end torque of the robotic arm and the gripper's clamping force, terminating the operation, and achieving safety protection during the delivery process of the elderly care robot.
[0017] Compared with existing technologies, this application has the following advantages: The multimodal large-scale model elderly care robot 6D pose estimation compliant delivery system of this application, by setting up an interactive recognition unit, collects multi-dimensional environmental parameter data of the operation scene in real time and completes preprocessing, providing a decision basis for the analysis and positioning unit. By analyzing and positioning the unit to complete the missing point cloud information due to occlusion, it avoids the target positioning failure problem caused by clutter in the home environment and incomplete object outlines. At the same time, the grasping and delivery unit uses a six-dimensional force sensor and a flexible tactile epidermal array sensor to build a multi-dimensional perception system, and combines the grasping stability scoring network to optimize the grasping strategy. While improving the grasping stability, it prevents fragile and flexible home items from being squeezed and damaged, thereby improving the compliance and reliability of the robot's grasping operation, and adapting to the complex operation environment of home-based elderly care and the actual use needs of elderly users.
[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0019] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0020] Figure 1 This is a schematic diagram of the overall modules of the present invention. Detailed Implementation
[0021] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0022] The following describes, with reference to the accompanying drawings, a multimodal large-scale elderly care robot 6D pose estimation compliant delivery system according to an embodiment of this application.
[0023] like Figure 1 As shown in this embodiment, a 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot is provided. This system may include:
[0024] Interaction recognition unit: used to obtain multi-dimensional environmental parameter data of the working scene when the elderly care robot receives open natural language interaction instructions input by humans.
[0025] The multidimensional environmental parameter data includes visual image data, depth distance data, and human-computer interaction posture data. After acquiring the multidimensional environmental parameter data of the work scene, the interaction recognition unit is further used to: filter, calibrate, align, and preprocess the multidimensional environmental parameter data.
[0026] For example, the open-ended natural language interaction commands input by humans are arbitrary and free semantic commands in the elderly care home scenario, including but not limited to natural language control commands with non-fixed sentence structures such as "hand me a glass of water", "hand me the medicine box on the table", and "pick up the tableware on the table and hand it to me". The robot can parse various open-ended semantic commands, obtain the corresponding task objectives and delivery requirements, and then collect multi-dimensional environmental parameter data that matches the current task scenario.
[0027] Analysis and positioning unit: used to determine the spatial positioning of the target object in complex occlusion scenarios based on the acquired multi-dimensional environmental parameter data during the movement of the elderly care robot.
[0028] The method for determining the spatial location of a target object in a complex occlusion scenario based on the acquired multidimensional environmental parameter data includes: using the Grounding DINO visual positioning model to determine the target object region in the scene corresponding to the semantic command; using the SAM segmentation model to perform pixel-level semantic segmentation of the target object, and based on the calibrated depth and distance data, performing a mapping transformation from the pixel coordinate system to the three-dimensional spatial coordinate system on the segmented target region to extract the original spatial point cloud data of the target object; using the Diffusion Policy diffusion strategy algorithm to perform dual optimization processing on the original target point cloud data based on the extracted original spatial point cloud data; and using a preset spatial coordinate calculation algorithm to calculate the spatial positioning information and pose parameters of the target object based on the optimized original target point cloud data to determine the spatial location of the target object in a complex occlusion scenario.
[0029] The Diffusion Policy algorithm is used to perform dual optimization processing on the original target point cloud, including: acquiring the original spatial point cloud data and performing point cloud denoising processing on the original spatial point cloud data; using the massive object shape prior knowledge learned by the pre-trained Diffusion Policy model, reasoning and restoring the spatial point information of occluded and missing objects, and reconstructing a complete target object point cloud model.
[0030] Training the Diffusion Policy diffusion strategy model includes: acquiring complete 3D point cloud samples containing multiple common household objects and partially occluded point cloud samples to construct a training dataset; using the partially occluded point cloud samples as network input and the complete 3D point cloud samples as ground truth labels, iteratively training the Diffusion Policy diffusion strategy network; and obtaining the pre-trained Diffusion Policy diffusion strategy model after the network parameters converge.
[0031] For example, when an elderly care robot performs object delivery tasks, there are common interferences in the scene, such as furniture occlusion, piled-up clutter, and partial occlusion by human bodies. Target objects, such as water cups, medicine boxes, tableware, and toiletries, which are commonly used household objects, are prone to having incomplete outlines and missing features. Through the multi-model collaborative perception and point cloud optimization and reconstruction method mentioned above, the point cloud missing problem caused by occlusion can be repaired, and the three-dimensional spatial coordinates and 6D pose parameters of the target object can be output. This enables the spatial positioning of various household objects in complex occlusion scenarios, providing data support for the robot's compliant grasping and safe delivery operations.
[0032] Training the Diffusion Policy diffusion strategy model includes: acquiring complete 3D point cloud samples containing multiple common household objects and partially occluded point cloud samples to construct a training dataset; using the partially occluded point cloud samples as network input and the complete 3D point cloud samples as ground truth labels, iteratively training the Diffusion Policy diffusion strategy network; and obtaining the pre-trained Diffusion Policy diffusion strategy model after the network parameters converge.
[0033] For example, the collected common household objects may include water cups, medicine boxes, tableware, tissue boxes, toiletries, and other frequently used objects in elderly care scenarios. The corresponding partial occlusion defect cloud samples simulate real-world interference conditions in the home, including point cloud defect samples caused by partial contour loss and point incompleteness due to occlusion by books, tableware, hands, and furniture. By enabling the Diffusion Policy network to fully learn the general three-dimensional contour features, topological structure, and morphological distribution patterns of various household objects, the accuracy of point cloud reconstruction and the robustness of localization of the elderly care robot in complex occlusion scenarios in real homes can be improved.
[0034] The grasping and delivery unit is used to control the movement of the robotic arm of the elderly care robot based on the judgment result, adjust the grasping reference position and posture, and realize the grasping and delivery of the target object. It includes: a data receiving and acquisition module, used to collect multi-dimensional perception data of the robotic arm gripper in real time, and to fuse and summarize the collected multi-dimensional perception data with the spatial positioning information and pose parameters of the target object output by the analysis and positioning unit; a data analysis and processing module, used to calculate and select the optimal grasping strategy based on the summarized data using a pre-trained grasping stability scoring network, including grasping posture and grasping force parameters; an adjustment module, used to adjust the robotic arm's motion trajectory and gripper holding parameters according to the optimal grasping strategy, controlling the robotic arm to complete the grasping of the target object; and a delivery path optimization module, used to identify the user's human posture and hand movement data, determine the user's human-computer interaction intention, optimize the robotic arm's delivery path, and achieve safety protection during the elderly care robot's delivery process.
[0035] The multidimensional sensing data includes real-time contact pressure and contact torque data of the gripper collected using a six-dimensional force sensor, and contact area and adhesion distribution data of the object collected using a flexible tactile epidermal array sensor.
[0036] For example, a six-dimensional force sensor provides real-time feedback on the contact pressure and torque changes between the gripper and the target object, while a flexible tactile skin array sensor provides real-time feedback on the object's contact area and contact distribution, distinguishing the contact differences between hard and flexible objects. Then, by fusing the above multi-dimensional sensing data with the target object's three-dimensional positioning and pose information, the system uses a gripping stability scoring network to adaptively match the optimal gripping posture and clamping force, avoiding problems such as object crushing and breakage, unstable gripping and slippage caused by rigid gripping.
[0037] The method of using a pre-trained grasping stability scoring network to calculate and select the optimal grasping strategy includes: when the target object is a rigid object, controlling the robotic arm gripper to use a pinching strategy to grasp the target object; when the target object is a flexible and easily deformable object, controlling the robotic arm gripper to use a holding strategy to grasp the target object.
[0038] For example, rigid objects can be specific items in a home setting such as medicine boxes, metal water cups, and hard tableware that are not easily deformed. By using a fixed-point alignment gripping method, the stability of the clamping contact point is ensured, and deviation and slippage are avoided. Flexible and easily deformable objects can be specific items such as tissues, sponges, cloth bags, and soft foods that are easily squeezed and deformed. By using a full-coverage gripping method, the contact area is increased and the clamping pressure is reduced, ensuring a secure grip while avoiding the object being squeezed, broken, or deformed and failing. This achieves differentiated and compliant gripping operations for home objects with different properties.
[0039] The user's body posture and hand movement data include the real-time position and movement trajectory of the user's hands, arms and torso. After recognizing the user's body posture and hand movement data, the delivery path optimization module is further used to: use an LSTM temporal intent classification network to extract the temporal features of hand movement speed, displacement amplitude and movement trend, and compare the extracted temporal features with a preset threshold to determine the user's interaction intent in real time.
[0040] The process of determining the user's human-computer interaction intent and optimizing the robotic arm delivery path includes: when an external collision force is detected to be greater than a preset safety threshold, controlling the robotic arm to stop in the opposite direction, releasing the end torque of the robotic arm, removing the gripper's clamping force, terminating the operation, and achieving safety protection during the delivery process of the elderly care robot.
[0041] For example, during the dynamic interaction of the robot delivering objects, by collecting continuous temporal motion information such as the speed of the elderly person's hand extension, the amplitude of hand displacement, and the forward-leaning posture of the torso, the system uses an LSTM temporal intent classification network to capture subtle differences in motion features. This distinguishes different interaction states, such as the elderly person actively reaching for the object, hesitating in their actions, or avoiding or refusing the object. The system then adaptively adjusts the delivery rhythm and movement path of the robotic arm. Furthermore, during the human-robot interaction delivery process, the system monitors the contact force state of the robotic arm in real time. When abnormalities such as accidental touch by the elderly person or sudden obstruction by the limb occur, an emergency stop and force relief protection mechanism can be triggered to avoid the robotic arm's rigid impact causing squeezing or bumping injuries to the elderly user. This improves the compliance and operational safety of the elderly care robot's human-robot interaction delivery.
[0042] In summary, the multimodal large-scale model elderly care robot 6D pose estimation compliant delivery system of this application, by setting up an interactive recognition unit, collects multi-dimensional environmental parameter data of the working scene in real time and completes preprocessing, providing a decision basis for the analysis and localization unit. By analyzing and localizing the missing point cloud information due to occlusion, it avoids the target localization failure problem caused by clutter in the home environment and incomplete object outlines. At the same time, the grasping and delivery unit uses a six-dimensional force sensor and a flexible tactile epidermal array sensor to construct a multi-dimensional perception system, and combines it with a grasping stability scoring network to optimize the grasping strategy. While improving the grasping stability, it prevents fragile and flexible home items from being squeezed and damaged, thereby improving the compliance and reliability of the robot's grasping operation, and adapting to the complex working environment of home-based elderly care and the actual use needs of elderly users.
[0043] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0044] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0045] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A compliant delivery system for 6D pose estimation of a multimodal large-scale elderly care robot, characterized in that, The multimodal large-scale elderly care robot 6D pose estimation compliant delivery system includes: Interaction recognition unit: used to acquire multi-dimensional environmental parameter data of the working scene when the elderly care robot receives open natural language interaction instructions input by humans; Analysis and positioning unit: used to determine the spatial positioning of the target object in complex occlusion scenarios based on the acquired multi-dimensional environmental parameter data during the movement of the elderly care robot; Grasping and delivery unit: Based on the judgment result, it controls the movement of the robotic arm of the elderly care robot, adjusts the grasping reference position and posture, and realizes the grasping and delivery of the target object.
2. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 1, characterized in that, The multidimensional environmental parameter data includes visual image data, depth distance data, and human-computer interaction posture data. After acquiring the multidimensional environmental parameter data of the work scene, the interaction recognition unit is further used for: The multidimensional environmental parameter data is filtered, calibrated, aligned, and preprocessed.
3. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 2, characterized in that, The method of determining the spatial location of a target object in a complex occlusion scenario based on the acquired multidimensional environmental parameter data includes: Using the Grounding DINO visual localization model, the target object region in the scene corresponding to the semantic command is determined; Using the SAM segmentation model, pixel-level semantic segmentation of the target object is performed. Based on the calibrated depth and distance data, the segmented target region is mapped from the pixel coordinate system to the three-dimensional spatial coordinate system to extract the original spatial point cloud data of the target object. Based on the extracted original spatial point cloud data, the original target point cloud is subjected to dual optimization processing using the Diffusion Policy algorithm. Based on the optimized original target point cloud data, the spatial positioning information and pose parameters of the target object are calculated using a preset spatial coordinate calculation algorithm to determine the spatial positioning of the target object in complex occlusion scenarios.
4. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 3, characterized in that, The process of performing dual optimization on the original target point cloud using the Diffusion Policy algorithm includes: Acquire raw spatial point cloud data and perform point cloud denoising processing on the raw spatial point cloud data; By leveraging the massive prior knowledge of object shapes learned by the pre-trained Diffusion Policy model, we can infer and reconstruct the spatial point information of occluded and missing objects, and reconstruct a complete point cloud model of the target object.
5. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 4, characterized in that, Training the Diffusion Policy model includes: Obtain complete 3D point cloud samples containing various common household objects and partially occluded point cloud samples to construct a training dataset; Using partially occluded point cloud samples as network input and complete 3D point cloud samples as ground truth labels, the Diffusion Policy network is iteratively trained. After the network parameters converge, a pre-trained Diffusion Policy model is obtained.
6. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 5, characterized in that, The grasping and delivery unit includes: Data receiving and acquisition module: used to collect multi-dimensional sensing data of the robotic arm gripper in real time, and to fuse and summarize the collected multi-dimensional sensing data with the spatial positioning information and pose parameters of the target object output by the analysis and positioning unit; Data analysis and processing module: Based on the aggregated data, it uses a pre-trained grasping stability scoring network to calculate and select the optimal grasping strategy, which includes grasping posture and grasping force parameters. Adjustment module: Used to adjust the robotic arm's motion trajectory and gripper parameters according to the optimal grasping strategy, and control the robotic arm to complete the grasping of the target object; Delivery path optimization module: used to identify the user's human posture and hand movement data, determine the user's human-computer interaction intention, optimize the delivery path of the robotic arm, and realize safety protection in the delivery process of the elderly care robot.
7. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 6, characterized in that, The multidimensional sensing data includes real-time contact pressure and contact torque data of the gripper collected using a six-dimensional force sensor, and contact area and adhesion distribution data of the object collected using a flexible tactile epidermal array sensor.
8. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 7, characterized in that, The process of using a pre-trained grasping stability scoring network to calculate and select the optimal grasping strategy includes: When the target object is a rigid object, the robotic arm gripper is controlled to use a pinching strategy to grasp the target object. When the target object is a flexible and easily deformable object, the robotic arm gripper is controlled to use a gripping strategy to grasp the target object.
9. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 8, characterized in that, The user's body posture and hand movement data include the real-time position and movement trajectories of the user's hands, arms, and torso. After recognizing the user's body posture and hand movement data, the delivery path optimization module is further used for: By using an LSTM temporal intent classification network, temporal features of hand movement speed, displacement amplitude, and movement trend are extracted, and the extracted temporal features are compared with preset thresholds to determine the user's interaction intent in real time.
10. The 6D pose estimation and compliant delivery system for a multimodal large-scale elderly care robot according to claim 9, characterized in that, The process of determining the user's human-computer interaction intent and optimizing the robotic arm delivery path includes: When an external collision force is detected to be greater than a preset safety threshold, the robotic arm is controlled to reverse and stop abruptly, and the end torque of the robotic arm is released to remove the gripping force of the gripper and terminate the operation, thus achieving safety protection during the delivery process of the elderly care robot.