A multi-robot collaborative operation method and system for elderly home care
By building a large multimodal model and virtual simulation platform for training, the problems of insufficient collaboration and safety of elderly care robots have been solved, and efficient, safe and personalized care services have been achieved through collaborative operations of multiple robots to meet the diverse needs of the elderly.
Patent Information
- Application Number
- CN202411317755.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Existing elderly care robots lack the ability and safety of multi-machine collaborative operations, lack medical expertise, are unable to provide personalized and efficient care services, and are unable to meet the diverse care needs of the elderly.
Build a large multimodal model, including motion planning, health monitoring and voice interaction models, train multi-robot collaborative work through a virtual simulation platform, combine motion planning with cost functions to evaluate, and use variational autoencoders to expand health data sets to achieve multi-robot collaborative work and health monitoring.
It realizes the safety, collaboration and precision of multi-robot collaborative operation, provides 24-hour uninterrupted multi-dimensional nursing services, can identify the health status of the elderly and respond to emergencies in a timely manner, and improve nursing efficiency and safety.
Smart Images

Figure CN119472635B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of nursing robots, and more specifically, relates to a multi-robot collaborative operation method and system for elderly home care. Background Art
[0002] With the global aging population intensifying, the need for home care for the elderly is becoming increasingly urgent, necessitating innovative solutions to improve both the quality and efficiency of care. Traditional care models, constrained by high labor costs and a shortage of specialized nursing talent, struggle to meet the high-quality, personalized care needs of the rapidly growing elderly population. Especially in this aging population, the elderly's needs for physical and mental health, daily care, and emergency medical response are becoming increasingly complex and diverse, placing higher demands on the speed, accuracy, and versatility of care services.
[0003] Currently, elderly care robots face numerous challenges. First, they have limitations in multi-robot collaborative operation capabilities and safety, making it difficult to effectively address the diverse care needs of the elderly and handle complex care situations. Furthermore, existing care robots often lack sufficient medical expertise and are unable to accurately and reliably assess the health status of the elderly. Finally, elderly care robots have relatively limited functions, making them unable to provide more precise and personalized services to the elderly. Summary of the Invention
[0004] In response to the above-mentioned defects or improvement needs of the prior art, the present invention provides a multi-robot collaborative operation method and system for home care of the elderly, with the aim of providing personalized, efficient, comprehensive, safe and intelligent care services for the elderly.
[0005] To achieve the above objectives, according to one aspect of the present invention, a multi-robot collaborative operation method for elderly home care is proposed, which includes a model construction phase and a model application phase, wherein:
[0006] The model building phase includes:
[0007] A virtual simulation platform is built, which includes a simulated robot, an elderly person, and a home environment. A large-scale motion planning model built based on PaLM-E is deployed on the simulated robot. The large-scale motion planning model is used to obtain unit motion plans for each robot based on multimodal information and nursing task information. The multimodal information includes visual perception information, tactile perception information, sound perception information, and temperature perception information.
[0008] Based on the virtual simulation platform, multi-robot collaborative operation training is carried out to train the motion planning model and obtain a trained motion planning model;
[0009] The model application phase includes:
[0010] Obtain multimodal information and nursing task information about the elderly and input it into the trained motion planning model to obtain the unit motion planning of each robot, thereby realizing home care for the elderly with multi-robot collaborative operation.
[0011] As a further preferred embodiment, multi-robot collaborative operation training is performed based on a virtual simulation platform, so that when the motion planning model is trained, the motion planning result is evaluated by a cost function, and the training of the motion planning model is completed when the cost function value is lower than a preset threshold;
[0012] The cost function J total The calculation formula is as follows:
[0013]
[0014] Where N is the number of robots, a i 、a j are the action sequences of robot i and robot j respectively;
[0015] J safe is the security cost, which is calculated as:
[0016]
[0017] Among them, d ij is the minimum distance between robot i and robot j, λ v is the speed penalty coefficient, ∥v i ∥ is the norm of the velocity of robot i, v max is the maximum speed of the robot;
[0018] J sim is the action similarity cost, which is calculated as:
[0019]
[0020] Among them, E is the action sequence of the robot in the expert data, α p ,α θ is the weight coefficient; Δp i ,Δp e are the position changes of robot i and robot e, Δθ i , Δθ e are the angle changes of robot i and robot e respectively;
[0021] J collaborate is the collaborative cost, which is calculated as follows:
[0022]
[0023] Among them, t (t i ,t j ) is the degree of overlap between the actions of robot i and robot j on the timeline, t i , t j are the action durations of robot i and robot j respectively.
[0024] As a further preferred method, before deploying the motion planning large model on the simulation robot, the motion planning large model is pre-trained and fine-tuned using a medical knowledge base;
[0025] The medical knowledge base includes books, journal articles, clinical guidelines, nursing manuals related to medical nursing and traditional Chinese medicine, text data in actual nursing records, and videos related to medical nursing, traditional Chinese medicine massage, and rehabilitation training.
[0026] As a further preferred embodiment, the operation training includes operation tasks preset in advance according to time and operation tasks provided according to actual conditions on site;
[0027] The work tasks preset in advance include scheduled meal preparation and delivery services, scheduled medicine preparation and delivery services, scheduled rehabilitation training services and acupoint massage tasks. The work tasks provided according to the actual situation on site include emergency assistance services, emergency cleaning services and dangerous goods sorting services.
[0028] As a further preferred embodiment, the model building stage further includes:
[0029] The variational autoencoder is trained using vital signs information from healthy and sub-healthy elderly people.
[0030] Based on the vital signs information of healthy and sub-healthy elderly people, the trained variational autoencoder generates reconstructed data labeled "healthy" and reconstructed data labeled "sub-healthy" respectively, thereby constructing a health monitoring dataset;
[0031] Train the text model based on the health monitoring dataset and use the trained text model as the health monitoring model.
[0032] The vital signs information includes heart rate, blood pressure, blood sugar, blood lipids and body temperature data;
[0033] The model application phase also includes:
[0034] The vital signs information of the elderly is collected and input into the health monitoring model to obtain health assessment results.
[0035] As a further preferred embodiment, the variational autoencoder includes an encoder and a decoder; when training the variational autoencoder, the method includes:
[0036] Based on vital sign information, a one-dimensional vector is obtained through the encoder, and then the one-dimensional vector is converted into a latent variable through linear mapping; the decoder is trained based on the latent variable so that the reconstructed data output by the decoder is as close as possible to the one-dimensional vector.
[0037] As a further preferred method, the loss function when training the decoder based on the latent variable is for:
[0038]
[0039] Among them, β is the weight coefficient, is the reconstruction error loss, is the KL divergence.
[0040] As a further preferred embodiment, the model building stage further includes:
[0041] The conversation information between the elderly and the caregiver is converted into text information, and the language model LLM is trained through the text information to obtain the voice interaction model;
[0042] The model application phase also includes:
[0043] After acquiring the elderly person's voice and converting it into text information, the text is input into the voice interaction model to obtain the dialogue result, which is then output in the form of voice through the microphone on the robot.
[0044] As a further preference, a direct preference optimization strategy is adopted to train the large language model LLM.
[0045] According to another aspect of the present invention, a multi-robot collaborative operation system for elderly home care is provided, comprising a processor for executing the multi-robot collaborative operation method for elderly home care.
[0046] In general, the above technical solutions conceived by the present invention have the following technical advantages compared with the existing technology:
[0047] 1. This invention leverages the planning and reasoning capabilities of a large multimodal model. By analyzing and reasoning the collected multimodal information, it can determine the optimal operational planning sequence for multiple robots and output the unit motion planning results for each robot. This system, while possessing self-analysis and reasoning capabilities, achieves the transition from multimodal sensory input to universal robot control instructions. This addresses the shortcomings of traditional collaborative planning involving multiple nursing robots, such as insufficient operational coherence, low automation, and poor generalization performance.
[0048] 2. Given the complexity and variability of nursing tasks, ensuring the safety, high collaboration, and accuracy of multi-robot nursing operations is challenging. To address this, the present invention constructs a training platform for multi-nursing robot motion simulation. A cost function that considers motion similarity, safety, and collaboration is applied to imitation learning tasks, enabling the robots to learn safe multi-robot collaboration strategies, ensuring the efficiency, safety, and accuracy of nursing tasks.
[0049] 3. The present invention proposes a health data set amplification method based on variational autoencoders, and uses it for training large health monitoring models. Specifically, accurately monitoring the health status of the elderly is a challenge, because the analysis and processing of multi-dimensional vital signs data is challenging; at the same time, deep learning monitoring models lack a large amount of training data, and the generalization performance of the model is insufficient. Therefore, the present invention combines variational autoencoders to learn the distribution characteristics of the vital signs data of the elderly, thereby obtaining a large amount of vital signs data sets, and uses them for training large multimodal health monitoring models, realizing elderly health monitoring with high generalization performance and high robustness.
[0050] 4. The multi-robot collaborative operation method proposed in this invention can not only effectively share the nursing load and ensure 24-hour uninterrupted care, but also provide multi-level and multi-dimensional services including but not limited to health monitoring, life assistance, emotional communication, and emergency assistance based on the specific needs of the elderly, greatly enriching the content and form of nursing services. More importantly, through deep learning, large model technology, and intelligent technology to achieve environmental adaptation, continuously optimize its service strategy and execution efficiency, achieve accurate identification and immediate response to the elderly's care needs, and achieve efficient collaboration. This can not only significantly improve the accuracy and safety of nursing care, but also provide faster and more professional assistance in emergency situations, thereby building a solid health defense line for the elderly. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flow chart of a multi-robot collaborative operation method for elderly home care according to an embodiment of the present invention;
[0052] Figure 2 This is a schematic diagram of a large health monitoring model according to an embodiment of the present invention;
[0053] Figure 3 Schematic diagram of a motion planning model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0055] The embodiment of the present invention provides a multi-robot collaborative operation method for elderly home care, such as Figure 1 As shown, the following steps are included:
[0056] The model building phase includes:
[0057] S1: Deploy microphones, physiological monitoring equipment, visual sensors, tactile sensors, and temperature sensors on multiple humanoid robots to collect sound information, vital sign information, visual information, tactile information, and temperature information respectively;
[0058] S2: Collect conversation information between caregivers and elderly people for natural language model training. Learn human preferences through direct preference optimization to obtain a large voice interaction model.
[0059] S3: Collect multimodal vital sign information of the elderly and construct a health monitoring dataset using a variational autoencoder for training a large health monitoring model;
[0060] S4: Build a virtual simulation training platform, fine-tune the multimodal motion planning model in the medical nursing knowledge base and the traditional Chinese medicine knowledge base, deploy the motion planning model on the simulation robot, and conduct multi-robot collaborative operation training based on the virtual simulation platform to train the motion planning model.
[0061] Furthermore, step S2 includes:
[0062] The sound information of the conversation between the caregiver and the elderly is collected and converted into text information to construct a dataset D. The natural language model LLM is trained with the dataset D to obtain a voice interaction model.
[0063] Specifically, based on the pre-trained natural language model LLM, a direct preference optimization strategy is adopted to obtain a large voice interaction model with human preferences. The direct preference optimization strategy is achieved by maximizing the objective function. The objective function J(π) is specifically expressed as follows:
[0064]
[0065] in, represents the expected value, (s, α1, α2) ~ D represents the state s and actions α1 and α2 sampled from the data set D, π is the current strategy, π ref is a known reference strategy, D is the collected dataset, s represents context information, a1 and a2 are two possible outcomes under state s respectively; According to the human preference in the data set, in the given data set D, under the condition of state s, the probability that action α1 is better than action α2, π(a1|s) represents the probability that strategy π selects action α1 in state s, and π(a2|s) represents the probability that strategy π selects action α2 in state s; D KL (π∥π ref ) is the KL divergence, which is used to ensure that the current policy is close to the reference policy, and λ is the regularization parameter.
[0066] By maximizing the above objective function, we can obtain a large voice interaction model that can better respond to human preference strategies.
[0067] Furthermore, in step S3, the multimodal vital signs information of the elderly includes heart rate, blood pressure, blood sugar, blood lipids, and body temperature, and the variational autoencoder includes an encoder and a decoder. First, the vital signs information of healthy elderly people and sub-healthy elderly people is collected to build a knowledge base, and the variational autoencoder is trained based on the vital signs information in the knowledge base. Then, a larger health monitoring dataset is constructed from the existing knowledge base through the trained variational autoencoder, such as Figure 2 As shown, the specific steps include:
[0068] S31: The encoder represents the input vital signs information, including heart rate, blood pressure, blood sugar, blood lipids, and body temperature data, using a one-dimensional vector x, where x = (x1, x2, x3, x4, x5);
[0069] S32: Convert the one-dimensional vector x into a latent variable z by linear mapping, which conforms to the standard normal distribution
[0070] S33: The decoder is a neural network model that converts the latent variable z into a reconstructed data that is close to the original data x.
[0071] Train the neural network model by minimizing the loss function;
[0072] The loss function for neural network model training is:
[0073]
[0074] Among them, β is the weight coefficient, is the reconstruction error loss, is the KL divergence, which is calculated as follows:
[0075]
[0076] The reconstruction loss ensures that the neural network can generate data that is similar to the original data, and the KL divergence makes the potential distribution q of the output φ (z|x) is close to a standard normal distribution.
[0077] S34: Based on the vital signs information of healthy elderly people in the knowledge base, a large amount of reconstructed data is generated through the trained variational autoencoder, and its label is "healthy"; based on the vital signs data of sub-healthy elderly people in the knowledge base, a large amount of reconstructed data is generated through the trained variational autoencoder, and its label is "sub-healthy"; the reconstructed data labeled "healthy" and the reconstructed data labeled "sub-healthy" constitute the health monitoring dataset.
[0078] S35: The text big model is trained using the health monitoring data set, and the trained text big model is used as the health monitoring big model. The output of the health monitoring big model is "healthy" or "sub-healthy" status.
[0079] Further, such as Figure 3 As shown, step S4 includes:
[0080] S41: Build a virtual simulation platform, which includes simulated robots, elderly people, and home environments such as beds, chairs, and dining tables to simulate real-world physics and object interactions for operational training of nursing robots.
[0081] S42: Build a large motion planning model based on PaLM-E and perform pre-training and fine-tuning on a medical knowledge base to enhance the capabilities of elderly care and healthcare services.
[0082] The medical knowledge base includes the medical nursing knowledge base and the traditional Chinese medicine knowledge base. The medical nursing knowledge base and the traditional Chinese medicine knowledge base include text data collected from books, journal articles, clinical guidelines, nursing manuals and actual nursing records related to medical nursing and traditional Chinese medicine, as well as videos related to medical nursing, traditional Chinese medicine massage, and rehabilitation training.
[0083] S43: Deploy the pre-trained motion planning model to the simulated robot. The simulated operation process includes three stages: multimodal information input, operation planning, and end-effector execution.
[0084] Multimodal information input: Input multimodal information collected by sensors installed on multiple robots, including visual perception information, tactile perception information, sound perception information, and temperature perception information;
[0085] Operation planning stage: After combining multimodal information and pre-input nursing task information, the multimodal motion planning model performs task analysis and reasoning, and outputs the planned action results of each robot in the form of motion state, posture change, start and end time. Specifically, the unit motion plan (S, Δx, Δy, Δz, Δθ) of each robot and its manipulator is generated in text form. x ,Δθ y ,Δθ z ,T S ,T E ), where S is the current state of the end, such as motion or stillness; Δp = (Δx, Δy, Δz) is the position change of the robot actuator, Δθ = (Δθ x ,Δθ y ,Δθ z ) is the angle change of the robot actuator; T S ,T E are the start time and end time of the action respectively;
[0086] End effector execution: Based on the unit motion plan generated for each robot and its robotic arm, the robot's actuator part executes the motion planning results to realize multi-robot nursing service.
[0087] S44: Based on the virtual simulation platform, multi-robot collaborative operation training is carried out through imitation learning, thereby training the motion planning model.
[0088] There are two types of work training. One type is work tasks that are preset in advance, including scheduled meal preparation and delivery services, scheduled medicine preparation and delivery services, scheduled rehabilitation training services, and acupoint massage tasks. The other type is work tasks provided based on actual on-site conditions, including emergency assistance services, emergency cleaning services, and dangerous goods sorting services.
[0089] During the training of the action planning model, the cost function J total Evaluate the motion planning results to continuously train the collaborative capabilities of multiple robots until the cost function is lower than the preset threshold and the training ends.
[0090] The cost function includes safety cost, action similarity cost, and collaboration cost, as follows:
[0091]
[0092] Among them, N represents the number of robots, and the action sequence of each robot is Each action contains four pieces of information: state information si, position change information Δp, angle change information Δθ, and duration T;
[0093] J safeis the security cost, which is calculated as:
[0094]
[0095] Among them, d ij is the minimum distance between robot i and robot j, λ v is the speed penalty coefficient, ∥v i ∥ is the norm of the velocity of robot i, v max is the maximum speed of the robot.
[0096] J sim is the action similarity cost, which is calculated as:
[0097]
[0098] Among them, E represents the action sequence in the expert data, α p ,α θ are weight coefficients respectively; expert data refers to the robot's action results under human supervision.
[0099] J collaborate is the collaborative cost, which is calculated as follows:
[0100]
[0101] Among them, t (t i ,t j ) represents the degree of overlap of actions on the timeline.
[0102] The model application phase includes:
[0103] S5: After acquiring the elderly person's voice and converting it into text information, it is input into the voice interaction model to obtain the conversation result, and then output in the form of voice through the microphone;
[0104] The collected vital signs information of the elderly is input into the health monitoring model to obtain the health assessment results of the elderly, that is, "healthy" or "sub-healthy".
[0105] The multimodal information and nursing task information obtained by multiple robot sensors are input into the motion planning model to generate unit motion planning results for each robot structure, realizing home care for the elderly with multi-robot collaborative operation.
[0106] S6: Abnormal reporting function: When the elderly’s health condition is assessed to be abnormal or unsafe behavior is identified, the system triggers the alarm function.
[0107] For example, when the health assessment result of the elderly is detected to be in a sub-healthy state in step S5, the data is uploaded to the caregiver of the monitoring center and an alarm is issued to the caregiver.
[0108] This invention deeply integrates artificial intelligence and robotics technology to achieve multi-robot collaborative operation for home care of the elderly, which can effectively meet the diverse care needs of the elderly and handle complex care situations.
[0109] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-robot collaborative operation method for elderly home care, characterized in that: It includes the model building phase and the model application phase, in which: The model building phase includes: A virtual simulation platform is built, which includes a simulated robot, an elderly person, and a home environment. A large-scale motion planning model built based on PaLM-E is deployed on the simulated robot. The large-scale motion planning model is used to obtain unit motion plans for each robot based on multimodal information and nursing task information. The multimodal information includes visual perception information, tactile perception information, sound perception information, and temperature perception information. Based on the virtual simulation platform, multi-robot collaborative operation training is carried out to train the motion planning model and obtain a trained motion planning model; The model application phase includes: Obtain multimodal information and nursing task information about the elderly and input it into the trained motion planning model to obtain the unit motion planning of each robot, thereby realizing home care for the elderly with multi-robot collaborative operation.
2. The multi-robot collaborative operation method for elderly home care according to claim 1, characterized in that: Based on the virtual simulation platform, multi-robot collaborative operation training is carried out to train the motion planning model. The motion planning results are evaluated by the cost function. When the cost function value is lower than the preset threshold, the training of the motion planning model is completed. The cost function J total The calculation formula is as follows: Where N is the number of robots, a i 、a j are the action sequences of robot i and robot j respectively; J safe is the security cost, which is calculated as: Among them, d ij is the minimum distance between robot i and robot j, λ v is the speed penalty coefficient, ∥v i ∥ is the norm of the velocity of robot i, v max is the maximum speed of the robot; J sim is the action similarity cost, which is calculated as: Among them, E is the action sequence of the robot in the expert data, α p ,α θ is the weight coefficient; Δp i ,Δp e are the position changes of robot i and robot e, Δθ i , Δθ e are the angle changes of robot i and robot e respectively; J collaborate is the collaborative cost, which is calculated as follows: Among them, t (t i ,t j ) is the degree of overlap between the actions of robot i and robot j on the timeline, t i , t j are the action durations of robot i and robot j respectively.
3. The multi-robot collaborative operation method for elderly home care according to claim 1, characterized in that: Before deploying the large motion planning model on the simulation robot, the large motion planning model is pre-trained and fine-tuned using the medical knowledge base. The medical knowledge base includes books, journal articles, clinical guidelines, nursing manuals related to medical nursing and traditional Chinese medicine, text data in actual nursing records, and videos related to medical nursing, traditional Chinese medicine massage, and rehabilitation training.
4. The multi-robot collaborative operation method for elderly home care according to claim 1, characterized in that: The operation training includes operation tasks preset in advance according to time and operation tasks provided according to actual conditions on site; The work tasks preset in advance include scheduled meal preparation and delivery services, scheduled medicine preparation and delivery services, scheduled rehabilitation training services and acupoint massage tasks. The work tasks provided according to the actual situation on site include emergency assistance services, emergency cleaning services and dangerous goods sorting services.
5. The multi-robot collaborative operation method for elderly home care according to claim 1, characterized in that: The model building phase also includes: The variational autoencoder is trained using vital signs information from healthy and sub-healthy elderly people. Based on the vital signs of healthy and sub-healthy elderly people, the trained variational autoencoder generates reconstructed data labeled "healthy" and "sub-healthy" respectively, thereby constructing a health monitoring dataset; Train the text model based on the health monitoring dataset and use the trained text model as the health monitoring model. The vital signs information includes heart rate, blood pressure, blood sugar, blood lipids and body temperature data; The model application phase also includes: The vital signs information of the elderly is collected and input into the health monitoring model to obtain health assessment results.
6. The multi-robot collaborative operation method for elderly home care according to claim 5, characterized in that: The variational autoencoder includes an encoder and a decoder; When training a variational autoencoder, it involves: Based on the vital sign information, a one-dimensional vector is obtained through the encoder, and then the one-dimensional vector is converted into a latent variable through linear mapping; The decoder is trained based on the latent variables so that the reconstructed data output by the decoder is as close as possible to a one-dimensional vector.
7. The multi-robot collaborative operation method for elderly home care according to claim 6, characterized in that: Loss function when training the decoder based on latent variables for: Among them, β is the weight coefficient, is the reconstruction error loss, is the KL divergence.
8. The multi-robot collaborative operation method for elderly home care according to any one of claims 1 to 7, characterized in that: The model building phase also includes: The conversation information between the elderly and the caregiver is converted into text information, and the language model LLM is trained through the text information to obtain the voice interaction model; The model application phase also includes: After acquiring the elderly person's voice and converting it into text information, the text is input into the voice interaction model to obtain the dialogue result, which is then output in the form of voice through the microphone on the robot.
9. The multi-robot collaborative operation method for elderly home care according to claim 8, characterized in that: The large language model LLM is trained using a direct preference optimization strategy.
10. A multi-robot collaborative operation system for elderly home care, characterized by: It includes a processor, which is used to execute the multi-robot collaborative operation method for elderly home care as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Human-machine cooperative control system for preventing fall of elderly person and control method thereof
CN108415250A
Multi-agent collaborative anti-collision picking method based on digital twinning and reinforcement learning
CN114942633A