A multi-task intelligent coordination execution method and system based on health care accompanying robot

By collecting multimodal perception data to generate fused feature vectors and calculating task priorities and resource coordination indexes, the multi-task scheduling problem of health care accompanying robots in complex scenarios is solved, and efficient and reliable task execution is achieved.

CN120494765BActive Publication Date: 2025-09-26XIAMEN QIUSHI INTELLIGENT NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510998832.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-26
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing health care accompanying robots find it difficult to respond to multiple needs simultaneously in complex scenarios, ignoring user emotional fluctuations, resource conflicts and equipment energy consumption, resulting in low multi-task scheduling efficiency, resource waste and poor system reliability.

Method used

Collect multimodal perception data (images, voice, vital signs), generate multimodal fusion feature vectors, calculate task priority and resource coordination index through feature decoupling, event extraction and clustering, use optimization algorithms to generate planning paths and action sequences, and monitor and trigger rescheduling in real time.

Benefits of technology

It achieves accurate capture of the user's comprehensive status, improves the accuracy of situation-emotion analysis, the scientific nature of task scheduling and resource utilization, and enhances the adaptability and reliability of the system in a dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494765B_ABST
    Figure CN120494765B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of task coordination and execution, and discloses a multi-task intelligent coordination and execution method and system based on a health care accompanying robot, comprising the following steps: Step 1, collecting multimodal perception data and integrating it to generate a multimodal fusion feature vector; Step 2, identifying the target event object and clustering it to obtain a context-emotion cluster; Step 3, performing task mapping on the target event object, extracting the time urgency and historical success rate, and weighted fusion to obtain a task priority index; Step 4, obtaining the robot resource usage of each task and calculating the inter-task resource coordination index; Step 5, determining the optimal task subset through an optimization algorithm, and generating a planned path and action sequence instructions; Step 6, triggering rescheduling when the deviation between the planned path and the actual execution path exceeds a preset deviation threshold. The present invention realizes efficient, accurate, and reliable multi-task execution of the robot in a dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of task coordination and execution, and specifically relates to a multi-task intelligent coordination and execution method and system based on a health care accompanying robot. Background Art

[0002] With the increasing demand for home healthcare, healthcare companion robots are gradually becoming an important direction for smart health services. Currently, most companion robots can only independently perform single tasks, such as health monitoring, walking assistance, or social interaction, and are unable to respond to multiple needs simultaneously in complex scenarios. At the same time, traditional task scheduling is mostly based on first-come, first-served or shortest path principles, ignoring multi-dimensional factors such as user mood swings, resource conflicts, and equipment energy consumption, resulting in slow response, waste of resources, and even service interruptions. In addition, existing technologies often rely on a single perception mode and are unable to integrate and analyze information such as vision, voice, and vital signs. They also lack a mechanism to quantify emotional states, making it difficult to provide timely and humane comfort when users are anxious or in pain. Summary of the Invention

[0003] The present invention provides a multi-task intelligent coordinated execution method and system based on a health care companion robot, which solves the technical problems in related technologies that it is difficult to integrate multimodal perception data while taking into account the user's emotional state, resource constraints and energy consumption management, resulting in low multi-task scheduling efficiency, unreasonable parallel execution and poor system reliability.

[0004] The present invention provides a multi-task intelligent coordinated execution method based on a health care accompanying robot, comprising the following steps:

[0005] Step 1: Collect multimodal perception data, where the multimodal perception data includes image data, voice stream data, and vital sign data, and synchronously integrate the multimodal perception data according to timestamps to generate a multimodal fusion feature vector;

[0006] Among them, vital sign data include: heart rate, respiratory rate, blood pressure and skin galvanic response;

[0007] Step 2: Identify and filter the multimodal fusion feature vectors to obtain target event objects. For each target event object, extract the spatial location information and emotional state vector, and cluster them to obtain situation-emotion clusters.

[0008] Step 3: Map the target event objects contained in each situation-emotion cluster to tasks to obtain a task set. For each task, extract the time urgency and historical success rate, and perform weighted fusion according to the preset weight ratio to obtain the task priority index.

[0009] Step 4: Obtain the robot resource usage of each task, calculate the resource coordination index between tasks, and then select task pairs that can be executed in parallel through the threshold comparison method;

[0010] Step 5: Based on the task priority index and the task pairs that can be executed in parallel, the optimization algorithm is used to determine the optimal task subset that meets the constraints, and generate the planning path and action sequence instructions;

[0011] Step 6: Obtain the actual execution path, monitor the deviation between the planned path and the actual execution path in real time, and trigger rescheduling when the deviation exceeds a preset deviation threshold.

[0012] Furthermore, the specific steps of identifying and screening the multimodal fusion feature vector include:

[0013] Step 11: Input the multimodal fusion feature vector into the pre-trained feature decoupling network to separate image features, speech features, and life features;

[0014] Step 12: Use the target detection algorithm to extract events from the image features to obtain three types of events: person set, joint point coordinates, and action type;

[0015] Step 13: Use the BERT classifier to extract events from speech features to obtain semantic type events;

[0016] Step 14, calculating the three events of average heart rate, average blood pressure and number of hypertension according to the characteristics of the life body;

[0017] In step 15, the character set, joint coordinates, action type, semantic type, average heart rate, average blood pressure, and number of hypertension events are aligned by timestamp and merged according to spatiotemporal constraints to obtain the target event object.

[0018] Furthermore, the spatiotemporal constraints include:

[0019] The timestamps of any two events are within a first preset threshold;

[0020] The Euclidean distance between the spatial positions corresponding to any two events is within a second preset threshold; wherein, the spatial position is obtained in the following ways: using the coordinates of the joint points as the spatial position of the event corresponding to the image feature; calculating the spatial position of the event corresponding to the voice feature through the sound source localization algorithm; and using the position uploaded by the vital sign data as the spatial position of the event corresponding to the vital feature.

[0021] Furthermore, the spatial position vector and emotional state vector of each target event object are extracted and clustered. The specific steps include:

[0022] Step 21, extracting the spatial position vector according to the joint point coordinates of the target event object;

[0023] Step 22, extracting the emotional state vector based on the pre-trained sentiment analysis model;

[0024] Step 23, calculating the Euclidean distance between the spatial position vectors and the cosine similarity between the emotional state vectors of any two target event objects;

[0025] Step 24 , clustering is performed using a K-means clustering method based on the Euclidean distance and the cosine similarity to satisfy clustering constraints, thereby obtaining a context-emotion cluster, wherein the clustering constraints include: the Euclidean distance is less than a third preset threshold; and the cosine similarity is greater than a fourth preset threshold.

[0026] Furthermore, the sentiment analysis model includes: a multimodal fusion layer, a multimodal attention layer and an output layer;

[0027] Among them, the multimodal fusion layer includes: image branch, speech branch and vital sign branch;

[0028] The image branch is used to extract the first emotion vector according to image features;

[0029] The speech branch is used to extract the second emotion vector according to the speech features;

[0030] The vital signs branch is used to extract the third emotion vector based on the vital signs characteristics;

[0031] The multimodal attention layer is used to obtain fusion features based on the first emotion vector, the second emotion vector, and the third emotion vector using a multi-head self-attention mechanism;

[0032] The output layer is used to output the emotional state vector based on the fusion features, where the emotional state vector is a three-dimensional vector.

[0033] Furthermore, the specific steps of obtaining the task priority index include:

[0034] Step 31: For each task in the task set, obtain the difference between the current time and the preset execution time of the task, and calculate the time urgency by calculating the ratio of the difference to the preset time window;

[0035] Step 32: Obtain the historical number of successful executions and the total number of executions of the task, and calculate the historical success rate by the ratio of the two;

[0036] Step 33: The time urgency and the historical success rate are weighted and fused according to preset weights to obtain a task priority index.

[0037] Furthermore, the robot resources include: robotic arm, language, chassis, and power;

[0038] The specific steps to obtain the resource synergy index include:

[0039] Step 41: construct a task-resource matrix based on the amount of robot resources occupied by each task, where the matrix elements represent the amount of robot resources occupied by each task;

[0040] Step 42: Take the smaller value of the resource usage of any two tasks and add them together to obtain the total collaborative usage.

[0041] Step 43: Take the larger value of the resource usage of any two tasks and add them up to get the total demand;

[0042] Step 44 , calculating the ratio of the total collaborative occupancy to the total demand to obtain a resource collaborative index.

[0043] Furthermore, the specific steps of step 5 include:

[0044] Step 51, setting a binary decision variable for each task, wherein the binary decision variable is used to represent the execution status of the task;

[0045] Step 52: Based on the task priority indices and the pairs of tasks that can be executed in parallel, an optimization model is constructed with the goal of maximizing the total priority and satisfying the constraints, wherein the total priority represents the sum of the priority indices of each task, and the objective function of the optimization model is obtained by accumulating the product of the binary decision variables of all tasks and the task priority indices;

[0046] Step 53: solving the optimization model using a genetic algorithm to obtain an optimal task subset, wherein the optimal task subset represents a task subset consisting of binary decision variables of all tasks that maximizes the total priority;

[0047] Step 54: Generate planning path and action sequence instructions based on the optimal task subset.

[0048] Furthermore, the constraints include:

[0049] The total amount of resources occupied by the selected tasks must not exceed the maximum preset value of the resources;

[0050] The total energy consumption of the selected tasks must not exceed the difference between the robot's current remaining power and the preset safety power level;

[0051] Only pairs of tasks that can be executed in parallel are allowed to be selected at the same time.

[0052] The present invention provides a multi-task intelligent coordinated execution system based on a health care accompanying robot, comprising:

[0053] The data acquisition and fusion module is used to collect multimodal perception data, where the multimodal perception data includes image data, voice stream data, and vital sign data, and synchronously integrate the multimodal perception data according to timestamps to generate a multimodal fusion feature vector;

[0054] Among them, vital sign data include: heart rate, respiratory rate, blood pressure and skin galvanic response;

[0055] The event processing and clustering module is used to identify and filter the multimodal fusion feature vectors to obtain the target event objects, extract the spatial location information and emotional state vector of each target event object, and cluster them to obtain the situation-emotion cluster;

[0056] The priority calculation module is used to map the target event objects contained in each situation-emotion cluster to tasks, obtain a task set, extract the time urgency and historical success rate of each task, and perform weighted fusion according to the preset weight ratio to obtain the task priority index;

[0057] The resource collaboration analysis module is used to obtain the robot resource usage of each task, calculate the resource collaboration index between tasks, and then select task pairs that can be executed in parallel through the threshold comparison method;

[0058] The task scheduling and path planning module is used to determine the optimal task subset that meets the constraints based on the task priority indicators and the task pairs that can be executed in parallel through the optimization algorithm, and generate the planned path and action sequence instructions;

[0059] The rescheduling module is used to obtain the actual execution path, monitor the deviation between the planned path and the actual execution path in real time, and trigger rescheduling when the deviation exceeds a preset threshold.

[0060] The beneficial effects of the present invention are as follows: the present invention accurately captures the comprehensive status of the user by collecting multimodal perception data and integrating it to generate a fusion feature vector; improves the accuracy of context-emotion analysis through feature decoupling, event extraction and clustering; realizes scientific scheduling of tasks by combining task mapping with time urgency and historical success rate to calculate priority; improves resource utilization and the rationality of parallel execution of tasks by calculating the resource synergy index and comparing it with the threshold; determines the optimal task subset by constructing an optimization model and solving it with a genetic algorithm, and generates planning paths and action sequence instructions; improves the adaptability of the system in a dynamic environment and the reliability of task execution by monitoring path deviations in real time and triggering rescheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flow chart of a multi-task intelligent coordination execution method based on a health care accompanying robot of the present invention. DETAILED DESCRIPTION

[0062] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. In addition, features described with respect to some examples may also be combined in other examples.

[0063] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in one or more embodiments of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprising" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, but do not exclude other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0064] like Figure 1 As shown, a multi-task intelligent coordinated execution method based on a health care accompanying robot includes the following steps:

[0065] Step 1: Collect multimodal perception data, where the multimodal perception data includes image data, voice stream data, and vital sign data, and synchronously integrate the multimodal perception data according to timestamps to generate a multimodal fusion feature vector;

[0066] Among them, vital sign data include: heart rate, respiratory rate, blood pressure and skin galvanic response;

[0067] Step 2: Identify and filter the multimodal fusion feature vectors to obtain target event objects. For each target event object, extract the spatial location information and emotional state vector, and cluster them to obtain situation-emotion clusters.

[0068] Step 3: Map the target event objects contained in each situation-emotion cluster to tasks to obtain a task set. For each task, extract the time urgency and historical success rate, and perform weighted fusion according to the preset weight ratio to obtain the task priority index.

[0069] Step 4: Obtain the robot resource usage of each task, calculate the resource coordination index between tasks, and then select task pairs that can be executed in parallel through the threshold comparison method;

[0070] Step 5: Based on the task priority index and the task pairs that can be executed in parallel, the optimization algorithm is used to determine the optimal task subset that meets the constraints, and generate the planning path and action sequence instructions;

[0071] Step 6: Obtain the actual execution path, monitor the deviation between the planned path and the actual execution path in real time, and trigger rescheduling when the deviation exceeds a preset deviation threshold.

[0072] In one embodiment of the present invention, image data is collected by a depth camera to obtain visual features such as the user's position, posture, and facial expressions; voice stream data is collected by a microphone array to recognize the user's voice requests and analyze acoustic features such as intonation, speaking speed, and pitch for emotion recognition; physiological data such as the user's heart rate, respiratory rate, blood pressure, and skin electrode response are obtained by wearable devices to determine the degree of tension, abnormal health status, and changes in sleep rhythm; the above data are all timestamped and integrated using a unified time base to generate a multimodal fusion feature vector at the corresponding time point, which serves as input for subsequent event recognition.

[0073] In one embodiment of the present invention, the specific steps of identifying and screening the multimodal fusion feature vector include:

[0074] Step 11: Input the multimodal fusion feature vector into a pre-trained feature decoupling network to separate image features, speech features, and life features. Specifically, the feature decoupling network is built based on the Transformer architecture and achieves feature separation through an attention mechanism.

[0075] In step 12, an object detection algorithm is used to extract events from image features, resulting in three types of events: person set, joint coordinates, and action type. Specifically, the YOLOv8 algorithm is combined with the OpenPose human pose estimation model to obtain the task bounding box set in the image. OpenPose is used to extract the coordinates of 18 joints. The joint sequence of consecutive frames is analyzed using a spatiotemporal graph convolutional network to identify action types, including standing, sitting, and falling.

[0076] Step 13: Use the BERT classifier to extract events from the speech features to obtain semantic type events. Specifically, the speech features are converted into text embeddings and input into the BERT classifier. After encoding through a multi-layer Transformer, the semantic type is obtained through classification through a fully connected layer.

[0077] Step 14: Calculate the three events of average heart rate, average blood pressure, and number of hypertension events based on the vital characteristics; specifically, calculate the average heart rate and average blood pressure using a sliding window method, and count the number of hypertension events by comparing the systolic blood pressure at each time point with a preset blood pressure threshold, preferably, the preset blood pressure threshold is set to 140 mmHg;

[0078] In step 15, the character set, joint coordinates, action type, semantic type, average heart rate, average blood pressure, and number of hypertension events are aligned by timestamp and merged according to spatiotemporal constraints to obtain the target event object.

[0079] Among them, the spatiotemporal constraints include:

[0080] The timestamps of any two events are within a first preset threshold;

[0081] The Euclidean distance between the spatial positions corresponding to any two events is within a second preset threshold; wherein, the spatial position is obtained in the following ways: using the coordinates of the joint points as the spatial position of the event corresponding to the image feature; calculating the spatial position of the event corresponding to the voice feature through the sound source localization algorithm; and using the position uploaded by the vital sign data as the spatial position of the event corresponding to the vital feature.

[0082] Through the above steps, this embodiment achieves effective decoupling of multimodal features and event extraction, can more accurately capture the user's comprehensive status information, and provide a more reliable basis for the intelligent decision-making of the health care accompanying robot.

[0083] In one embodiment of the present invention, the spatial position vector and the emotional state vector are extracted for each target event object and clustered. The specific steps include:

[0084] Step 21, extracting a spatial position vector based on the joint point coordinates of the target event object; specifically, determining the spatial position vector by calculating the spatial centroid of all relevant node coordinates to represent the overall spatial position of the target event object;

[0085] Step 22, extracting the emotional state vector based on the pre-trained sentiment analysis model;

[0086] The sentiment analysis model includes: a multimodal fusion layer, a multimodal attention layer, and an output layer, which process different modal data respectively;

[0087] Among them, the multimodal fusion layer includes: image branch, speech branch and vital sign branch;

[0088] The image branch is used to extract the first emotion vector based on the image features. Specifically, the image branch captures the expressions, actions, etc. in the image features through a convolutional neural network and processes them to obtain the first emotion vector.

[0089] The speech branch is used to extract the second emotion vector based on speech features. Specifically, the speech features are converted into acoustic features through the Mel-spectrogram, which are then input into the Transformer-based speech encoder to output the second emotion vector.

[0090] The vital signs branch is used to extract the third emotion vector based on the characteristics of the life form. Specifically, the third emotion vector is extracted through a long short-term memory network combined with a threshold judgment mechanism.

[0091] The multimodal attention layer is used to process the first, second, and third emotion vectors using a multi-head self-attention mechanism to obtain fused features. The multi-head self-attention mechanism can adaptively assign weights to each modality in emotional expression.

[0092] The output layer is used to output the emotional state vector based on the fusion features, where the emotional state vector is a three-dimensional vector; the output layer is mapped to the three-dimensional emotional state space through a two-layer fully connected neural network, and each dimension is normalized to the range of -1 to 1 through the Tanh function.

[0093] The above sentiment analysis model achieves accurate capture of complex emotional states through hierarchical design and multimodal fusion mechanism, avoids the limitations of single-modal data, and provides a sentiment analysis basis for subsequent task priority indicators.

[0094] Step 23, calculate the Euclidean distance of the spatial position vectors and the cosine similarity of the emotional state vectors of any two target event objects; the Euclidean distance is used to measure the spatial position difference of the target event objects, and the smaller the distance, the closer the spatial positions; the cosine similarity is used to measure the similarity of emotional states, and the larger the cosine similarity, the more consistent the emotional tendency. By combining spatial distance and similarity, the spatiotemporal correlation and emotional consistency of the target event objects can be captured simultaneously.

[0095] Step 24 uses the K-means clustering method to cluster the context-emotion clusters that meet the clustering constraints, based on the Euclidean distance and cosine similarity. The clustering constraints include: the Euclidean distance is less than a third preset threshold; the cosine similarity is greater than a fourth preset threshold. This step, through the clustering constraints, prevents the incorrect clustering of target event objects with long spatial distances and conflicting emotions. It generates context-emotion clusters with high semantic consistency, transforming complex multimodal events into structured units, and significantly improving the accuracy and efficiency of subsequent task priority determination and resource scheduling.

[0096] In one embodiment of the present invention, task mapping is performed on the target event objects contained in each context-emotion cluster to achieve the transition from event understanding to task execution. Specifically, each context-emotion cluster represents a group of target event objects with similar spatial locations and emotional states. For example, a context-emotion cluster may contain target event objects such as image actions, speech semantics, and abnormal vital signs related to "an elderly person sitting on a sofa expressing discomfort." Task mapping involves matching and associating these context-emotion clusters with preset task templates, each of which corresponds to specific context-emotion features. For example, if a target event object in a context-emotion cluster indicates an elderly person's abnormal heart rate and speech semantics containing the word "uncomfortable," the cluster is mapped to the "Health Monitoring and Emergency Call" task template through a feature matching algorithm. If the context-emotion cluster reflects that the elderly person is in a static state and in a negative mood, it is mapped to the "Companion Chat" task template. Through this mapping mechanism, this embodiment converts abstract context-emotion information into specific, executable tasks, resulting in a task set. This task set serves as input for the calculation and scheduling of subsequent task priority indicators, providing a basis for the robot to execute intelligent decisions.

[0097] In one embodiment of the present invention, the specific steps of obtaining the task priority indicator include:

[0098] Step 31: For each task in the task set, obtain the difference between the current time and the preset execution time of the task, and calculate the time urgency by calculating the ratio of the difference to the preset time window. The calculation formula of the time urgency is: , p represents time urgency, which is used to measure the urgency of the task. Indicates the preset execution time of the task, Indicates the current time, Indicates the preset time window, max indicates the maximum value operation, when If it is less than or equal to 0, it means that the task has timed out. At this time, the time urgency is 0. Equal to the preset time window, p is 1, indicating that the task is in the most relaxed time state;

[0099] Step 32: Obtain the historical number of successful executions and the total number of executions of the task, and calculate the historical success rate by the ratio of the two. The historical success rate reflects the reliability of the task in the historical execution process and can effectively avoid priority misjudgment caused by differences in task execution difficulty.

[0100] Step 33 , the time urgency and the historical success rate are weighted and fused according to preset weights to obtain a task priority index, wherein a larger task priority index indicates a higher task priority.

[0101] The above steps provide a scientific basis for task prioritization by quantifying time urgency and historical execution data. Time urgency transforms task time constraints into comparable numerical indicators, ensuring that urgent tasks are prioritized. Historical success rate, based on past task performance, prioritizes reliable tasks to reduce the risk of failure. The weighted fusion of these two factors creates a task priority indicator that reflects both the time sensitivity of tasks and their execution stability, effectively guiding system resource allocation and task scheduling.

[0102] In one embodiment of the present invention, the robot resources include: a robotic arm, a language, a chassis, and power;

[0103] The specific steps to obtain the resource synergy index include:

[0104] Step 41: construct a task-resource matrix based on the robot resource occupation of each task. The matrix elements represent the robot resource occupation of each task. The occupation value ranges from 0 to 1. For example, when the robot arm occupation of a task is 0.8, it means that the task occupies 80% of the robot arm resources.

[0105] Step 42: Take the smaller value of the resource usage of any two tasks and add them together to obtain the total collaborative usage.

[0106] Step 43: Take the larger value of the resource usage of any two tasks and add them up to get the total demand;

[0107] Step 44: Calculate the ratio of the total collaborative occupancy to the total demand to obtain a resource collaborative index. The calculation formula of the resource collaborative index is: , represents the resource coordination index between the j-th task and the k-th task, represents the element in the jth row and mth column of the task-resource matrix, i.e., the amount of the mth resource occupied by the jth task. represents the element in the kth row and mth column of the task-resource matrix, min represents the minimum operation, max represents the maximum operation, j and k represent the index of the task, and m represents the index of the resource;

[0108] The resource synergy index ranges from 0 to 1. The closer the value is to 1, the more compatible the two tasks are in resource usage and the more suitable they are for parallel execution. The closer the value is to 0, the more serious the resource conflict is and the two tasks are not suitable for parallel execution. Parallel execution can effectively improve resource utilization. When the resource synergy index is greater than the preset synergy threshold, the two tasks are determined to be a task pair that can be executed in parallel.

[0109] In one embodiment of the present invention, the specific steps of step 5 include:

[0110] Step 51: Set a binary decision variable for each task, where the binary decision variable is used to represent the execution status of the task. A binary decision variable of 1 indicates that the task is selected for execution, and a binary decision variable of 0 indicates that the task is not selected for execution.

[0111] Step 52: Based on the task priority indices and the pairs of tasks that can be executed in parallel, an optimization model is constructed with the goal of maximizing the total priority and satisfying the constraints. The total priority represents the sum of the priority indices of each task, and the objective function of the optimization model is obtained by accumulating the product of the binary decision variables of all tasks and the task priority indices. That is, for tasks with a binary decision variable of 0, the task priority indices are not included in the objective function.

[0112] The constraints include:

[0113] The total amount of resources occupied by the selected tasks must not exceed the maximum preset value of the resources;

[0114] The total energy consumption of the selected tasks must not exceed the difference between the robot's current remaining power and the preset safety power level;

[0115] Only pairs of tasks that can be executed in parallel are allowed to be selected at the same time.

[0116] Step 53: solving the optimization model using a genetic algorithm to obtain an optimal task subset, wherein the optimal task subset represents a task subset consisting of binary decision variables of all tasks that maximizes the total priority;

[0117] Specifically, the individuals of the genetic algorithm are task subsets composed of binary decision variables of all tasks, representing a task execution plan. The fitness value of each individual is calculated through the objective function, and new individuals are generated through selection, crossover and mutation. After repeating N times, the individual with the highest fitness value is selected as the optimal task subset.

[0118] Step 54, generate a planned path and action sequence instructions based on the optimal task subset; specifically, for tasks involving movement, use the Dijkstra algorithm, combined with the robot's working environment map, to calculate the shortest path from the current position to the task execution position as the planned path, and convert the planned path into motion parameters of the robot joints according to the task requirements to generate continuous action sequence instructions.

[0119] In one embodiment of the present invention, the robot may be offset due to dynamic obstacles or sensor positioning errors when actually performing a task. The present invention evaluates the execution status by continuously calculating the deviation between the planned path and the actual execution path; the deviation calculation adopts dual measurements of spatial distance and direction angle, the spatial distance measures the position offset by the Euclidean distance between the actual path point and the corresponding point of the planned path, and the direction angle evaluates the degree of path deviation by comparing the angle between the actual travel direction and the planned path direction. When either of the two exceeds the preset deviation threshold, the current execution status is determined to be abnormal and the task is triggered to be rescheduled; the rescheduling will be combined with the remaining robot resources to re-evaluate the priority of the task, and regenerate the planned path based on the updated environmental map to ensure that the robot can dynamically adapt to environmental changes and resume task execution.

[0120] This embodiment also provides a multi-task intelligent coordinated execution system based on a health care accompanying robot, including:

[0121] The data acquisition and fusion module is used to collect multimodal perception data, where the multimodal perception data includes image data, voice stream data, and vital sign data, and synchronously integrate the multimodal perception data according to timestamps to generate a multimodal fusion feature vector;

[0122] Among them, vital sign data include: heart rate, respiratory rate, blood pressure and skin galvanic response;

[0123] The event processing and clustering module is used to identify and filter the multimodal fusion feature vectors to obtain the target event objects, extract the spatial location information and emotional state vector of each target event object, and cluster them to obtain the situation-emotion cluster;

[0124] The priority calculation module is used to map the target event objects contained in each situation-emotion cluster to tasks, obtain a task set, extract the time urgency and historical success rate of each task, and perform weighted fusion according to the preset weight ratio to obtain the task priority index;

[0125] The resource collaboration analysis module is used to obtain the robot resource usage of each task, calculate the resource collaboration index between tasks, and then select task pairs that can be executed in parallel through the threshold comparison method;

[0126] The task scheduling and path planning module is used to determine the optimal task subset that meets the constraints based on the task priority indicators and the task pairs that can be executed in parallel through the optimization algorithm, and generate the planned path and action sequence instructions;

[0127] The rescheduling module is used to obtain the actual execution path, monitor the deviation between the planned path and the actual execution path in real time, and trigger rescheduling when the deviation exceeds a preset threshold.

[0128] It should be noted that the intervals and thresholds are set for ease of comparison. The threshold size depends on the amount of sample data and the cardinality set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless numerical calculations. These formulas are derived from software simulations of the most recent real-world conditions using large amounts of data. The preset parameters in these formulas are set by those skilled in the art based on actual conditions.

[0129] The above describes the embodiments of the present invention, but the present invention is not limited to the above specific implementation methods. The above specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make many forms based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A multi-task intelligent coordination execution method based on a health care accompanying robot, characterized in that: The following steps are involved: Step 1: Collect multimodal perception data, where the multimodal perception data includes image data, voice stream data, and vital sign data, and synchronously integrate the multimodal perception data according to timestamps to generate a multimodal fusion feature vector; Among them, vital sign data include: heart rate, respiratory rate, blood pressure and skin galvanic response; Step 2: Identify and filter the multimodal fusion feature vectors to obtain target event objects. For each target event object, extract the spatial location information and emotional state vector, and cluster them to obtain situation-emotion clusters. Step 3: Map the target event objects contained in each situation-emotion cluster to tasks to obtain a task set. For each task, extract the time urgency and historical success rate, and perform weighted fusion according to the preset weight ratio to obtain the task priority index. Step 4: Obtain the robot resource usage of each task, calculate the resource coordination index between tasks, and then select task pairs that can be executed in parallel through the threshold comparison method; Step 5: Based on the task priority index and the task pairs that can be executed in parallel, the optimization algorithm is used to determine the optimal task subset that meets the constraints, and generate the planning path and action sequence instructions; Step 6: Obtain the actual execution path, monitor the deviation between the planned path and the actual execution path in real time, and trigger rescheduling when the deviation exceeds a preset deviation threshold.

2. A multi-task intelligent coordination execution method based on a health care accompanying robot according to claim 1, characterized in that: The specific steps for identifying and screening multimodal fusion feature vectors include: Step 11: Input the multimodal fusion feature vector into the pre-trained feature decoupling network to separate image features, speech features, and life features; Step 12: Use the target detection algorithm to extract events from the image features to obtain three types of events: person set, joint point coordinates, and action type; Step 13: Use the BERT classifier to extract events from speech features to obtain semantic type events; Step 14, calculating the three events of average heart rate, average blood pressure and number of hypertension according to the characteristics of the life body; In step 15, the character set, joint coordinates, action type, semantic type, average heart rate, average blood pressure, and number of hypertension events are aligned by timestamp and merged according to spatiotemporal constraints to obtain the target event object.

3. The multi-task intelligent coordination execution method based on the health care accompanying robot according to claim 2 is characterized in that: The spatiotemporal constraints include: The timestamps of any two events are within a first preset threshold; The Euclidean distance between the spatial positions corresponding to any two events is within a second preset threshold; wherein, the spatial position is obtained in the following ways: using the coordinates of the joint points as the spatial position of the event corresponding to the image feature; calculating the spatial position of the event corresponding to the voice feature through the sound source localization algorithm; and using the position uploaded by the vital sign data as the spatial position of the event corresponding to the vital feature.

4. The multi-task intelligent coordination execution method based on the health care accompanying robot according to claim 2 is characterized in that: Extract the spatial position vector and emotional state vector of each target event object and perform clustering. The specific steps include: Step 21, extracting the spatial position vector according to the joint point coordinates of the target event object; Step 22, extracting the emotional state vector based on the pre-trained sentiment analysis model; Step 23, calculating the Euclidean distance between the spatial position vectors and the cosine similarity between the emotional state vectors of any two target event objects; Step 24 , clustering is performed using a K-means clustering method based on the Euclidean distance and the cosine similarity to satisfy clustering constraints, thereby obtaining a context-emotion cluster, wherein the clustering constraints include: the Euclidean distance is less than a third preset threshold; and the cosine similarity is greater than a fourth preset threshold.

5. The multi-task intelligent coordination execution method based on the health care accompanying robot according to claim 4 is characterized in that: The sentiment analysis model includes: a multimodal fusion layer, a multimodal attention layer and an output layer; Among them, the multimodal fusion layer includes: image branch, speech branch and vital sign branch; The image branch is used to extract the first emotion vector according to image features; The speech branch is used to extract the second emotion vector according to the speech features; The vital signs branch is used to extract the third emotion vector based on the vital signs characteristics; The multimodal attention layer is used to obtain fusion features based on the first emotion vector, the second emotion vector, and the third emotion vector using a multi-head self-attention mechanism; The output layer is used to output the emotional state vector based on the fusion features, where the emotional state vector is a three-dimensional vector.

6. The multi-task intelligent coordination execution method based on a health care accompanying robot according to claim 1 is characterized in that: The specific steps of obtaining the task priority index include: Step 31: For each task in the task set, obtain the difference between the current time and the preset execution time of the task, and calculate the time urgency by calculating the ratio of the difference to the preset time window; Step 32: Obtain the historical number of successful executions and the total number of executions of the task, and calculate the historical success rate by the ratio of the two; Step 33: The time urgency and the historical success rate are weighted and fused according to preset weights to obtain a task priority index.

7. The multi-task intelligent coordination execution method based on a health care accompanying robot according to claim 1 is characterized in that: The robot resources include: robotic arm, language, chassis, and power; The specific steps to obtain the resource synergy index include: Step 41: construct a task-resource matrix based on the amount of robot resources occupied by each task, where the matrix elements represent the amount of robot resources occupied by each task; Step 42: Take the smaller value of the resource usage of any two tasks and add them together to obtain the total collaborative usage. Step 43: Take the larger value of the resource usage of any two tasks and add them up to get the total demand; Step 44 , calculating the ratio of the total collaborative occupancy to the total demand to obtain a resource collaborative index.

8. The multi-task intelligent coordination execution method based on a health care accompanying robot according to claim 1 is characterized in that: The specific steps of step 5 include: Step 51, setting a binary decision variable for each task, wherein the binary decision variable is used to represent the execution status of the task; Step 52: Based on the task priority indices and the pairs of tasks that can be executed in parallel, an optimization model is constructed with the goal of maximizing the total priority and satisfying the constraints, wherein the total priority represents the sum of the priority indices of each task, and the objective function of the optimization model is obtained by accumulating the product of the binary decision variables of all tasks and the task priority indices; Step 53: solving the optimization model using a genetic algorithm to obtain an optimal task subset, wherein the optimal task subset represents a task subset consisting of binary decision variables of all tasks that maximizes the total priority; Step 54: Generate planning path and action sequence instructions based on the optimal task subset.

9. The multi-task intelligent coordination execution method based on a health care accompanying robot according to claim 1 is characterized in that: The constraints include: The total amount of resources occupied by the selected tasks must not exceed the maximum preset value of the resources; The total energy consumption of the selected tasks must not exceed the difference between the robot's current remaining power and the preset safety power level; Only pairs of tasks that can be executed in parallel are allowed to be selected at the same time.

10. A multi-task intelligent coordination execution system based on a health care accompanying robot, characterized in that: A multi-task intelligent coordinated execution method based on a health care accompanying robot as described in any one of claims 1 to 9 is adopted, comprising: The data acquisition and fusion module is used to collect multimodal perception data, where the multimodal perception data includes image data, voice stream data, and vital sign data, and synchronously integrate the multimodal perception data according to timestamps to generate a multimodal fusion feature vector; Among them, vital sign data include: heart rate, respiratory rate, blood pressure and skin galvanic response; The event processing and clustering module is used to identify and filter the multimodal fusion feature vectors to obtain the target event objects, extract the spatial location information and emotional state vector of each target event object, and cluster them to obtain the situation-emotion cluster; The priority calculation module is used to map the target event objects contained in each situation-emotion cluster to tasks, obtain a task set, extract the time urgency and historical success rate of each task, and perform weighted fusion according to the preset weight ratio to obtain the task priority index; The resource collaboration analysis module is used to obtain the robot resource usage of each task, calculate the resource collaboration index between tasks, and then select task pairs that can be executed in parallel through the threshold comparison method; The task scheduling and path planning module is used to determine the optimal task subset that meets the constraints based on the task priority indicators and the task pairs that can be executed in parallel through the optimization algorithm, and generate the planned path and action sequence instructions; The rescheduling module is used to obtain the actual execution path, monitor the deviation between the planned path and the actual execution path in real time, and trigger rescheduling when the deviation exceeds a preset threshold.

Citation Information

Patent Citations

  • A multi-task collaborative identification method and system

    CN109947954A

  • Quadruped robot form optimization method based on Monte Carlo tree search space migration

    CN119250445A