Robot control method and device, system and robot

By introducing task planning module and federated learning module into the robot control system, combining the server's model construction and inference module, the problems of robot controllability and flexibility are solved, and higher performance and reliability are achieved to meet the diverse needs of users.

CN118721204BActive Publication Date: 2025-05-06QINGDAO SMART LONGEVITY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410975689.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2025-05-06
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

The controllability of the robot is too limited and not flexible enough, resulting in low performance and reliability, which cannot better meet user needs.

Method used

It provides a robot's control system, including a task planning module and a federated learning module, and works in collaboration with the robot through the model construction and reasoning module in the server. The system can respond to target demand instructions, analyze and plan multiple target tasks, and dynamically adjust task execution plans to meet user needs.

Benefits of technology

It improves the controllability and flexibility of the robot, enhances performance and reliability, better meets users' diverse needs, and improves user comfort and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118721204B_ABST
    Figure CN118721204B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a control method and device for a robot, and a robot. The control system of the robot includes a robot and a server, wherein the robot includes a task planning module and a federated learning module, and the server includes a model building and reasoning module. The present application uses the global model parameters of the trained second embodied intelligent model on the server side, and sends the global model parameters to the federated learning module of the robot. After the robot side obtains the trained first embodied intelligent model based on the global model parameters, the target demand instruction is analyzed to obtain the target task execution plan corresponding to the target demand instruction. The target task execution plan will be dynamically adjusted according to the feedback of the service user and the current environment to better meet the user's needs and improve the user's comfort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a robot control method and device, a system and a robot. Background Art

[0002] A robot is an intelligent machine that can work semi-autonomously. It can perform tasks such as operation or movement through programming and automatic control. It can assist or even replace humans to complete part of the designated work, improve work efficiency and quality, serve human life, and expand or extend the scope of human activities and capabilities.

[0003] At present, robots are controlled through human-computer interaction systems, that is, the robot's body, head and neck are controlled through hearing or vision. As a result, the robot's controllability is too limited and not flexible enough, resulting in low performance and reliability of the robot, which cannot better meet user needs. Summary of the invention

[0004] The present disclosure provides a control method, device, electronic device and storage medium for a robot, the main purpose of which is to solve the problem that the controllability of the robot is too limited and not flexible enough, resulting in low performance and reliability of the robot.

[0005] According to a first aspect of the present disclosure, a control system of a robot is provided, the system comprising a robot and a server, wherein the robot comprises a task planning module and a federated learning module, and the server comprises a model building and reasoning module;

[0006] The task planning module is used to respond to the target demand instruction, call the trained first embodied intelligent model to analyze the target demand instruction, obtain the multiple target tasks corresponding to the target demand instruction, obtain the current environment information and the status information of the user served by the target demand instruction, and obtain multiple task execution plans composed of the multiple target tasks based on the current environment information and the status information, and determine the target task execution plan from the multiple task execution plans according to the priority order preset by the user and the target demand;

[0007] The federated learning module is configured to receive the global model parameters of the first embodied intelligence model sent by the server, and perform local training on the first embodied intelligence model and user historical data according to the global model parameters;

[0008] The model construction and reasoning module is used to input multimodal training data into the second embodied intelligence model for training after constructing the second embodied intelligence model, obtain corresponding global model parameters, and send the global model parameters to the federated learning module.

[0009] In some embodiments, the mission planning module includes: a target analysis unit, a mission planning unit, a mission execution unit, an environment perception and adjustment unit, and a decision-making unit;

[0010] The target analysis unit is used to analyze the target demand instruction in response to the target demand instruction to obtain the multiple target tasks corresponding to the target demand instruction;

[0011] The task planning unit is used to obtain the current environment information and the status information of the user served by the target demand instruction, and obtain multiple task execution plans consisting of the multiple target tasks based on the current environment information and the status information;

[0012] The decision unit is used to determine the target task execution plan from the multiple task execution plans according to the priority order and target requirements preset by the user;

[0013] The task execution unit is used to control the execution of each target task in the target task execution plan;

[0014] The environment perception and adjustment unit is used to dynamically adjust the target task execution plan based on the user's feedback information on the target task and / or the update of the target requirement until all target tasks in the target task execution plan are executed.

[0015] In some embodiments, the server includes: a multimodal data fusion module and an embodied intelligence modeling engine;

[0016] The multimodal data fusion module is used to perform feature extraction and fusion processing on the multimodal training data to obtain fused training data, and input the fused training data into the embodied intelligent modeling engine;

[0017] The embodied intelligence modeling engine is used to construct the second embodied intelligence model, and train the second embodied intelligence model based on the fused training data to obtain the global model parameters corresponding to the trained second embodied intelligence model, and send the global model parameters to the federated learning module.

[0018] According to a second aspect of the present disclosure, a robot control method is provided, the method being applied to a robot in the robot control system according to any one of the first aspects, comprising:

[0019] In response to the target demand instruction, calling the trained first embodied intelligent model to analyze the target demand instruction to obtain the multiple target tasks corresponding to the target demand instruction, wherein the global model parameters of the first embodied intelligent model are configured by the server and obtained by local training based on the global model parameters and user historical data;

[0020] Acquire current environment information and status information of a user served by the target demand instruction, and obtain multiple task execution plans consisting of the multiple target tasks based on the current environment information and the status information;

[0021] Determining a target task execution plan from the plurality of task execution plans according to a priority order and target requirements preset by a user;

[0022] During the process of executing each target task according to the target task execution plan, the target task execution plan is dynamically adjusted based on the user's feedback information on the target task and / or the update of the target requirement until all the target tasks in the target task execution plan are executed.

[0023] In some embodiments, the method further comprises:

[0024] The model parameters corresponding to the trained first embodied intelligent model are uploaded to the server, so that the server can perform federated learning according to the different model parameters of each robot, obtain updated global model parameters, and send the updated global model parameters to each robot.

[0025] In some embodiments, obtaining a plurality of task execution plans consisting of the plurality of target tasks based on the current environment information and the state information includes:

[0026] Planning execution steps for each target task based on the current environment information and the state information in the trained first embodied intelligent model;

[0027] Based on the current environment information, the execution strategy of each target task in each task execution plan and the serial and parallel order of each target task are adjusted to obtain the multiple task execution plans.

[0028] In some embodiments, determining the target task execution plan from the plurality of task execution plans according to the priority order and target requirements preset by the user includes:

[0029] Call the value assessment standard to evaluate each target task in each task execution plan and obtain the assessment result;

[0030] According to the preset priority order and the target requirement, the best evaluation result among the evaluation results is used as the target task execution plan.

[0031] In some embodiments, after uploading the model parameters corresponding to the trained first embodied intelligence model to the server, the method further includes:

[0032] Receiving updated global model parameters sent by the server;

[0033] The first embodied intelligence model is trained using the updated global model parameters, user historical data, and user historical feedback information to obtain the trained first embodied intelligence model.

[0034] According to a third aspect of the present disclosure, a robot control method is provided, the method being applied to a server in a control system of the robot according to any one of the first aspects, comprising:

[0035] Constructing a second embodied intelligence model;

[0036] Inputting the multimodal training data into the second embodied intelligent model for training, thereby obtaining a trained second embodied intelligent model, wherein the trained second embodied intelligent model corresponds to a global model parameter;

[0037] The global model parameters are sent to the robot so that the robot can locally train the second embodied intelligence model in the robot according to the global model parameters and the user's historical data.

[0038] In some embodiments, inputting the multimodal training data into the second embodied intelligence model for training to obtain a trained second embodied intelligence model comprises:

[0039] In the second embodied intelligent model, the multimodal training data is subjected to feature extraction and fusion processing to obtain fused training data; the multimodal training data includes at least two of camera training data, lidar training data, sound training data, text training data, temperature training data, smell training data and tactile training data;

[0040] Cross-modal reasoning is performed based on the fused training data to obtain the trained second embodied intelligence model.

[0041] In some embodiments, after inputting the multimodal training data into the second embodied intelligence model for training, the method further includes:

[0042] Receive new global model parameters sent by each robot;

[0043] Inputting the multimodal training data into the second embodied intelligence model for training includes:

[0044] Aggregate all new global model parameters based on a preset aggregation algorithm to obtain aggregated global model parameters;

[0045] The second embodied intelligent model is trained based on the aggregated global model parameters to obtain updated global model parameters, and the updated global model parameters are sent to all robots respectively.

[0046] According to a fourth aspect of the present disclosure, there is provided a control device for a robot, the device being configured in the robot in the control system of the robot according to any one of the first aspects, comprising:

[0047] an analysis unit, configured to respond to a target demand instruction and call a trained first embodied intelligent model to analyze the target demand instruction to obtain the multiple target tasks corresponding to the target demand instruction, wherein the global model parameters of the first embodied intelligent model are configured by the server and are obtained by local training based on the global model parameters and user historical data;

[0048] An acquisition unit, used to acquire current environment information and status information of users served by the target demand instruction;

[0049] A composition unit, configured to obtain a plurality of task execution plans composed of the plurality of target tasks based on the current environment information and the state information;

[0050] A determination unit, configured to determine a target task execution plan from the plurality of task execution plans according to a priority order and target requirements preset by a user;

[0051] The adjustment unit is used to dynamically adjust the target task execution plan during the execution of each target task in the target task execution plan based on the user's feedback information on the target task and / or the update of the target requirements until all target tasks in the target task execution plan are executed.

[0052] In some embodiments, the apparatus further comprises:

[0053] The uploading unit is used to upload the model parameters corresponding to the trained embodied intelligent model to the server, so that the server can perform federated learning according to the different model parameters of each robot, obtain updated global model parameters, and send the updated global model parameters to each robot.

[0054] In some embodiments, the constituent unit is further used for:

[0055] Planning execution steps for each target task based on the current environment information and the state information in the trained first embodied intelligent model;

[0056] Based on the current environment information, the execution strategy of each target task in each task execution plan and the serial and parallel order of each target task are adjusted to obtain the multiple task execution plans.

[0057] In some embodiments, the determining unit is further configured to:

[0058] Call the value assessment standard to evaluate each target task in each task execution plan and obtain the assessment result;

[0059] According to the preset priority order and the target requirement, the best evaluation result among the evaluation results is used as the target task execution plan.

[0060] In some embodiments, the apparatus further comprises:

[0061] a receiving unit, configured to receive updated global model parameters sent by the server after the uploading unit uploads the model parameters corresponding to the trained first embodied intelligence model to the server;

[0062] A training unit is used to train the first embodied intelligence model using the updated global model parameters, user historical data, and user historical feedback information to obtain the trained first embodied intelligence model.

[0063] According to a fifth aspect of the present disclosure, a robot control method is provided, the method being applied to a server in a control system of the robot according to any one of the first aspects, comprising:

[0064] A construction unit, used for constructing a second embodied intelligence model;

[0065] A training unit, configured to input the multimodal training data into the second embodied intelligent model for training, and obtain a trained second embodied intelligent model, wherein the trained second embodied intelligent model corresponds to a global model parameter;

[0066] A sending unit is used to send the global model parameters to the robot so that the robot can locally train the first embodied intelligent model in the robot according to the global model parameters and user historical data.

[0067] In some embodiments, the training unit is further used to:

[0068] In the embodied intelligence model, the multimodal training data is subjected to feature extraction and fusion processing to obtain fused training data; the multimodal training data includes at least two of camera training data, lidar training data, sound training data, text training data, temperature training data, smell training data and tactile training data;

[0069] Cross-modal reasoning is performed based on the fused training data to obtain the trained second embodied intelligence model.

[0070] In some embodiments, the apparatus further comprises:

[0071] A receiving unit, configured to receive new global model parameters respectively sent by each robot after the training unit inputs the multimodal training data into the second embodied intelligent model for training;

[0072] The training unit is further used to aggregate all new global model parameters based on a preset aggregation algorithm to obtain aggregated global model parameters;

[0073] The second embodied intelligent model is trained based on the aggregated global model parameters to obtain updated global model parameters, and the updated global model parameters are sent to all robots respectively.

[0074] According to a sixth aspect of the present disclosure, a robot is provided, comprising the robot control device as described in the fifth aspect.

[0075] According to a seventh aspect of the present disclosure, there is provided an electronic device, including:

[0076] at least one processor; and

[0077] a memory communicatively connected to the at least one processor; wherein,

[0078] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the second aspect or the third aspect.

[0079] According to an eighth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the second aspect or the third aspect.

[0080] According to a ninth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the second aspect or the third aspect.

[0081] The present disclosure provides a control method and device, system and robot for a robot, wherein the control system of the robot includes a robot and a server, wherein the robot includes a task planning module and a federated learning module, and the server includes a model building and reasoning module; the task planning module is used to respond to a target demand instruction, call a trained second embodied intelligent model to analyze the target demand instruction, obtain the multiple target tasks corresponding to the target demand instruction, obtain current environment information and status information of the user served by the target demand instruction, and obtain multiple task execution plans composed of the multiple target tasks based on the current environment information and the status information, and determine the target task execution plan from the multiple task execution plans according to the priority order and target demand preset by the user; the federated learning module is used to receive the global model parameters of the first embodied intelligent model sent by the server, and locally train the first embodied intelligent model and user historical data according to the global model parameters; the model building and reasoning module is used to input the multimodal training data into the second embodied intelligent model for training after constructing the second embodied intelligent model, obtain the corresponding global model parameters, and send the global model parameters to the federated learning module. The present application uses the global model parameters of the trained second embodied intelligent model on the server side and sends the global model parameters to the federated learning module of the robot. After the robot side obtains the trained first embodied intelligent model based on the global model parameters, the target demand instructions are analyzed to obtain the target task execution plan corresponding to the target demand instructions. The target task execution plan will be dynamically adjusted according to the feedback from service users and the current environment to better meet user needs and improve user comfort.

[0082] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0084] Figure 1 A schematic diagram of the structure of a control system of a robot provided by an embodiment of the present disclosure;

[0085] Figure 2 A schematic diagram of the structure of another robot control system provided by an embodiment of the present disclosure;

[0086] Figure 3 A schematic flow chart of a robot control method provided by an embodiment of the present disclosure;

[0087] Figure 4 A schematic flow chart of another robot control method provided by an embodiment of the present disclosure;

[0088] Figure 5 A schematic diagram of the structure of a robot control device provided by an embodiment of the present disclosure;

[0089] Figure 6 A schematic diagram of the structure of another robot control device provided by an embodiment of the present disclosure;

[0090] Figure 7 A schematic diagram of the structure of another robot control device provided by an embodiment of the present disclosure;

[0091] Figure 8 A schematic diagram of the structure of another robot control device provided by an embodiment of the present disclosure;

[0092] Fig. 9 A schematic block diagram of an exemplary electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION

[0093] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0094] The following describes the robot control method, device, and robot according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0095] Figure 1 A schematic diagram of the structure of a robot control system provided in an embodiment of the present disclosure.

[0096] like Figure 1 As shown, the system includes a robot 1 and a server 2, wherein the robot 1 includes a task planning module 11 and a federated learning module 12;

[0097] The task planning module 11 is used to respond to the target requirement instruction, call the trained embodied intelligent model to analyze the target requirement instruction, obtain the multiple target tasks corresponding to the target requirement instruction, obtain the current environment information and the status information of the user served by the target requirement instruction, and obtain multiple task execution plans composed of the multiple target tasks based on the current environment information and the status information, and determine the target task execution plan from the multiple task execution plans according to the priority order and target requirements preset by the user.

[0098] The task planning module 11 is used to understand the user's target needs and decide on the target task execution plan corresponding to the target needs. For example, when the user initiates a target demand instruction through voice, such as the user's bathing demand, the robot receives the target demand instruction, parses the target demand instruction, and obtains the corresponding multiple target tasks. The multiple target tasks corresponding to the bathing demand include but are not limited to: moving the user to a shower chair or bathtub, adjusting the water temperature, adjusting the bathroom temperature, applying shower gel, washing hair, washing the body, drying, and moving the user out of the shower chair or bathtub.

[0099] In some embodiments, when the target demand instruction is to cook, the robot receives the target demand instruction, parses the target demand instruction, and obtains the corresponding multiple target tasks. The multiple target tasks corresponding to the cooking demand include but are not limited to: moving to the refrigerator, taking vegetables, washing vegetables, cutting vegetables, checking gas and starting it, turning on the fire, cooking, seasoning, plating and other target tasks.

[0100] The process of parsing the target demand instruction into multiple target tasks can be performed by calling the trained first embodied intelligent model, which has a certain learning ability. Through learning, the corresponding relationship between the target demand instruction and the multiple target tasks can be obtained. The above is an exemplary explanation of the user bathing and cooking for ease of understanding. It should be clear that this explanation method is not intended to limit the target demand instruction to only bathing or cooking for the user. When there are differences in the target demand instructions, the corresponding target tasks are also different. The embodiment of this application does not limit the target tasks either.

[0101] After determining the target task, the task planning module 11 also needs to determine multiple task execution plans consisting of each target task based on the user's actual needs, current environment, and user status information, and determine the target task execution plan from the multiple task execution plans according to the user's preset priority order and target requirements.

[0102] For example, assume that the task execution plan includes task execution plan 1 and task execution plan 2. Task execution plan 1: move the user to the shower chair - adjust the water temperature - adjust the bathroom temperature - wash the hair - wash the body - dry the user - move the user out of the bathroom; Task execution plan 2: move the user to the shower chair - adjust the water temperature - adjust the bathroom temperature - wash the hair - apply shampoo - wash the body - apply shower gel - dry the user out of the bathroom.

[0103] According to the priority order preset by the user, one of the two task execution plans is selected as the target task execution plan. Assume that when the user sets not to use shower gel, the target task execution plan finally determined is task execution plan 1.

[0104] It should be noted that the target task execution plan described in the embodiment of the present application is not immutable after being determined. During the execution process, it is necessary to monitor the current ambient temperature and user feedback information in real time to dynamically adjust the target task execution plan. For example, when the robot detects that the fog in the bathroom is too much, or there is a lack of oxygen, the robot will add the ventilation function of the bathroom to the target task execution plan; or, when the user feedbacks that the water temperature is too low, the target task execution plan will add the task of raising the water temperature, etc. The above examples are only for the convenience of understanding, and are not intended to limit specific application scenarios and adjustment ranges.

[0105] The federated learning module 12 is used to receive the global model parameters of the first embodied intelligence model sent by the server, and perform local training on the first embodied intelligence model and user historical data according to the global model parameters;

[0106] In some embodiments, the federated learning module 12 is located in the robot 1, but in actual applications, the federated learning module 12 can also be located in the server 2. Because both the server 2 and the robot 1 have embodied intelligence models, although the embodiments of the present application adopt the first and second writing forms, their essence is embodied intelligence models. The difference is that the second embodied intelligence model in the server is a global embodied intelligence model, while the first embodied intelligence model in the robot is an embodied intelligence model customized in the robot according to user needs.

[0107] The server 2 includes a model building and reasoning module 21, which is used to input multimodal training data into the second embodied intelligent model for training after building the second embodied intelligent model, obtain corresponding global model parameters, and send the global model parameters to the federated learning module. The multimodal data described in the embodiment of the present application includes but is not limited to text data, temperature data, smell data and tactile data, sound data, image data, etc.

[0108] In some embodiments, please refer to Figure 2 , the task planning module 11 described in the embodiment of the present application includes: a target analysis unit 111, a task planning unit 112, a task execution unit 113, an environment perception and adjustment unit 114 and a decision unit 115;

[0109] The target analysis unit 111 is used to respond to the target demand instruction, analyze the target demand instruction, and obtain the multiple target tasks corresponding to the target demand instruction; the target analysis unit 111 understands the user's high-level goals, that is, recognizes the target demand instruction triggered by the user in any way, such as through voice, gesture, action, keyword or pre-set task plan, etc., to trigger the target demand instruction. And split the target demand into multiple target tasks, the split analysis process can be determined by the preset corresponding relationship, or by the result of learning and training, and the specific embodiment of the present application is not limited.

[0110] The task planning unit 112 is used to obtain the current environment information and the status information of the user served by the target demand instruction, and obtain multiple task execution plans composed of the multiple target tasks based on the current environment information and the status information; the task planning unit 112 needs to determine the execution order (task execution plan) of the target tasks analyzed by the target analysis unit 111. In addition to the execution order, the parallel or serial execution paths or steps of different target tasks will also be planned.

[0111] The decision unit 115 is used to determine the target task execution plan from the multiple task execution plans according to the priority order and target requirements preset by the user; in some embodiments, the decision unit 115 evaluates the pros and cons of the execution orders of each target task in different task execution plans based on the value function, and the pros and cons are determined by the evaluation score of the value function. The higher the evaluation score, the greater the probability of being selected as the target task execution plan. In the embodiment of the present application, the task execution plan with the highest evaluation score is determined as the target task execution plan.

[0112] It should be noted that the evaluation score of the value function comprehensively considers the user's priority order and target requirements, such as comfort, safety, task completion time, and other factors to make the best decision.

[0113] The task execution unit 113 is used to control the execution of each target task in the target task execution plan; control the robot to execute each target task, or the robot controls other devices to execute the target task. Specifically, the embodiment of the present application does not limit the execution process of the target task.

[0114] The environment perception and adjustment unit 114 is used to dynamically adjust the target task execution plan based on the user's feedback information on the target task and / or the update of the target demand until all the target tasks in the target task execution plan are executed. In the process of executing each target task in the target task execution plan, the system needs to continuously perceive the changes in the environment and make dynamic adjustments based on the user's real-time feedback to better meet the user's actual needs.

[0115] Please continue reading Figure 2 , the server 2 includes: a multimodal data fusion module 22 and an embodied intelligent modeling engine 23;

[0116] The multimodal data fusion module 22 is used to perform feature extraction and fusion processing on the multimodal training data to obtain fused training data, and input the fused training data into the embodied intelligent modeling engine;

[0117] The feature vectors obtained by different sensors are fused. For example, the visual features extracted from the image and the audio features extracted from the speech can be fused by concatenation, weighted averaging or other methods.

[0118] The integration specifically includes:

[0119] 1. Data preprocessing

[0120] Single-modal preprocessing: Perform separate preprocessing on the data of each modality, such as image denoising and normalization, text cleaning and word segmentation, audio denoising and feature extraction, etc.

[0121] Synchronization and alignment: If data from multiple modalities have a temporal or spatial correspondence (such as images and audio in a video), synchronization and alignment are required.

[0122] 2. Feature extraction

[0123] Unimodal feature extraction: Extract features from each modality that can represent the original data well.

[0124] Feature selection / dimensionality reduction (optional): If the feature dimension is too high, feature selection or dimensionality reduction may be required to reduce the amount of computation and improve model efficiency.

[0125] 3. Multimodal Fusion Strategy

[0126] Early fusion: After feature extraction, the features of different modalities are concatenated or fused immediately, and then input into a unified embodied intelligence model (second embodied intelligence model) for processing. This method is suitable for situations where the relationship between modalities is close and difficult to distinguish.

[0127] Mid-term fusion: Fusion is performed at the middle layer of the embodied intelligence model, allowing the embodied intelligence model to interact after partially learning the specific information of each modality.

[0128] Late fusion: Each modality is trained separately and then fused at the decision layer, such as through weighted voting, averaging, or more complex fusion algorithms. This method is highly flexible, but may ignore the interaction information between modalities.

[0129] The embodied intelligence modeling engine 23 is used to construct the second embodied intelligence model, and train the second embodied intelligence model based on the fused training data, obtain the global model parameters corresponding to the trained second embodied intelligence model, and send the global model parameters to the federated learning module.

[0130] The embodied intelligent modeling engine described in the embodiment of the present application is a comprehensive system for fusing, reasoning, planning and executing interrelated tasks of multimodal data. The following are the main subsystems of the embodied intelligent modeling engine:

[0131] 2.1 Data Collection Subsystem

[0132] This subsystem is responsible for collecting multimodal data from various sources such as sensors, databases, networks, etc.

[0133] 2.11 Visual Perception

[0134] Responsible for acquiring visual information such as images and point clouds from sensors such as cameras and lidars.

[0135] In particular, when used to achieve the user's bathing goal, the visual sensor should include infrared sensors, binocular cameras, ToF sensors, lidar, millimeter-wave radar, and structured light sensors.

[0136] Based on the visual sensor, a three-dimensional human body model of the user is constructed.

[0137] Based on the user's three-dimensional human body model, the motion trajectory of the bathing robot arm is planned in accordance with the robot arm operation principles.

[0138] The operation principles of the robotic arm should include:

[0139] Based on the user's three-dimensional human body model, a three-dimensional model of the robotic arm's suspended operation is constructed by extending outward in the normal direction of the body surface, with a certain distance as the robotic arm's suspension distance.

[0140] The robotic arm suspension distance can be set by the user or based on environmental information, such as water temperature, bath fluid used, etc.

[0141] The robot arm nozzle should run along the robot arm suspension running three-dimensional model, or run outside the robot arm suspension running three-dimensional model.

[0142] The robot arm nozzle must pass through specific position points on the three-dimensional model of the robot arm suspension operation. These points can be set by the user or by default, such as the forehead, armpit, crotch, etc.

[0143] 2.12 Sound Perception

[0144] Responsible for obtaining audio information from sensors such as microphones.

[0145] 2.13 Other Perceptions

[0146] It may also include other sensors such as touch, temperature, smell, etc. Data pre-processing subsystem: This subsystem is responsible for cleaning, standardizing and transforming the raw data to facilitate subsequent analysis and modeling.

[0147] 2.2 Feature extraction and fusion subsystem

[0148] This subsystem is responsible for fusing data from different sources to provide a unified view, involving technologies such as data alignment, data synchronization, and data fusion.

[0149] 2.21 Knowledge Base

[0150] Store knowledge about objects, environments, tasks, and rules. You can apply for but are not limited to graphs, ontologies, databases, etc. Use ontology technology to formally represent the attributes, relationships, and functions of objects.

[0151] 2.22 Cross-modal Knowledge Fusion

[0152] Integrate knowledge from different modalities into a unified knowledge representation and fuse data from different sensors to form a comprehensive environmental perception.

[0153] 2.3 Model building subsystem

[0154] This subsystem is responsible for using the fused data to create and update the embodied intelligence model, using but not limited to machine learning, deep learning, and reinforcement learning techniques.

[0155] 2.31 Natural Language Understanding

[0156] Convert text information into machine-understandable form.

[0157] 2.32 Cross-modal Reasoning

[0158] Based on the knowledge base and perception data, reasoning is performed, such as understanding the attributes, relationships, and functions of objects.

[0159] 2.33 Physical Simulation

[0160] Use a physical simulation environment to allow robots to experiment and learn in a virtual world. Through physical simulation, robots can simulate the movement and interaction of different objects, thereby gaining an intuitive understanding of the physical world.

[0161] The present application also provides a robot control method, which is applied to a robot in a robot control system, such as Figure 3 As shown, the method includes:

[0162] Step 101, in response to a target demand instruction, calling a trained first embodied intelligent model to analyze the target demand instruction to obtain the multiple target tasks corresponding to the target demand instruction, wherein the global model parameters of the first embodied intelligent model are configured by the server and are obtained by local training based on the global model parameters and user historical data;

[0163] This embodiment is described by taking the target demand instruction of a user taking a bath as an example, but it should be clear that this description method is not intended to limit the target demand of the robot to only taking a bath for the user.

[0164] For the correspondence between target requirement instructions and target tasks, please refer to Figure 1 The detailed description of the embodiments of the present application will not be repeated here one by one.

[0165] Step 102, obtaining current environment information and status information of the user served by the target demand instruction, and obtaining multiple task execution plans consisting of the multiple target tasks based on the current environment information and the status information;

[0166] Based on the environmental information and user status, a detailed execution path and steps are planned for each target task, and the execution strategy of the target task is adjusted in real time using the perceived environmental dynamics. In addition, the parallel and serial relationships between target tasks are coordinated to ensure the smooth progress of the overall bathing process.

[0167] In the trained first embodied intelligent model, execution steps are planned for each target task based on the current environmental information and the state information. The current environmental information mentioned here can be determined by, but not limited to, the multimodal data collected in the data collection subsystem of the above-mentioned embodiment 2.1.

[0168] Step 103, determining a target task execution plan from the plurality of task execution plans according to a priority order and target requirements preset by a user;

[0169] In the trained first embodied intelligent model, execution steps are planned for each target task based on the current environment information and the status information; the execution strategy of each target task in each task execution plan and the serial and parallel order of each target task are adjusted based on the current environment information to obtain the multiple task execution plans.

[0170] The process of parsing the target demand instruction into multiple target tasks can be executed by calling the trained first embodied intelligent model. The first embodied intelligent model has a certain learning ability, and the correspondence between the target demand instruction and the multiple target tasks can be obtained through learning.

[0171] Step 104, during the process of executing each target task in the target task execution plan, dynamically adjust the target task execution plan based on the user's feedback information on the target task and / or the update of the target requirement until all target tasks in the target task execution plan are executed.

[0172] In some embodiments, the value assessment standard is called to evaluate each target task in each task execution plan to obtain an assessment result, and the best assessment result among the assessment results is used as the target task execution plan according to the preset priority order and the target requirements.

[0173] According to the value function or value assessment standard, the advantages and disadvantages of different decisions (task execution plans) and the target task execution order are evaluated. During the specific implementation process, it is necessary to comprehensively consider factors such as user comfort, safety, task completion time, etc. to make the best decision and determine the task execution plan with the highest evaluation score as the target task execution plan.

[0174] During the execution of the target task execution plan, the decision-making effects are continuously evaluated and adjustments are made based on user feedback.

[0175] The present application uses the global model parameters of the trained second embodied intelligent model on the server side and sends the global model parameters to the federated learning module of the robot. After the robot side obtains the trained first embodied intelligent model based on the global model parameters, the target demand instructions are analyzed to obtain the target task execution plan corresponding to the target demand instructions. The target task execution plan will be dynamically adjusted according to the feedback from service users and the current environment to better meet user needs and improve user comfort.

[0176] In some embodiments, the method further includes: uploading the model parameters corresponding to the trained first embodied intelligent model to the server, so that the server performs federated learning according to the different model parameters of each robot, obtains updated global model parameters, and sends the updated global model parameters to each robot. Each intelligent bathing robot uses its own data (such as historical bathing data, user feedback, etc.) to train the first embodied intelligent model locally. The training process is performed locally on the robot to ensure data privacy.

[0177] After the training is completed, the robot sends the global model parameter updates (such as weight updates) of the first embodied intelligence model back to the server in encrypted form.

[0178] In some embodiments, after uploading the model parameters corresponding to the trained first embodied intelligence model to the server, the method further includes: receiving updated global model parameters sent by the server; and training the first embodied intelligence model using the updated global model parameters, user historical data, and user historical feedback information to obtain the trained first embodied intelligence model.

[0179] Each robot uploads the model parameter updates corresponding to the locally trained first embodied intelligence model to the server, and the server collects the updates of all participants. These updated model parameters are integrated using an aggregation algorithm (such as weighted average) to update the global model parameters. In some embodiments, the aggregation process aims to combine the contributions of all participants (robots) while protecting their data privacy.

[0180] The server sends the updated global model parameters to each robot again, and the intelligent bathing robot continues to train locally using the new global model parameters and uploads new updates. The whole process is iterated continuously until the global model (the second embodied intelligent model) converges or reaches the preset performance indicators.

[0181] The present application also provides a robot control method, which is applied to a server in a robot control system, such as Figure 4 As shown, including:

[0182] Step 201, constructing a second embodied intelligence model;

[0183] Collect fused data from the multimodal data fusion subsystem in real time, clean, convert and load the data for subsequent analysis and modeling, and build it using machine learning, deep learning and other technologies, such as user behavior models, environmental models, etc.

[0184] During this process, the embodied intelligence modeling engine will combine the global model parameters of the second embodied intelligence model provided by the federated learning module to improve the accuracy and generalization ability of the model, and perform cross-modal reasoning and prediction based on the fused data to guide the execution of the bathing task.

[0185] Step 202, inputting the multimodal training data into the second embodied intelligence model for training, obtaining a trained second embodied intelligence model, wherein the trained second embodied intelligence model corresponds to global model parameters;

[0186] For the training process of the second embodied intelligent model, please refer to the above-mentioned related instructions, and the embodiments of this application will not be repeated here.

[0187] Step 203: Send the global model parameters to the robot so that the robot can perform local training on the first embodied intelligent model in the robot according to the global model parameters and the user's historical data.

[0188] The present application uses the global model parameters of the trained second embodied intelligent model on the server side and sends the global model parameters to the federated learning module of the robot. After the robot side obtains the trained first embodied intelligent model based on the global model parameters, the target demand instructions are analyzed to obtain the target task execution plan corresponding to the target demand instructions. The target task execution plan will be dynamically adjusted according to the feedback from service users and the current environment to better meet user needs and improve user comfort.

[0189] In some embodiments, inputting the multimodal training data into the second embodied intelligence model for training to obtain a trained second embodied intelligence model comprises:

[0190] In the second embodied intelligent model, the multimodal training data is subjected to feature extraction and fusion processing to obtain fused training data; the multimodal training data includes at least two of camera training data, lidar training data, sound training data, text training data, temperature training data, smell training data and tactile training data;

[0191] The source of the multimodal training data is obtained through the data collection subsystem in the above embodiment 2.1, or the format and content of the multimodal training data are the same as the multimodal data in the data collection subsystem in 2.1.

[0192] Cross-modal reasoning is performed based on the fused training data to obtain the trained second embodied intelligence model.

[0193] In some embodiments, after inputting the multimodal training data into the second embodied intelligence model for training, the method further includes:

[0194] Receive new global model parameters sent by each robot;

[0195] Inputting the multimodal training data into the second embodied intelligence model for training includes:

[0196] Aggregate all new global model parameters based on a preset aggregation algorithm to obtain aggregated global model parameters;

[0197] The second embodied intelligent model is trained based on the aggregated global model parameters to obtain updated global model parameters, and the updated global model parameters are sent to all robots respectively.

[0198] After the bathing task is completed, the system will provide the user with feedback on the task completion status and important information about the bathing process, such as water temperature, bathing duration, etc. At the same time, the system will automatically detect and confirm the user's safety status to ensure that no accidents occur during the bathing process.

[0199] From a machine perspective, the system will send satisfaction surveys to users or request direct feedback from users to collect evaluations and suggestions on bathing services. These feedback data are crucial to improving the quality of bathing services and meeting user needs. The system will organize and analyze the collected user feedback data to identify the strengths and weaknesses of the service. The analysis results will be used to guide the improvement and optimization of subsequent services and improve user satisfaction and experience.

[0200] Based on user feedback and the actual execution of bathing tasks, the global model of the embodied intelligence model and the federated learning module is optimized. The accuracy and generalization ability of the model are improved by adjusting model parameters, improving model structure or introducing new algorithms and technologies. The optimized model will be used for the execution and decision support of subsequent bathing tasks.

[0201] Based on user feedback and the actual execution effect of bathing tasks, the hardware and software systems of the intelligent bathing robot are optimized. Possible optimization directions include improving the accuracy and stability of sensors, optimizing task execution processes, improving the human-computer interaction interface, etc. System optimization will improve the performance and reliability of the robot and further enhance the user experience.

[0202] It should be noted that the robot will continuously collect new data and learn and train to adapt to the bathing needs of different users and environmental changes. Through continuous learning, the robot will continuously improve its intelligence level and adaptability, and regularly release system updates and upgrades to fix potential security vulnerabilities and improve system performance and stability. Updates and upgrades will include new functions, algorithms and interface designs to meet the ever-changing needs and expectations of users.

[0203] Corresponding to the above-mentioned robot control method, the present invention also provides a robot control device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be repeated in the present invention.

[0204] Figure 5 A schematic diagram of the structure of a robot control device provided by an embodiment of the present disclosure, such as Figure 5 As shown, the device is configured in a robot in a control system of the robot, and includes:

[0205] The analysis unit 31 is used to respond to the target demand instruction and call the trained first embodied intelligent model to analyze the target demand instruction to obtain the multiple target tasks corresponding to the target demand instruction, wherein the global model parameters of the first embodied intelligent model are configured by the server and are obtained by local training based on the global model parameters and user historical data;

[0206] An acquisition unit 32, used to acquire current environment information and status information of the user served by the target demand instruction;

[0207] A composition unit 33, configured to obtain a plurality of task execution plans composed of the plurality of target tasks based on the current environment information and the state information;

[0208] A determination unit 34, configured to determine a target task execution plan from the plurality of task execution plans according to a priority order and target requirements preset by a user;

[0209] The adjustment unit 35 is used to dynamically adjust the target task execution plan during the process of executing each target task in the target task execution plan based on the user's feedback information on the target task and / or the update of the target requirements until all target tasks in the target task execution plan are executed.

[0210] The present application uses the global model parameters of the trained second embodied intelligent model on the server side and sends the global model parameters to the federated learning module of the robot. After the robot side obtains the trained first embodied intelligent model based on the global model parameters, the target demand instructions are analyzed to obtain the target task execution plan corresponding to the target demand instructions. The target task execution plan will be dynamically adjusted according to the feedback from service users and the current environment to better meet user needs and improve user comfort.

[0211] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the device also includes:

[0212] The uploading unit 36 ​​is used to upload the model parameters corresponding to the trained first embodied intelligent model to the server, so that the server can perform federated learning according to the different model parameters of each robot, obtain updated global model parameters, and send the updated global model parameters to each robot.

[0213] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the component unit 33 is also used for:

[0214] Based on the current environment information, the execution strategy of each target task in each task execution plan and the serial and parallel order of each target task are adjusted to obtain the multiple task execution plans.

[0215] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the determining unit 34 is further used for:

[0216] Call the value assessment standard to evaluate each target task in each task execution plan and obtain the assessment result;

[0217] According to the preset priority order and the target requirement, the best evaluation result among the evaluation results is used as the target task execution plan.

[0218] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 6 As shown, the device also includes:

[0219] A receiving unit 37, configured to receive updated global model parameters sent by the server after the uploading unit 36 ​​uploads the model parameters corresponding to the trained first embodied intelligence model to the server;

[0220] The training unit 38 is used to train the first embodied intelligence model using the updated global model parameters, user historical data, and user historical feedback information to obtain the trained first embodied intelligence model.

[0221] The present application also provides a robot control device, which is applied to a server in a robot control system, such as Figure 7 As shown, including:

[0222] A construction unit 41, configured to construct a second embodied intelligence model;

[0223] A training unit 42, configured to input the multimodal training data into the second embodied intelligence model for training, and obtain a trained second embodied intelligence model, wherein the trained second embodied intelligence model corresponds to a global model parameter;

[0224] The sending unit 43 is used to send the global model parameters to the robot so that the robot can locally train the first embodied intelligent model in the robot according to the global model parameters and the user's historical data.

[0225] The present application uses the global model parameters of the trained second embodied intelligent model on the server side and sends the global model parameters to the federated learning module of the robot. After the robot side obtains the trained first embodied intelligent model based on the global model parameters, the target demand instructions are analyzed to obtain the target task execution plan corresponding to the target demand instructions. The target task execution plan will be dynamically adjusted according to the feedback from service users and the current environment to better meet user needs and improve user comfort.

[0226] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 8 As shown, the training unit 42 is also used for:

[0227] In the second embodied intelligent model, the multimodal training data is subjected to feature extraction and fusion processing to obtain fused training data; the multimodal training data includes at least two of camera training data, lidar training data, sound training data, text training data, temperature training data, smell training data and tactile training data;

[0228] Cross-modal reasoning is performed based on the fused training data to obtain the trained second embodied intelligence model.

[0229] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 8 As shown, the device also includes:

[0230] A receiving unit 44, configured to receive new global model parameters respectively sent by each robot after the training unit inputs the multimodal training data into the second embodied intelligent model for training;

[0231] The training unit 42 is further used to aggregate all new global model parameters based on a preset aggregation algorithm to obtain aggregated global model parameters;

[0232] The second embodied intelligent model is trained based on the aggregated global model parameters to obtain updated global model parameters, and the updated global model parameters are sent to all robots respectively.

[0233] The present application also provides a robot, the robot comprising: Figure 6 or Figure 7 The control device of the robot shown.

[0234] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present disclosure, and the principle is the same, which is no longer limited in the embodiment of the present disclosure.

[0235] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0236] Fig. 9 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0237] like Fig. 9 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 to a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0238] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0239] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as a control method for a robot. For example, in some embodiments, the control method for a robot may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the aforementioned robot control method in any other appropriate manner (eg, by means of firmware).

[0240] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor that may be a special purpose or general purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0241] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0242] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory) or a flash memory, an optical fiber, a CD-ROM (Compact Dis sc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0243] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0244] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0245] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0246] It should be noted that artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.

[0247] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0248] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A robot control system, characterized in that: The system includes a robot and a server, wherein the robot includes a task planning module and a federated learning module, and the server includes a model building and reasoning module; The task planning module is used to respond to the target demand instruction, call the trained first embodied intelligent model to analyze the target demand instruction, obtain multiple target tasks corresponding to the target demand instruction, obtain current environment information and status information of the user served by the target demand instruction, and obtain multiple task execution plans composed of the multiple target tasks based on the current environment information and the status information, and determine the target task execution plan from the multiple task execution plans according to the priority order preset by the user and the target demand; The federated learning module is configured to receive the global model parameters of the first embodied intelligence model sent by the server, and perform local training on the first embodied intelligence model and user historical data according to the global model parameters; The model building and reasoning module is used to input the multimodal training data into the second embodied intelligent model for training after building the second embodied intelligent model, obtain corresponding global model parameters, and send the global model parameters to the federated learning module; the multimodal training data includes at least two of camera training data, lidar training data, sound training data, text training data, temperature training data, smell training data and tactile training data; The task planning module includes: a target analysis unit, a task planning unit, a task execution unit, an environment perception and adjustment unit and a decision-making unit; The target analysis unit is used to analyze the target requirement instruction in response to the target requirement instruction to obtain the multiple target tasks corresponding to the target requirement instruction; The task planning unit is used to obtain the current environment information and the status information of the user served by the target demand instruction, and obtain multiple task execution plans consisting of the multiple target tasks based on the current environment information and the status information; The decision unit is used to determine the target task execution plan from the multiple task execution plans according to the priority order and target requirements preset by the user; The task execution unit is used to control the execution of each target task in the target task execution plan; The environment perception and adjustment unit is used to dynamically adjust the target task execution plan based on the user's feedback information on the target task and / or the update of the target requirement until all target tasks in the target task execution plan are executed.

2. The system according to claim 1, characterized in that The server includes: a multimodal data fusion module and an embodied intelligent modeling engine; The multimodal data fusion module is used to perform feature extraction and fusion processing on the multimodal training data to obtain fused training data, and input the fused training data into the embodied intelligent modeling engine; The embodied intelligence modeling engine is used to build the embodied intelligence model, and train the second embodied intelligence model based on the fused training data to obtain the global model parameters corresponding to the trained second embodied intelligence model, and send the global model parameters to the federated learning module.

3. A robot control method, characterized in that: The method is applied to a robot in a control system of a robot according to any one of claims 1 to 2, comprising: In response to the target demand instruction, calling the trained first embodied intelligent model to analyze the target demand instruction to obtain the multiple target tasks corresponding to the target demand instruction, wherein the global model parameters of the first embodied intelligent model are configured by the server and obtained by local training based on the global model parameters and user historical data; Acquire current environment information and status information of a user served by the target demand instruction, and obtain multiple task execution plans consisting of the multiple target tasks based on the current environment information and the status information; Determining a target task execution plan from the plurality of task execution plans according to a priority order and target requirements preset by a user; During the process of executing each target task according to the target task execution plan, the target task execution plan is dynamically adjusted based on the user's feedback information on the target task and / or the update of the target requirement until all the target tasks in the target task execution plan are executed.

4. The method according to claim 3, characterized in that The method further comprises: The model parameters corresponding to the trained first embodied intelligent model are uploaded to the server, so that the server can perform federated learning according to the different model parameters of each robot, obtain updated global model parameters, and send the updated global model parameters to each robot.

5. The method according to claim 3, characterized in that: The obtaining of a plurality of task execution plans consisting of the plurality of target tasks based on the current environment information and the state information comprises: Planning execution steps for each target task based on the current environment information and the state information in the trained first embodied intelligent model; Based on the current environment information, the execution strategy of each target task in each task execution plan and the serial and parallel order of each target task are adjusted to obtain the multiple task execution plans.

6. The method according to claim 3, characterized in that Determining the target task execution plan from the multiple task execution plans according to the priority order and target requirements preset by the user includes: Call the value assessment standard to evaluate each target task in each task execution plan and obtain the assessment result; According to the preset priority order and the target requirement, the best evaluation result among the evaluation results is used as the target task execution plan.

7. The method according to claim 4, characterized in that After uploading the model parameters corresponding to the trained first embodied intelligence model to the server, the method further includes: Receiving updated global model parameters sent by the server; The first embodied intelligence model is trained using the updated global model parameters, user historical data, and user historical feedback information to obtain the trained first embodied intelligence model.

8. A robot control method, characterized in that: The method is applied to a server in a control system of a robot according to any one of claims 1 to 2, comprising: Constructing a second embodied intelligence model; Inputting the multimodal training data into the second embodied intelligent model for training to obtain a trained second embodied intelligent model, wherein the trained second embodied intelligent model corresponds to a global model parameter; Sending the global model parameters to the robot so that the robot performs local training on the first embodied intelligence model in the robot according to the global model parameters and the user's historical data; The step of inputting the multimodal training data into the second embodied intelligence model for training to obtain a trained second embodied intelligence model comprises: In the second embodied intelligent model, the multimodal training data is subjected to feature extraction and fusion processing to obtain fused training data containing the extracted features; the multimodal training data includes at least two of camera training data, lidar training data, sound training data, text training data, temperature training data, smell training data and tactile training data; Cross-modal training is performed based on the fused training data to obtain the trained second embodied intelligence model.

9. The method according to claim 8, characterized in that After inputting the multimodal training data into the second embodied intelligence model for training, the method further includes: Receive new global model parameters sent by each robot; Inputting the multimodal training data into the second embodied intelligence model for training includes: Aggregate all new global model parameters based on a preset aggregation algorithm to obtain aggregated global model parameters; The second embodied intelligent model is trained based on the aggregated global model parameters to obtain updated global model parameters, and the updated global model parameters are sent to all robots respectively.

10. A robot control device, characterized in that: The device is configured in a robot in a control system of a robot according to any one of claims 1 to 2, and comprises: an analysis unit, configured to respond to a target demand instruction and call a trained first embodied intelligent model to analyze the target demand instruction to obtain the multiple target tasks corresponding to the target demand instruction, wherein the global model parameters of the first embodied intelligent model are configured by the server and are obtained by local training based on the global model parameters and user historical data; An acquisition unit, used to acquire current environment information and status information of the user served by the target demand instruction; A composition unit, configured to obtain a plurality of task execution plans composed of the plurality of target tasks based on the current environment information and the state information; A determination unit, configured to determine a target task execution plan from the plurality of task execution plans according to a priority order and target requirements preset by a user; The adjustment unit is used to dynamically adjust the target task execution plan during the execution of each target task in the target task execution plan based on the user's feedback information on the target task and / or the update of the target requirements until all target tasks in the target task execution plan are executed.

11. A robot control device, characterized in that: The device is configured in a server in a control system of a robot according to any one of claims 1 to 2, and comprises: A construction unit, used for constructing a second embodied intelligence model; a training unit, configured to input the multimodal training data into the second embodied intelligence model for training, so as to obtain a trained second embodied intelligence model, wherein the trained second embodied intelligence model corresponds to a global model parameter; a sending unit, configured to send the global model parameters to the robot, so that the robot performs local training on the second embodied intelligent model in the robot according to the global model parameters and the user's historical data; The training unit is further used to perform feature extraction and fusion processing on the multimodal training data in the second embodied intelligent model to obtain fused training data containing the extracted features; the multimodal training data includes at least two of camera training data, lidar training data, sound training data, text training data, temperature training data, smell training data and tactile training data; Cross-modal training is performed based on the fused training data to obtain the trained second embodied intelligence model.

12. A robot, characterized in that: The robot comprises the control device of the robot according to claim 10.

13. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 3-7 or 8-9.

14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 3-7 or 8-9.

15. A computer program product, characterized in that It comprises a computer program which, when executed by a processor, implements the method according to any one of claims 3-7 or 8-9.

Citation Information

Patent Citations

  • Federal learning-based model training method, device and system, and storage medium

    CN114819190A

  • Control method and system of intelligent robot with body, electronic equipment and storage medium

    CN117885082A