Dual-arm robot control method, device, equipment and medium
Patent Information
- Application Number
- CN202511055831.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-07-29
AI Technical Summary
[0005]鉴于以上内容,有必要提供一种双臂机器人控制方法、装置、设备及介质,旨在解决由于动作决策模型泛化性能低而导致的双臂机器人对新环境适应性差的问题
[0024]由以上技术方案可以看出,本发明能够利用双臂机器人配备的多类型传感器采集当前环境数据以训练视觉语言动作模型,使模型初步具备基础的泛化能力;利用双臂机器人动作决策模型的特征提取模块,基于注意力机制动态提取实时环境信息的多模态特征,以快速捕获适应动态环境的多模态特征表示,提高模型对新环境的感知力,从而提高模型泛化能力;将解析得到的任务目标、约束条件及多模态特征输入至双臂机器人动作决策模型的策略网络,得到双臂机器人的动作序列,并控制双臂机器人执行动作序列,以实现对新环境下任务的准确理解及决策,提高了模型对于新环境的适应性。
Smart Images

Figure CN120921372B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a dual-arm robot control method, device, equipment, and medium. Background Technology
[0002] In robotic systems equipped with two arms, the decision-making ability based on the Vision-Language-Action (VLA) model is limited by specific environments and task scenarios. Traditional models lack continuous learning and dynamic adaptation capabilities, making it difficult to make reasonable decisions in complex and ever-changing environments.
[0003] For example, when a robot in a medical laboratory moves from the laboratory environment to a home setting, it may make decision-making errors or be unable to perform the task due to changes in the environment layout, object characteristics, and task requirements. Similarly, when a new self-service area is put into use in a bank lobby, the model may also make decision-making errors or be unable to perform the task due to changes in the environment layout, object characteristics, and task requirements.
[0004] In view of the above problems, it is urgent to improve the environmental adaptability of the action decision model of dual-arm robots in order to improve their generalization performance. Summary of the Invention
[0005] In view of the above, it is necessary to provide a dual-arm robot control method, device, equipment and medium, which aims to solve the problem of poor adaptability of dual-arm robots to new environments due to the low generalization performance of motion decision models.
[0006] A dual-arm robot control method, the dual-arm robot control method comprising:
[0007] The current environmental data is collected by using multiple types of sensors equipped on the dual-arm robot, and the current environmental data is used to train a visual language action model to obtain a dual-arm robot action decision model.
[0008] In response to control commands to the dual-arm robot in the new environment, real-time environmental information of the new environment is collected through the various types of sensors;
[0009] Using the feature extraction module of the dual-arm robot motion decision model, multimodal features of the real-time environmental information are dynamically extracted based on an attention mechanism;
[0010] Using the natural language processing module of the dual-arm robot motion decision model, the control commands are parsed based on the multimodal features to obtain the task objectives and constraints;
[0011] The task objective, the constraints, and the multimodal features are input into the policy network of the dual-arm robot motion decision model to obtain the motion sequence of the dual-arm robot.
[0012] Control the dual-arm robot to execute the sequence of actions.
[0013] A dual-arm robot control device, the dual-arm robot control device comprising:
[0014] The training unit is used to collect current environmental data using multiple types of sensors equipped on the dual-arm robot, and to train a visual language action model using the current environmental data to obtain a dual-arm robot action decision model.
[0015] The data acquisition unit is used to respond to control commands to the dual-arm robot in a new environment by acquiring real-time environmental information of the new environment through the multiple types of sensors.
[0016] The extraction unit is used to dynamically extract multimodal features of the real-time environmental information based on the attention mechanism using the feature extraction module of the dual-arm robot motion decision model.
[0017] The parsing unit is used to use the natural language processing module of the dual-arm robot motion decision model to parse the control instructions based on the multimodal features to obtain the task objective and constraints.
[0018] The decision unit is used to input the task objective, the constraints and the multimodal features into the policy network of the dual-arm robot action decision model to obtain the action sequence of the dual-arm robot.
[0019] A control unit is used to control the dual-arm robot to execute the action sequence.
[0020] A computer device, the computer device comprising:
[0021] Memory, storing at least one instruction; and
[0022] The processor executes the instructions stored in the memory to implement the dual-arm robot control method.
[0023] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the dual-arm robot control method.
[0024] As can be seen from the above technical solutions, this invention can utilize the multiple types of sensors equipped on a dual-arm robot to collect current environmental data to train a visual language action model, enabling the model to initially possess basic generalization capabilities. By utilizing the feature extraction module of the dual-arm robot's action decision model, multimodal features of real-time environmental information are dynamically extracted based on an attention mechanism to quickly capture multimodal feature representations adapted to dynamic environments, improving the model's perception of new environments and thus enhancing its generalization ability. The parsed task objectives, constraints, and multimodal features are input into the policy network of the dual-arm robot's action decision model to obtain the action sequence of the dual-arm robot, and the dual-arm robot is controlled to execute the action sequence, thereby achieving accurate understanding and decision-making regarding tasks in new environments and improving the model's adaptability to new environments. Attached Figure Description
[0025] Figure 1 This is a flowchart of a preferred embodiment of the dual-arm robot control method of the present invention.
[0026] Figure 2 This is a functional block diagram of a preferred embodiment of the dual-arm robot control device of the present invention.
[0027] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the dual-arm robot control method of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the dual-arm robot control method of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0030] The dual-arm robot control method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0031] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0032] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0033] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0034] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0035] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0036] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0037] S10: Collect current environmental data using multiple types of sensors equipped on the dual-arm robot, and use the current environmental data to train a Vision-Language-Action (VLA) model to obtain a dual-arm robot action decision model.
[0038] In this embodiment, the dual-arm robot may include, but is not limited to, service robots in financial scenarios and medical device grasping robots in medical and health scenarios.
[0039] In this embodiment, the multiple types of sensors may include, but are not limited to, one or a combination of the following sensors:
[0040] Binocular cameras, microphones, joint angle sensors, tactile sensors, etc.
[0041] In this embodiment, the current environment data may include, but is not limited to, one or more of the following combinations of data:
[0042] Image data acquired by the binocular camera, audio data acquired by the microphone, arm movement data acquired by the joint angle sensor, and tactile data acquired by the tactile sensor.
[0043] The arm movement data and the tactile data together constitute the arm movement and tactile data.
[0044] In this embodiment, training the visual language action model using the current environmental data to obtain the dual-arm robot action decision model includes:
[0045] The first feature is obtained by extracting features from the image data in the current environmental data using the Swing Transformer model.
[0046] The text data in the current environment data is encoded using the BERT (BIDIRECTIONAL ENCODER REPRESENTATIONS FROM TRANSFORMERS) model to obtain the second feature;
[0047] Mel frequency cepstral coefficients of the audio data in the current environment data are extracted, and the Mel frequency cepstral coefficients are processed using a recurrent neural network to obtain the third feature;
[0048] The arm movement and tactile data in the current environmental data are processed using a multilayer perceptron to obtain the fourth feature;
[0049] The first feature, the second feature, the third feature, and the fourth feature are deeply fused using a self-attention mechanism to obtain a fused feature;
[0050] The visual language action model is trained using the fused features to obtain the action decision model of the dual-arm robot.
[0051] The fusion features can reflect the typical characteristics of the environment. For example, through the fusion features of multimodal data, data can be collected and multimodal datasets can be constructed for different tasks such as handling, assembly, cleaning, and instrument grasping in typical environments such as industrial workshops, family living rooms, warehouses, and operating rooms. This allows the constructed multimodal datasets to comprehensively cover various attributes of the environment.
[0052] For example, in a financial self-service center environment, cameras can capture images of users operating self-service equipment, microphones can collect users' voice commands, and sensors can record the operational status of the self-service equipment. A dataset can then be built, enabling the trained model to understand the user's intentions for depositing or withdrawing money and the equipment's status. Similarly, in a hospital ward environment, cameras can capture images of the ward environment and patient activities, microphones can collect instructions from medical staff and patient calls, and sensors can record operational data from medical equipment. A dataset can then be built, enabling the trained model to understand the ward environment, the status of medical equipment, and simple medical instructions.
[0053] Through the above embodiments, a large-scale multimodal dataset is first constructed, and then the visual language action model is trained using the dataset, enabling the model to learn general multimodal feature representations and basic decision-making logic, possess basic generalization ability, and be able to process and understand basic information of different environments and tasks.
[0054] S11, in response to control commands to the dual-arm robot in the new environment, real-time environmental information of the new environment is collected through the multiple types of sensors.
[0055] In this embodiment, the control command can be triggered when a specific voice instruction is detected. For example, when the voice command "hemostats" is received in the operating room, the control command is triggered.
[0056] In this embodiment, once the control command is received, real-time environmental information of the new environment can be collected immediately through the multiple types of sensors.
[0057] S12, using the feature extraction module of the dual-arm robot motion decision model, dynamically extract the multimodal features of the real-time environmental information based on the attention mechanism.
[0058] In this embodiment, since the action decision model of the dual-arm robot is trained using existing environmental data, in order to adapt it to the new environment, the weights of each modal data need to be dynamically adjusted according to the attributes of the new environment.
[0059] Specifically, the feature extraction module utilizing the dual-arm robot motion decision model to dynamically extract multimodal features of the real-time environmental information based on an attention mechanism includes:
[0060] The environmental characteristics of the new environment are determined based on the real-time environmental information;
[0061] Based on the environmental characteristics, the weights of each modality are dynamically adjusted according to the attention mechanism.
[0062] The feature extraction module is used to extract the multimodal features based on the weight of each modality.
[0063] Among these features, the pre-trained parameters of the dual-arm robot motion decision model can be used to quickly extract multimodal features of the new environment.
[0064] The attention mechanism described above can be used to dynamically adjust the weights of each modality based on the attributes of the new environment. For example, when the lighting in the new environment is complex, the visual weight can be increased; when the new environment is noisy, audio recognition can be strengthened.
[0065] For example, in financial scenarios, in a new self-service area of a bank lobby, the robot can collect real-time images of the device locations and user flow, detect sounds in the lobby (such as user inquiries and device prompts), and dynamically extract these features. This allows the model to quickly adapt to the new self-service environment and accurately perceive the status of users and devices. In healthcare scenarios, when the robot enters a new operating room, it can collect real-time images of the operating room's equipment layout and the locations of medical staff, detect instrument sounds and conversations among medical staff during surgery, and dynamically extract features to help the model quickly understand the new operating room environment and provide environmental data for assisting the surgery.
[0066] Through the above embodiments, multimodal feature representations adapted to dynamic environments can be quickly obtained, making the model's perception of new environments more accurate and comprehensive, and providing a reliable environmental information foundation for subsequent decision-making.
[0067] S13, using the natural language processing module of the dual-arm robot motion decision model, the control command is parsed according to the multimodal features to obtain the task objective and constraints.
[0068] In this embodiment, the control instructions can be semantically parsed using natural language processing algorithms, and the task objectives and constraints can be understood in conjunction with multimodal environment features.
[0069] The constraints mentioned above are a type of safety constraint, which prevents collisions and other problems from occurring during the movement of the dual-arm robot.
[0070] For example, for the control instruction "hand the surgical instrument to the doctor", semantic parsing can yield the task objective "coordinated use of both arms to retrieve the surgical instrument and hand it to the doctor", with the constraint "avoid obstacles during the process of retrieving and handing the surgical instrument to the doctor to avoid the risk of collision". For the control instruction "place the red box on the table on the second shelf", the model identifies "the red box and the shelf as the operation object and target location" through visual features, with the constraint "avoid obstacles during the process of retrieving and placing the box to avoid the risk of collision".
[0071] S14, the task objective, the constraints, and the multimodal features are input into the policy network of the dual-arm robot motion decision model to obtain the motion sequence of the dual-arm robot.
[0072] In this embodiment, before inputting the task objective, the constraints, and the multimodal features into the policy network of the dual-arm robot motion decision model, the method further includes:
[0073] Obtain a pre-built knowledge base; wherein the knowledge base is used to store the correspondence between environmental information, task information and action decisions;
[0074] The task objective, the constraints, and the multimodal features are used to perform similarity retrieval in the knowledge base to obtain similar knowledge.
[0075] The action sequence is generated based on the similarity knowledge.
[0076] For example, the knowledge base can store knowledge about object attributes, task processes, and other information relevant to different environments and tasks. This way, when the robot encounters new medical consumables, it can leverage its knowledge of using similar consumables to assist medical staff more quickly and effectively.
[0077] Through the above embodiments, when entering a new scene, similar historical experiences are retrieved first through similar knowledge to transfer relevant knowledge, such as reusing the grasping strategy of similar objects, which accelerates the decision-making process in the new environment. This allows the model to learn new knowledge continuously without forgetting existing knowledge, avoid making repeated decisions on similar problems, improve decision-making efficiency, and continuously enhance its adaptability to new environments and tasks.
[0078] S15, control the dual-arm robot to execute the action sequence.
[0079] In this embodiment, the action sequence may include joint angle changes, gripping force control, motion path planning, etc.
[0080] In this embodiment, controlling the dual-arm robot to execute the action sequence includes:
[0081] The dual-arm robot is controlled to move according to the action sequence, and during the movement, the dual-arm robot is controlled to avoid obstacles based on the constraints.
[0082] The above embodiments can introduce safety constraints to ensure that the arm movements will not collide with objects in the environment or cause danger.
[0083] In this embodiment, the method further includes:
[0084] During the process of controlling the dual-arm robot to execute the action sequence, the execution deviation is continuously detected, and the actual action data and environmental change information of the dual-arm robot are continuously collected;
[0085] Feedback data is generated based on the execution deviation;
[0086] The dual-arm robot motion decision model is incrementally updated based on the feedback data, the actual motion data, and the environmental change information.
[0087] The execution deviations may include fetching failures, path obstruction, etc.
[0088] The actual action data and the environmental change information can be compared with the generated decision, i.e., the action sequence.
[0089] Incremental learning algorithms can be used to incrementally update the motion decision model of the dual-arm robot.
[0090] Through the above embodiments, while retaining existing knowledge, the model learns the unique characteristics and decision-making patterns in new environments and tasks, avoiding catastrophic forgetting; by adjusting model parameters, the model can better adapt to new scenarios, thereby continuously expanding its knowledge reserves and decision-making capabilities.
[0091] In this embodiment, after controlling the dual-arm robot to execute the action sequence, the method further includes:
[0092] Construct a multi-dimensional model evaluation system;
[0093] The decision-making performance of the dual-arm robot motion decision-making model in the new environment is evaluated based on the multi-dimensional model evaluation system, and the evaluation results are obtained.
[0094] Based on the evaluation results, an adjustment strategy is generated for the motion decision model of the dual-arm robot.
[0095] The evaluation indicators in the multi-dimensional model evaluation system may include, but are not limited to, one or more of the following indicators:
[0096] Task completion rate, decision-making accuracy, smoothness of action, execution efficiency, etc.
[0097] Among these, the decision-making performance of the model in a new environment can be evaluated based on the multi-dimensional model evaluation system, and the advantages and disadvantages of the model can be analyzed.
[0098] After obtaining the evaluation results, the reinforcement learning reward function can be adjusted, the multimodal feature fusion method can be optimized, the knowledge transfer and incremental learning strategies can be improved, and the model can be retrained and tested regularly to ensure that the model continues to learn and remains available.
[0099] For example, in the medical field, if an evaluation determines that a dual-arm robot's movements are not smooth enough when assisting in the delivery of surgical tools, the model can be optimized and iterated to make the dual-arm robot's movements more stable and efficient during the surgical process, thereby improving the quality of care.
[0100] For example, in the financial sector, when an assessment determines that a collision occurs while a dual-arm robot is moving items in a service hall, the model can be optimized to make the path of the dual-arm robot more reasonable during the assisted handling process.
[0101] Through the above embodiments, the model can continuously improve its generalization performance during the continuous learning process, enabling it to stably and efficiently control the robot's arms through reasonable decision-making in various complex environments.
[0102] As can be seen from the above technical solutions, this invention can utilize the multiple types of sensors equipped on a dual-arm robot to collect current environmental data to train a visual language action model, enabling the model to initially possess basic generalization capabilities. By utilizing the feature extraction module of the dual-arm robot's action decision model, multimodal features of real-time environmental information are dynamically extracted based on an attention mechanism to quickly capture multimodal feature representations adapted to dynamic environments, improving the model's perception of new environments and thus enhancing its generalization ability. The parsed task objectives, constraints, and multimodal features are input into the policy network of the dual-arm robot's action decision model to obtain the action sequence of the dual-arm robot, and the dual-arm robot is controlled to execute the action sequence, thereby achieving accurate understanding and decision-making regarding tasks in new environments and improving the model's adaptability to new environments.
[0103] like Figure 2The diagram shown is a functional block diagram of a preferred embodiment of the dual-arm robot control device of the present invention. The dual-arm robot control device 11 includes a training unit 110, a data acquisition unit 111, an extraction unit 112, a parsing unit 113, a decision-making unit 114, and a control unit 115. The module / unit referred to in this invention is a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0104] The training unit 110 is used to collect current environmental data using multiple types of sensors equipped on the dual-arm robot, and to train a Vision-Language-Action (VLA) model using the current environmental data to obtain a dual-arm robot action decision model.
[0105] In this embodiment, the dual-arm robot may include, but is not limited to, service robots in financial scenarios and medical device grasping robots in medical and health scenarios.
[0106] In this embodiment, the multiple types of sensors may include, but are not limited to, one or a combination of the following sensors:
[0107] Binocular cameras, microphones, joint angle sensors, tactile sensors, etc.
[0108] In this embodiment, the current environment data may include, but is not limited to, one or more of the following combinations of data:
[0109] Image data acquired by the binocular camera, audio data acquired by the microphone, arm movement data acquired by the joint angle sensor, and tactile data acquired by the tactile sensor.
[0110] The arm movement data and the tactile data together constitute the arm movement and tactile data.
[0111] In this embodiment, the training unit 110 uses the current environmental data to train the visual language action model to obtain the dual-arm robot action decision model, including:
[0112] The first feature is obtained by extracting features from the image data in the current environmental data using the Swing Transformer model.
[0113] The text data in the current environment data is encoded using the BERT (BIDIRECTIONAL ENCODER REPRESENTATIONS FROM TRANSFORMERS) model to obtain the second feature;
[0114] Mel frequency cepstral coefficients of the audio data in the current environment data are extracted, and the Mel frequency cepstral coefficients are processed using a recurrent neural network to obtain the third feature;
[0115] The arm movement and tactile data in the current environmental data are processed using a multilayer perceptron to obtain the fourth feature;
[0116] The first feature, the second feature, the third feature, and the fourth feature are deeply fused using a self-attention mechanism to obtain a fused feature;
[0117] The visual language action model is trained using the fused features to obtain the action decision model of the dual-arm robot.
[0118] The fusion features can reflect the typical characteristics of the environment. For example, through the fusion features of multimodal data, data can be collected and multimodal datasets can be constructed for different tasks such as handling, assembly, cleaning, and instrument grasping in typical environments such as industrial workshops, family living rooms, warehouses, and operating rooms. This allows the constructed multimodal datasets to comprehensively cover various attributes of the environment.
[0119] For example, in a financial self-service center environment, cameras can capture images of users operating self-service equipment, microphones can collect users' voice commands, and sensors can record the operational status of the self-service equipment. A dataset can then be built, enabling the trained model to understand the user's intentions for depositing or withdrawing money and the equipment's status. Similarly, in a hospital ward environment, cameras can capture images of the ward environment and patient activities, microphones can collect instructions from medical staff and patient calls, and sensors can record operational data from medical equipment. A dataset can then be built, enabling the trained model to understand the ward environment, the status of medical equipment, and simple medical instructions.
[0120] Through the above embodiments, a large-scale multimodal dataset is first constructed, and then the visual language action model is trained using the dataset, enabling the model to learn general multimodal feature representations and basic decision-making logic, possess basic generalization ability, and be able to process and understand basic information of different environments and tasks.
[0121] The acquisition unit 111 is used to collect real-time environmental information of the new environment through the multi-type sensors in response to control commands to the dual-arm robot in the new environment.
[0122] In this embodiment, the control command can be triggered when a specific voice instruction is detected. For example, when the voice command "hemostats" is received in the operating room, the control command is triggered.
[0123] In this embodiment, once the control command is received, real-time environmental information of the new environment can be collected immediately through the multiple types of sensors.
[0124] The extraction unit 112 is used to dynamically extract the multimodal features of the real-time environmental information based on the attention mechanism using the feature extraction module of the dual-arm robot motion decision model.
[0125] In this embodiment, since the action decision model of the dual-arm robot is trained using existing environmental data, in order to adapt it to the new environment, the weights of each modal data need to be dynamically adjusted according to the attributes of the new environment.
[0126] Specifically, the extraction unit 112 utilizes the feature extraction module of the dual-arm robot motion decision model to dynamically extract multimodal features of the real-time environmental information based on an attention mechanism, including:
[0127] The environmental characteristics of the new environment are determined based on the real-time environmental information;
[0128] Based on the environmental characteristics, the weights of each modality are dynamically adjusted according to the attention mechanism.
[0129] The feature extraction module is used to extract the multimodal features based on the weight of each modality.
[0130] Among these features, the pre-trained parameters of the dual-arm robot motion decision model can be used to quickly extract multimodal features of the new environment.
[0131] The attention mechanism described above can be used to dynamically adjust the weights of each modality based on the attributes of the new environment. For example, when the lighting in the new environment is complex, the visual weight can be increased; when the new environment is noisy, audio recognition can be strengthened.
[0132] For example, in financial scenarios, in a new self-service area of a bank lobby, the robot can collect real-time images of the device locations and user flow, detect sounds in the lobby (such as user inquiries and device prompts), and dynamically extract these features. This allows the model to quickly adapt to the new self-service environment and accurately perceive the status of users and devices. In healthcare scenarios, when the robot enters a new operating room, it can collect real-time images of the operating room's equipment layout and the locations of medical staff, detect instrument sounds and conversations among medical staff during surgery, and dynamically extract features to help the model quickly understand the new operating room environment and provide environmental data for assisting the surgery.
[0133] Through the above embodiments, multimodal feature representations adapted to dynamic environments can be quickly obtained, making the model's perception of new environments more accurate and comprehensive, and providing a reliable environmental information foundation for subsequent decision-making.
[0134] The parsing unit 113 is used to use the natural language processing module of the dual-arm robot motion decision model to parse the control instructions based on the multimodal features to obtain the task objective and constraints.
[0135] In this embodiment, the control instructions can be semantically parsed using natural language processing algorithms, and the task objectives and constraints can be understood in conjunction with multimodal environment features.
[0136] The constraints mentioned above are a type of safety constraint, which prevents collisions and other problems from occurring during the movement of the dual-arm robot.
[0137] For example, for the control instruction "hand the surgical instrument to the doctor", semantic parsing can yield the task objective "coordinated use of both arms to retrieve the surgical instrument and hand it to the doctor", with the constraint "avoid obstacles during the process of retrieving and handing the surgical instrument to the doctor to avoid the risk of collision". For the control instruction "place the red box on the table on the second shelf", the model identifies "the red box and the shelf as the operation object and target location" through visual features, with the constraint "avoid obstacles during the process of retrieving and placing the box to avoid the risk of collision".
[0138] The decision unit 114 is used to input the task objective, the constraints and the multimodal features into the policy network of the dual-arm robot action decision model to obtain the action sequence of the dual-arm robot.
[0139] In this embodiment, the decision unit 114 inputs the task objective, the constraints, and the multimodal features into the policy network of the dual-arm robot action decision model to obtain a pre-built knowledge base; wherein, the knowledge base is used to store the correspondence between environmental information, task information, and action decisions;
[0140] The task objective, the constraints, and the multimodal features are used to perform similarity retrieval in the knowledge base to obtain similar knowledge.
[0141] The action sequence is generated based on the similarity knowledge.
[0142] For example, the knowledge base can store knowledge about object attributes, task processes, and other information relevant to different environments and tasks. This way, when the robot encounters new medical consumables, it can leverage its knowledge of using similar consumables to assist medical staff more quickly and effectively.
[0143] Through the above embodiments, when entering a new scene, similar historical experiences are retrieved first through similar knowledge to transfer relevant knowledge, such as reusing the grasping strategy of similar objects, which accelerates the decision-making process in the new environment. This allows the model to learn new knowledge continuously without forgetting existing knowledge, avoid making repeated decisions on similar problems, improve decision-making efficiency, and continuously enhance its adaptability to new environments and tasks.
[0144] The control unit 115 is used to control the dual-arm robot to execute the action sequence.
[0145] In this embodiment, the action sequence may include joint angle changes, gripping force control, motion path planning, etc.
[0146] In this embodiment, the control unit 115 controls the dual-arm robot to execute the action sequence, including:
[0147] The dual-arm robot is controlled to move according to the action sequence, and during the movement, the dual-arm robot is controlled to avoid obstacles based on the constraints.
[0148] The above embodiments can introduce safety constraints to ensure that the arm movements will not collide with objects in the environment or cause danger.
[0149] In this embodiment, during the process of controlling the dual-arm robot to execute the action sequence, execution deviations are continuously detected, and the actual action data and environmental change information of the dual-arm robot are continuously collected;
[0150] Feedback data is generated based on the execution deviation;
[0151] The dual-arm robot motion decision model is incrementally updated based on the feedback data, the actual motion data, and the environmental change information.
[0152] The execution deviations may include fetching failures, path obstruction, etc.
[0153] The actual action data and the environmental change information can be compared with the generated decision, i.e., the action sequence.
[0154] Incremental learning algorithms can be used to incrementally update the motion decision model of the dual-arm robot.
[0155] Through the above embodiments, while retaining existing knowledge, the model learns the unique characteristics and decision-making patterns in new environments and tasks, avoiding catastrophic forgetting; by adjusting model parameters, the model can better adapt to new scenarios, thereby continuously expanding its knowledge reserves and decision-making capabilities.
[0156] In this embodiment, after controlling the dual-arm robot to execute the action sequence, a multi-dimensional model evaluation system is constructed;
[0157] The decision-making performance of the dual-arm robot motion decision-making model in the new environment is evaluated based on the multi-dimensional model evaluation system, and the evaluation results are obtained.
[0158] Based on the evaluation results, an adjustment strategy is generated for the motion decision model of the dual-arm robot.
[0159] The evaluation indicators in the multi-dimensional model evaluation system may include, but are not limited to, one or more of the following indicators:
[0160] Task completion rate, decision-making accuracy, smoothness of action, execution efficiency, etc.
[0161] Among these, the decision-making performance of the model in a new environment can be evaluated based on the multi-dimensional model evaluation system, and the advantages and disadvantages of the model can be analyzed.
[0162] After obtaining the evaluation results, the reinforcement learning reward function can be adjusted, the multimodal feature fusion method can be optimized, the knowledge transfer and incremental learning strategies can be improved, and the model can be retrained and tested regularly to ensure that the model continues to learn and remains available.
[0163] For example, in the medical field, if an evaluation determines that a dual-arm robot's movements are not smooth enough when assisting in the delivery of surgical tools, the model can be optimized and iterated to make the dual-arm robot's movements more stable and efficient during the surgical process, thereby improving the quality of care.
[0164] For example, in the financial sector, when an assessment determines that a collision occurs while a dual-arm robot is moving items in a service hall, the model can be optimized to make the path of the dual-arm robot more reasonable during the assisted handling process.
[0165] Through the above embodiments, the model can continuously improve its generalization performance during the continuous learning process, enabling it to stably and efficiently control the robot's arms through reasonable decision-making in various complex environments.
[0166] As can be seen from the above technical solutions, this invention can utilize the multiple types of sensors equipped on a dual-arm robot to collect current environmental data to train a visual language action model, enabling the model to initially possess basic generalization capabilities. By utilizing the feature extraction module of the dual-arm robot's action decision model, multimodal features of real-time environmental information are dynamically extracted based on an attention mechanism to quickly capture multimodal feature representations adapted to dynamic environments, improving the model's perception of new environments and thus enhancing its generalization ability. The parsed task objectives, constraints, and multimodal features are input into the policy network of the dual-arm robot's action decision model to obtain the action sequence of the dual-arm robot, and the dual-arm robot is controlled to execute the action sequence, thereby achieving accurate understanding and decision-making regarding tasks in new environments and improving the model's adaptability to new environments.
[0167] like Figure 3 The diagram shown is a schematic diagram of the computer device for implementing the dual-arm robot control method of the present invention.
[0168] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a dual-arm robot control program.
[0169] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0170] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0171] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code of a dual-arm robot control program, but also to temporarily store data that has been output or will be output.
[0172] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing a dual-arm robot control program) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0173] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes these applications to implement the steps in the various embodiments of the dual-arm robot control method described above, for example... Figure 1 The steps are shown.
[0174] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a training unit 110, a data acquisition unit 111, an extraction unit 112, a parsing unit 113, a decision-making unit 114, and a control unit 115.
[0175] The integrated unit, implemented as a software functional module, can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, a computer device, or a network device, etc.) or processor to execute portions of the dual-arm robot control method described in the various embodiments of the present invention.
[0176] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0177] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0178] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0179] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0180] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0181] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0182] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.
[0183] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0184] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0185] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0186] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a dual-arm robot control method, and the processor 13 can execute the multiple instructions to achieve:
[0187] The current environmental data is collected by using multiple types of sensors equipped on the dual-arm robot, and the current environmental data is used to train a visual language action model to obtain a dual-arm robot action decision model.
[0188] In response to control commands to the dual-arm robot in the new environment, real-time environmental information of the new environment is collected through the various types of sensors;
[0189] Using the feature extraction module of the dual-arm robot motion decision model, multimodal features of the real-time environmental information are dynamically extracted based on an attention mechanism;
[0190] Using the natural language processing module of the dual-arm robot motion decision model, the control commands are parsed based on the multimodal features to obtain the task objectives and constraints;
[0191] The task objective, the constraints, and the multimodal features are input into the policy network of the dual-arm robot motion decision model to obtain the motion sequence of the dual-arm robot.
[0192] Control the dual-arm robot to execute the sequence of actions.
[0193] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0194] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0195] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0196] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0197] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0198] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0199] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0200] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0201] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A control method for a dual-arm robot, characterized in that, The dual-arm robot control method includes: A dual-arm robot collects current environmental data using multiple types of sensors and trains a visual-language-action model using this data to obtain a dual-arm robot action decision model. The process includes: extracting features from image data in the current environmental data using a Swing Transformer model to obtain a first feature; encoding text data in the current environmental data using a BERT model to obtain a second feature; extracting Mel-frequency cepstral coefficients from audio data in the current environmental data and processing these coefficients using a recurrent neural network to obtain a third feature; processing dual-arm motion and tactile data in the current environmental data using a multilayer perceptron to obtain a fourth feature; deeply fusing the first, second, third, and fourth features using a self-attention mechanism to obtain a fused feature; and training the visual-language-action model using the fused feature to obtain the dual-arm robot action decision model. In response to control commands to the dual-arm robot in the new environment, real-time environmental information of the new environment is collected through the various types of sensors; Using the feature extraction module of the dual-arm robot motion decision model, multimodal features of the real-time environmental information are dynamically extracted based on an attention mechanism; Using the natural language processing module of the dual-arm robot motion decision model, the control commands are parsed based on the multimodal features to obtain the task objectives and constraints; The task objective, the constraints, and the multimodal features are input into the policy network of the dual-arm robot motion decision model to obtain the motion sequence of the dual-arm robot. Control the dual-arm robot to execute the sequence of actions.
2. The dual-arm robot control method as described in claim 1, characterized in that, The feature extraction module utilizing the dual-arm robot motion decision model dynamically extracts multimodal features of the real-time environmental information based on an attention mechanism, including: The environmental characteristics of the new environment are determined based on the real-time environmental information; Based on the environmental characteristics, the weights of each modality are dynamically adjusted according to the attention mechanism. The feature extraction module is used to extract the multimodal features according to the weight of each modality.
3. The dual-arm robot control method as described in claim 1, characterized in that, Before inputting the task objective, the constraints, and the multimodal features into the policy network of the dual-arm robot motion decision model, the method further includes: Obtain a pre-built knowledge base; wherein the knowledge base is used to store the correspondence between environmental information, task information and action decisions; The task objective, the constraints, and the multimodal features are used to perform similarity retrieval in the knowledge base to obtain similar knowledge. The action sequence is generated based on the similarity knowledge.
4. The dual-arm robot control method as described in claim 1, characterized in that, The control of the dual-arm robot to execute the action sequence includes: The dual-arm robot is controlled to move according to the action sequence, and during the movement, the dual-arm robot is controlled to avoid obstacles based on the constraints.
5. The dual-arm robot control method as described in claim 1, characterized in that, The method further includes: During the process of controlling the dual-arm robot to execute the action sequence, the execution deviation is continuously detected, and the actual action data and environmental change information of the dual-arm robot are continuously collected; Feedback data is generated based on the execution deviation; The dual-arm robot motion decision model is incrementally updated based on the feedback data, the actual motion data, and the environmental change information.
6. The dual-arm robot control method as described in claim 1, characterized in that, After controlling the dual-arm robot to execute the action sequence, the method further includes: Construct a multi-dimensional model evaluation system; The decision-making performance of the dual-arm robot motion decision-making model in the new environment is evaluated based on the multi-dimensional model evaluation system, and the evaluation results are obtained. Based on the evaluation results, an adjustment strategy is generated for the motion decision model of the dual-arm robot.
7. A control device for a dual-arm robot, characterized in that, The dual-arm robot control device includes: The training unit is used to collect current environmental data using multiple types of sensors equipped on the dual-arm robot, and to train a visual language action model using the current environmental data to obtain a dual-arm robot action decision model. The training includes: extracting features from image data in the current environmental data using a SwingTransformer model to obtain a first feature; encoding text data in the current environmental data using a BERT model to obtain a second feature; extracting Mel-frequency cepstral coefficients from audio data in the current environmental data and processing the Mel-frequency cepstral coefficients using a recurrent neural network to obtain a third feature; processing dual-arm motion and tactile data in the current environmental data using a multilayer perceptron to obtain a fourth feature; deeply fusing the first, second, third, and fourth features using a self-attention mechanism to obtain a fused feature; and training the visual language action model using the fused feature to obtain the dual-arm robot action decision model. The data acquisition unit is used to respond to control commands to the dual-arm robot in a new environment by acquiring real-time environmental information of the new environment through the multiple types of sensors. The extraction unit is used to dynamically extract multimodal features of the real-time environmental information based on the attention mechanism using the feature extraction module of the dual-arm robot motion decision model. The parsing unit is used to use the natural language processing module of the dual-arm robot motion decision model to parse the control instructions based on the multimodal features to obtain the task objective and constraints. The decision unit is used to input the task objective, the constraints and the multimodal features into the policy network of the dual-arm robot action decision model to obtain the action sequence of the dual-arm robot. A control unit is used to control the dual-arm robot to execute the action sequence.
8. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the dual-arm robot control method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the dual-arm robot control method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Brain-like multi-mode emotion recognition network, brain-like multi-mode emotion recognition method and emotion robot
CN115169507A
Human attention mechanism imitation learning method in social scene
CN115994576A