Multi-modal feedback integration type knowledge-search enhanced robot control system

The multimodal robot control system integrates language and sensorimotor feedback to enhance adaptability and execution of complex tasks in unpredictable environments, addressing limitations of conventional systems by enhancing task decomposition, environmental understanding, and safety.

JP2025108598APending Publication Date: 2025-07-23NYU-YO-KU ZENERAL GURU-PU INKU
View PDF 0 Cites 15 Cited by

Patent Information

Application Number
JP2025067322
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Conventional robot systems struggle with executing complex tasks in unpredictable environments due to limited adaptability to environmental changes and uncertainties, difficulty in understanding and executing high-level abstract instructions, insufficient integration of visual and force feedback, inefficient utilization of knowledge bases, and limited ability to plan and execute long-term tasks.

Method used

A multimodal feedback integrated robot control system that combines a large language model, retrieval-augmented generation technology, vision system, force sensing module, speech recognition/synthesis module, and robot control system to enhance adaptability, understanding, and execution of complex tasks.

Benefits of technology

The system effectively decomposes high-level instructions into subtasks, integrates multimodal feedback for richer environmental understanding, adapts to changes and uncertainties, and ensures scalability and safety, enabling the execution of complex and long-term tasks.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To realize improvement of complicated task execution ability under unpredictable circumstance, understanding and execution of high-level abstract instructions, effective integration of multi-modal feedback, enhancement of ability to adapt itself to change of environment and uncertainty, efficient use of knowledge base, and securing expandability.SOLUTION: The present invention is directed to a robot control system which integrates a large-scale language model, retrieval-enhanced generation technique, a visual system, a force sense module, a voice recognition and synthesizing module, a multi-modal integration module, and a robot control system. A language processing component recognizes high-level instructions, and dynamically selects and adapts relevant examples from a knowledge base by using the retrieval-enhanced generation technique. The visual system generates a three-dimensional expression of an environment. The force sense module measures force of an end effector. The multi-modal integration module integrates information from different modalities to realize environment understanding and adaptive action generation. Thus, it is possible to adapt the system to execution of complicated tasks and change of the environment under an unpredictable environment.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and robotics, and more particularly to a robot control system that integrates visual and force feedback using large language models (LLMs) and retrieval-augmented generation (RAG) technology. More specifically, it relates to the realization of a robot system that can understand high-level instructions in human natural language and execute complex long-term tasks in an uncertain and unpredictable environment. The present invention is inspired by the concept of "embodied cognition" in the human cognitive process, and improves the adaptability and intelligence of robots by integrating language processing capabilities and sensorimotor feedback.

Background Art

[0002] In recent years, in the fields of artificial intelligence and robotics, there has been a demand for improving the ability of robots to execute complex tasks. Conventional robot systems rely on pre-programmed responses and have limited adaptability to environmental changes and uncertainties. In particular, it has been difficult to execute complex tasks in unpredictable environments such as homes and medical sites. For example, even a seemingly simple task such as "making coffee" requires dealing with many subtasks and environmental uncertainties, such as identifying the position of the mug, opening and closing the drawer, scooping coffee beans, and adjusting the amount of hot water to be poured.

[0003] Human intelligence is often understood as "embodied cognition," and cognitive processes such as attention, language, learning, memory, and perception are essentially linked to how the body interacts with the surrounding environment. According to the research of Lakoff & Johnson (1999), the human conceptual system is rooted in bodily experience, and even abstract thinking is constructed through bodily metaphors. Also, Wilson (2002) argues that cognition is formed through interaction with the environment and is not limited to abstract information processing in the brain. Furthermore, Shapiro (2011) has shown that the morphology and motor abilities of the body affect cognitive processes. These studies provide evidence that human intelligence is based on sensorimotor processes, which gives important implications for the design of machine intelligence.

[0004] In the prior art, the development of the robot's sensorimotor capabilities and artificial intelligence has progressed in parallel, but attempts to effectively integrate them have been limited. On the one hand, research focusing on improving the robot's sensorimotor capabilities has advanced, and technologies such as object recognition and grasping using visual feedback (Levine et al., 2018), and precise manipulation using force feedback (Hogan, 1985; Siciliano & Khatib, 2016) have been developed. On the other hand, in the field of artificial intelligence, the development of large language models (LLMs) has enabled improvements in natural language understanding, generation, and inference capabilities (Brown et al., 2020; Devlin et al., 2019). However, the development of a robot system that integrates these technologies and combines language understanding capabilities with sensorimotor feedback has faced many technical challenges and has not advanced sufficiently.

[0005] Approaches such as reinforcement learning and imitation learning have been shown to be effective in executing complex tasks. For example, research has been carried out on robot motion control using reinforcement learning (Schulman et al., 2017; Haarnoja et al., 2018) and skill acquisition using imitation learning (Finn et al., 2016; Ho & Ermon, 2016). However, these approaches have challenges in adapting to new tasks and diverse scenarios. Reinforcement learning requires a large amount of data and trial-and-error, and learning in the real world is time-consuming and resource-intensive. In addition, since imitation learning depends on human demonstrations, it may be difficult to generalize to new situations.

[0006] In recent years, research on robot control using large language models (LLMs) has been progressing. For example, VoxPoser (Huang et al., 2023) utilizes an LLM to execute daily operation tasks, and Robotics Transformer (RT-2) (Brohan et al., 2023) has demonstrated the adaptability to execute tasks beyond training scenarios by leveraging large-scale web data and robot learning data. Also, Hierarchical diffusion policy (Chi et al., 2023) has introduced a model structure that generates context-responsive motion trajectories to enhance task-specific motions from high-level LLM decision inputs.

[0007] However, these approaches have issues such as complex prompt requirements, lack of real-time feedback, insufficient utilization of force feedback, and inefficient pipelines that hinder task execution. In particular, in robot control using large language models (LLMs), the design of prompts is complex, and desirable results are often not obtained without appropriate instructions. Also, there is a lack of a feedback mechanism to respond to real-time environmental changes, and the adaptability in unpredictable situations is limited. Furthermore, in many approaches, emphasis is placed on visual feedback, and the utilization of force feedback is insufficient. As a result, the adaptability to physical interactions that cannot be captured by visual information alone (e.g., adjustment of the injection volume of a liquid, adjustment of the force required to open and close a door) has been restricted.

[0008] Also, the application of retrieval-augmented generation (RAG) technology to robotics has not been fully explored despite its potential. RAG is a technology that retrieves relevant information from an external knowledge base and incorporates it into the generation process to improve the output of language models (Lewis et al., 2020). This technology has been shown to reduce hallucinations (generation of information different from facts) in language models and generate more accurate and reliable outputs. However, the application of RAG to robot control, especially to motion generation for complex task execution, has been limited.

[0009] Against this background, there is a need to develop a robot system that effectively integrates language understanding ability and sensorimotor feedback to enable complex task execution in unpredictable environments. The present invention provides a robot control system that integrates a large language model (LLM), retrieval-augmented generation (RAG) technology, visual and force feedback to address these issues.

Prior Art Documents

Non-Patent Documents

[0010]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0011] The present invention aims to solve the following problems: 1. Improvement in the ability to execute complex tasks in unpredictable environments: Conventional robot systems rely on pre-programmed responses and have limited adaptability to environmental changes and uncertainties. In particular, it has been difficult to execute complex tasks in unpredictable environments such as homes and medical sites. The present invention aims to improve the ability to execute complex tasks while coping with environmental changes and uncertainties. 2. Understanding and execution of high-level abstract instructions: Conventional robot systems require specific operation instructions and have difficulty understanding and decomposing abstract instructions such as "make coffee" into appropriate subtasks for execution. The present invention aims to improve the ability to understand high-level abstract instructions in human natural language and decompose them into appropriate subtasks for execution. 3. Effective integration of visual and force feedback: In conventional robot systems, emphasis has been placed on visual feedback, and the utilization of force feedback has been insufficient. Also, the mechanisms for integrating and utilizing these feedbacks have been limited. The present invention aims to effectively integrate visual and force feedback and improve the accuracy and adaptability of interaction with the environment. 4. Enhancement of the ability to adapt to environmental changes and uncertainties: Conventional robot systems have limited adaptability to environmental changes and uncertainties and have difficulty coping with unexpected situations (e.g., movement of objects, appearance of obstacles). The present invention aims to incorporate real-time feedback and enhance the adaptability to environmental changes and uncertainties. 5. Efficient utilization of the knowledge base and ensuring scalability: In conventional robot systems, the utilization of the knowledge base is limited, and it is difficult to add or update new knowledge. Also, the performance degradation associated with the expansion of the knowledge base has been an issue. The present invention aims to efficiently utilize the knowledge base using Retrieval-Augmented Generation (RAG) technology and ensure the scalability of the system. 6. Planning and execution of long-term tasks: Conventional robot systems focus on the execution of short-term tasks and have difficulty planning and executing long-term tasks composed of multiple subtasks. The present invention aims to improve the ability to plan and execute long-term tasks considering the dependencies between subtasks. 7. Generation of creative actions: Conventional robot systems rely on pre-programmed action patterns and have limited ability to generate creative actions in response to new situations or requirements. The present invention aims to improve the ability to generate new action patterns by combining a language model and a visual generation model. 8. Ensuring safety and reliability: With the advancement of robot systems, ensuring safety and reliability has become an important issue. The present invention aims to incorporate safety constraints and ensure the safety and reliability of the system. 9. Integration and processing of multimodal information: In conventional robot systems, the ability to effectively integrate and process information from different modalities (vision, force sense, language, etc.) has been limited. The present invention aims to integrate multimodal information and achieve richer environmental understanding and adaptive behavior generation. 10. Natural interaction with humans: In conventional robot systems, natural interaction with humans has been limited. The present invention aims to realize natural interaction with humans through natural language understanding and generation, and recognition and generation of non-verbal communication (gestures, expressions, etc.). By solving these problems, the present invention aims to contribute to the realization of an "intelligent robot" that can cooperate with humans and execute complex tasks.

Means for Solving the Problems

[0012] The present invention provides a multimodal feedback integrated knowledge retrieval enhanced robot control system (hereinafter, this system). This system integrates a large language model (LLM), a retrieval augmented generation (RAG) infrastructure, a vision system, a force sense module, a speech recognition / synthesis module, and a robot control system. The main components of this system are as follows: 1. Language processing component: The language processing component is responsible for processing user queries and environmental data and decomposing them into a series of steps capable of executing complex tasks. This component uses a large language model (e.g., GPT-4) to perform natural language understanding and generation, task decomposition and planning, and code generation. Additionally, it uses retrieval augmented generation (RAG) technology to dynamically select and adapt relevant examples from a curated knowledge base. The main functions of the language processing component are as follows: - Understanding and interpreting natural language: Understanding the user's instructions and extracting their intentions and goals. - Task decomposition and planning: Decompose complex tasks into executable subtasks and make a plan considering the dependencies between subtasks. - Code generation: Generate the code required for the execution of subtasks and send it to the robot control system. - Knowledge retrieval and adaptation: Use RAG to retrieve relevant examples from the knowledge base and adapt them to the current task. - Integration of multimodal information: Integrate information from different modalities such as vision, force sense, and voice to achieve a richer understanding of the environment. - Dialogue management: Manage the dialogue with the user and ask questions or confirmations as necessary. The language processing component represents the dependencies between subtasks as conditional probabilities and optimizes the planning and execution of tasks. For example, P(L2A, L2B|L1) specifies the probability of proceeding to task L2A or L2B after the successful execution of task L1. This allows for dynamically selecting the next subtask according to the progress of the task. 2. Vision system: The vision system is responsible for generating a three-dimensional representation of the environment and identifying the positions of objects. This system uses a depth camera (e.g., Azure Kinect DK) to obtain three-dimensional information about the environment and identifies objects based on language instructions through a language-vision module (e.g., Grounded-Segment-Anything). The main functions of the vision system are as follows: - Generation of a three-dimensional representation of the environment: Use a depth camera to generate a three-dimensional representation of the environment. - Object detection and segmentation: Use a language-vision module to detect and segment objects. - Pose estimation of objects: Estimate the position and orientation of the detected objects. - Object tracking: Track the movement of objects and update the position information in real time. - Occlusion handling: Estimate the position of an object even when it is hidden by other objects. - Detection of environmental changes: Detect changes in the environment (e.g., movement of objects, appearance of new objects). The vision system performs object segmentation and pose estimation, providing information necessary for the robot's grasping and operation. It also detects changes in the environment (e.g., movement of objects) in real time and reflects them in the robot's actions. 3. Force sensing module: The force sensing module measures the forces received by the robot's end effector and plays a role in improving the accuracy of object manipulation. This module uses a multi-axis force sensor (e.g., ATI force sensor) to measure the forces and torques received by the robot's end effector. The main functions of the force sensing module are as follows: - Measurement of force and torque: Use a multi-axis force sensor to measure the forces and torques received by the robot's end effector. - Conversion and interpretation of force: Convert the measured force to the robot's base coordinate system and interpret its meaning. - Force control: Control to apply appropriate forces in the interaction with objects. - Contact detection: Detect contact with objects and determine their states. - Estimation of object characteristics: Estimate object characteristics such as weight, hardness, and friction from force feedback. - Anomaly detection: Detect unexpected force changes and take safety measures. The force sensing module uses force feedback to improve the accuracy of object manipulation (e.g., liquid injection volume, opening and closing of drawers). Also, even in situations where visual information is limited (e.g., when the field of view is blocked), operations can be continued using force feedback. 4. Speech recognition and synthesis module: The speech recognition and synthesis module is responsible for recognizing the user's voice instructions and outputting the system's responses in voice. This module consists of a speech recognition engine and a speech synthesis engine. The main functions of the speech recognition and synthesis module are as follows: - Speech recognition: Recognize the user's voice instructions and convert them into text. - Speech synthesis: Convert the system's response from text to voice. - Speaker identification: Identify the voices of multiple users and provide personalized responses. - Emotion recognition: Recognize the emotion from the user's voice and provide appropriate responses. - Ambient sound recognition: Recognize ambient sounds (e.g., warning sounds, door opening and closing sounds) to understand the situation. - Speech dialogue management: Manage the speech dialogue with the user to achieve natural conversations. The speech recognition and synthesis module enables natural interaction with the user and allows for hands-free operation. Additionally, by combining with visual information, it enables multimodal dialogue. 5. Robot control system: The robot control system integrates feedback from the language processing, vision system, force sensing module, and speech recognition and synthesis module, and is responsible for controlling the robot's movements. This system is based on ROS (Robot Operating System) and provides a control mechanism incorporating safety constraints. The main functions of the robot control system are as follows: - Integration of multimodal feedback: Integrate feedback from different modalities and reflect it in the robot's movements. - Motion planning and execution: Plan and execute the movements necessary for task execution. - Implementation of safety constraints: Implement safety constraints such as maximum speed and force limits, and boundaries of the working space. - Obstacle avoidance: Detect obstacles in the environment and plan a path to avoid them. - Error detection and recovery: Detect errors during operation and perform appropriate recovery processing. - Adaptive control: Adapt to changes and uncertainties in the environment and adjust control parameters. The robot control system incorporates safety constraints, sets limits on maximum speed and force, and defines the boundaries of the working space. This ensures the safety and reliability of the system. Additionally, based on real-time feedback, the operations can be adjusted. 6. Knowledge Base: The knowledge base is a curated database that includes examples of verified low-level and high-level actions. This database systematically organizes code examples, motion primitives, examples of task decomposition, etc. The main components of the knowledge base are as follows: - Basic motion primitives: Basic operations such as linear motion, rotational motion, opening and closing of the gripper. - Motion primitives for object manipulation: Operations related to object manipulation such as grasping, placing, lifting, and transporting. - Special motion primitives: Special operations such as liquid injection, powder scooping, opening and closing of drawers, drawing. - High-level task examples: Complex tasks such as making coffee, decorating dishes, passing objects. - Error handling and recovery examples: Examples for dealing with and recovering from errors during operations. - Environment adaptation examples: Examples for adapting to environmental changes and uncertainties. The knowledge base includes motion examples incorporating known uncertainties and can handle various scenarios. Also, it has a structure that allows for easy addition and update of new knowledge, ensuring the extensibility of the system. 7. Multimodal Integration Module: The multimodal integration module is responsible for integrating information from different modalities such as vision, force sensing, and voice, and realizing richer environmental understanding and adaptive behavior generation. This module provides a mechanism to appropriately weight and integrate information from each modality. The main functions of the multimodal integration module are as follows: - Information integration across modalities: Integrate information from different modalities to generate a consistent environmental representation. - Modality weighting: Dynamically adjust the importance of each modality according to the task and situation. - Cross-modal learning: Learn the relationships between different modalities to achieve information complementarity. - Multimodal inference: Make inferences about the environment and situation based on the integrated information. - Uncertainty management: Consider the uncertainty of each modality and prioritize more reliable information. - Completion of missing information: When information from some modalities is missing, use information from other modalities to complement it. The multimodal integration module performs processes such as identifying the position of an object from visual information, estimating the weight of the object from haptic information, and integrating this information to determine the appropriate grasping force and position. This enables more precise and adaptable object manipulation. The operation flow of this system is as follows: 1. The user gives a high-level instruction (e.g., "Make coffee") in voice or text. 2. The speech recognition / synthesis module converts the voice instruction into text and sends it to the language processing component. 3. The language processing component interprets the instruction and combines it with environmental data (images from visual sensors) to decompose the task into a series of subtasks (e.g., "Find the mug", "Scoop the coffee", "Pour the hot water"). 4. For each subtask, use the RAG to search for relevant examples from the knowledge base. At this time, based on vector embedding and cosine similarity, select the most relevant example. 5. Based on the retrieved examples, the language model generates executable code. This code includes function calls for using visual and haptic feedback. 6. The generated code is sent to the robot control system and executed. 7. The robot utilizes multi-modal feedback such as vision, force sense, and voice to execute tasks based on the generated code. - Using vision feedback to identify the position of an object and plan grasping and operation. - Using force sense feedback to control the interaction with an object (e.g., the injection volume of a liquid, the opening and closing of a drawer). - Using voice feedback to continue the interaction with the user and ask questions or confirm as necessary. 8. The multi-modal integration module integrates information from different modalities to achieve a richer environmental understanding and generate adaptive behaviors. 9. Incorporate real-time feedback and adjust operations in response to changes and uncertainties in the environment. - When an object moves, use vision feedback to identify the new position and adjust the operation. - When an unexpected obstacle appears, take avoidance actions. - When the force sense feedback is different from the prediction (e.g., a drawer is heavier than expected), adjust the force. - When the user gives a new instruction, update the task plan.

[0013] 10. After the task is completed, report the result to the user and wait for the next instruction. The characteristic points of this system are as follows: 1. Efficient utilization of the knowledge base using RAG: This system ensures the scalability of the system by efficiently utilizing the knowledge base using RAG. In the conventional approach, since the content of the knowledge base was statically incorporated into the context window of the language model, there was a problem of performance degradation as the knowledge base expanded. In this system, since relevant examples are dynamically selected from the knowledge base using RAG, performance can be maintained even when the knowledge base expands. 2. Integration of Multimodal Feedback: This system realizes richer environmental understanding and adaptive behavior generation by integrating feedback from different modalities such as vision, force sense, and voice. Each modality provides different types of information, and by appropriately integrating them, it is possible to handle complex situations that cannot be captured by a single modality. 3. Creative Motion Generation: This system can generate new motion patterns by combining a language model and a visual generation model. For example, creative tasks such as generating a silhouette based on keywords using a visual generation model such as DALL-E, extracting its outline, and drawing it on a physical surface can be executed. 4. Incorporation of Safety Constraints: This system ensures the safety and reliability of the system by incorporating safety constraints. Constraints such as maximum speed, force limits, and the boundaries of the working space are encoded in the basic motion primitives, so errors in the language model do not override these constraints. 5. Natural User Interaction: This system realizes natural interaction with the user through the speech recognition / synthesis module. The user can give instructions in natural language, and the system can also respond in natural language. Also, during the execution of a task, interaction with the user can be continued, and questions and confirmations can be made as needed. 6. Hardware-independent design: This system has a design that is independent of specific hardware and is applicable to various robot platforms. Also, since the RAG infrastructure can be flexibly selected, it is possible to use open-source tools such as Haystack and vebra, or the RAG process of OpenAI. Due to these features, this system improves the ability to execute complex tasks in unpredictable environments, and can solve problems such as understanding and executing high-level abstract instructions, effectively integrating visual and force feedback, enhancing the ability to adapt to environmental changes and uncertainties, efficiently utilizing and ensuring the scalability of the knowledge base, planning and executing long-term tasks, generating creative actions, ensuring safety and reliability, integrating and processing multimodal information, and having natural interaction with humans.

Advantages of the Invention

[0014] The present invention provides the following advantages: 1. Understanding and execution of high-level abstract instructions: This system can understand high-level abstract instructions in the user's natural language (e.g., "Make coffee", "I'm tired, so make a hot drink") and break them down into appropriate subtasks for execution. This enables the user to execute tasks just by conveying the purpose without giving specific operation instructions. For example, for a complex instruction like "A friend who is tired and will come to eat cake soon, so make a hot drink and draw a random animal on a plate", the system can select coffee as "a hot drink for a tired person", find a mug, scoop the coffee, pour hot water, and further break it down into a series of subtasks such as drawing an animal picture on the plate for execution. This effect is achieved by the natural language understanding ability of large language models (LLMs) and the efficient utilization of knowledge bases through retrieval-augmented generation (RAG) technology. LLMs can understand context and intent and break down abstract instructions into specific subtasks. Additionally, by using RAG, relevant examples can be retrieved from the knowledge base and adapted to the current task. As a result, the system can handle various abstract instructions, break them down into appropriate subtasks, and execute them. As a specific example, when the user instructs "Prepare breakfast", the system can consider environmental data (e.g., the contents of the refrigerator, the availability of cooking utensils) and break it down into subtasks such as "Make toast", "Cook eggs", "Brew coffee", etc., and execute each subtask in an appropriate order. Also, for the instruction "Tidy up the room", it can be broken down into subtasks such as "Pick up the items scattered on the floor", "Organize the tabletop", "Throw away the garbage", etc., and executed. This high-level ability to understand and execute abstract instructions significantly improves the usability and practicality of the robot. Users can give instructions to the robot in natural language without the need for specialized knowledge or detailed instructions. This enables various users, such as the elderly, disabled, and children, to easily use the robot. 2. Completion of Long-Term Tasks: This system can consistently execute long-term tasks composed of multiple subtasks. By representing the dependency relationships between subtasks as conditional probabilities and optimizing task planning and execution, it becomes possible to efficiently complete complex tasks. For example, in the task of making coffee, a plan can be made considering the dependency that if the mug is not found, open the drawer. Also, the plan can be adjusted according to the progress of the task to handle unexpected situations (e.g., when there is a shortage of coffee beans). This effect is achieved by the task decomposition and planning ability of the language processing component, as well as the adaptive control ability of the robot control system. The language processing component decomposes tasks into executable subtasks and makes plans considering the dependencies between subtasks. Also, the robot control system adapts to changes and uncertainties in the environment and adjusts control parameters. As a result, the system can consistently execute long-term tasks and handle unexpected situations. As a specific example, considering the long-term task of "preparing dinner", the system decomposes this into a number of subtasks such as "taking ingredients out of the refrigerator", "washing vegetables", "cutting vegetables", "cooking", "preparing tableware", and "plating the dish". Each subtask has dependencies. For example, "cutting vegetables" needs to be executed after "washing vegetables" is completed. The system makes plans considering these dependencies and executes tasks efficiently. Also, if it is found that there is a shortage of ingredients during the process, the system can adjust the plan and take actions such as using alternative ingredients or asking the user for confirmation. This ability to complete such long-term tasks greatly improves the practicality and usefulness of the robot. The robot can not only repeat simple operations but also consistently execute complex and multi-step tasks. As a result, it is expected that the utilization of robots will progress in various application fields such as housework support, caregiving support, and complex assembly work in manufacturing. 3. Integration and Utilization of Multimodal Feedback: This system realizes richer environmental understanding and adaptive behavior generation by integrating feedback from different modalities such as vision, force sense, and voice. It can recognize the position and shape of an object using visual feedback, control the physical interaction with the object using force sense feedback, and continue the dialogue with the user using voice feedback. This makes it possible to handle complex situations that cannot be grasped by a single modality. This effect is achieved through the cooperation of the visual system, force sensing module, speech recognition and synthesis module, and multimodal integration module. The visual system performs object detection and segmentation, and pose estimation. The force sensing module measures force and torque and performs force control. The speech recognition and synthesis module performs speech recognition and synthesis and dialogue management. The multimodal integration module integrates this information and generates a consistent environmental representation. As a result, the system can achieve a richer environmental understanding and generate adaptive behaviors. As a specific example, considering the task of pouring a liquid, the system uses visual feedback to identify the position of the cup and force feedback to control the pouring amount. If the cup moves, visual feedback is used to identify the new position and adjust the operation. Also, even if the cup is hidden from view during pouring, the pouring amount can be continuously monitored using force feedback and pouring can be stopped at an appropriate timing. Furthermore, if the user gives an instruction such as "pour a little more" verbally, the instruction can be understood using speech feedback and additional liquid can be poured. This ability to integrate and utilize such multimodal feedback greatly improves the robot's environmental understanding and adaptability. The robot can go beyond the limitations of a single sensor and integrate multiple senses like a human to understand the environment and act adaptively. As a result, it becomes possible to execute complex tasks in an unpredictable environment, and the application range of the robot is significantly expanded. 4. Adaptation to environmental changes and uncertainties: This system can incorporate real-time feedback and improve its adaptability to environmental changes and uncertainties. When an object moves, its new position can be identified using visual feedback and the operation can be adjusted. Also, when an unexpected obstacle appears, avoidance behavior can be taken. Furthermore, when the force feedback is different from what is expected (e.g., a drawer is heavier than expected), the force can be adjusted. As a result, it becomes possible to continue executing the task even in an unpredictable environment. This effect is realized by the environmental change detection ability of the vision system, the force control ability of the force sensing module, and the adaptive control ability of the robot control system. The vision system detects environmental changes in real time, and the force sensing module appropriately controls the interaction with the object. Based on these feedbacks, the robot control system dynamically adjusts the control parameters. As a result, the system can adapt to environmental changes and uncertainties and continue to execute tasks. As a specific example, considering the task of tidying up an object on a table, the system uses visual feedback to identify the position of the object, grasp it, and place it at the designated location. If the object moves during the process, the system uses visual feedback to identify the new position and adjusts the motion plan. Also, if the object is heavier than expected, the system can use force feedback to detect its weight and take corresponding actions such as increasing the grasping force or lifting it with both hands. Furthermore, if an obstacle appears at the placement location, the system can use visual feedback to detect the obstacle and take actions such as avoiding it or placing it in another location. This ability to adapt to environmental changes and uncertainties greatly improves the usefulness of robots in the real world. Since the real world is constantly changing and full of uncertainties, robots that can adapt to environmental changes and uncertainties can perform more practical tasks. As a result, it is expected that the utilization of robots in various fields such as home, medical, and manufacturing will progress. 5. Efficient utilization and scalability of the knowledge base: This system can efficiently utilize the knowledge base using RAG and ensure the scalability of the system. In the conventional approach, since the content of the knowledge base was statically incorporated into the context window of the language model, there was a problem of performance degradation as the knowledge base expanded. In this system, since relevant examples are dynamically selected from the knowledge base using RAG, performance can be maintained even as the knowledge base expands. This makes it easier to add and update new knowledge and enables continuous improvement of the system's capabilities. According to the experimental results, it has been confirmed that by using RAG, the faithfulness score of GPT-4 improves from 0.74 to 0.88, and the faithfulness score of GPT-3.5-turbo improves from 0.78 to 0.86. This effect is realized by the RAG function of the language processing component. RAG converts the user's query and the current task state into vector embeddings, calculates the cosine similarity with each chunk of the knowledge base, and selects the most relevant chunk. The selected chunk is added to the input of the language model, enabling the generation of more accurate and relevant responses. Thereby, the system can efficiently utilize the knowledge base and ensure scalability. As a specific example, consider the case of adding a new task (e.g., "making pancakes") to the system. In the conventional approach, it was necessary to incorporate all the knowledge about pancake making into the context window of the language model, which might exceed the limit of the context window when combined with the existing knowledge. In this system, simply adding the knowledge about pancake making to the knowledge base allows RAG to search for relevant knowledge as needed and provide it to the language model. Thereby, the performance of the system can be maintained even as the knowledge base expands. The efficient utilization and scalability of such a knowledge base significantly improve the learning and adaptation capabilities of the robot. The robot can continuously add new knowledge and become capable of handling various tasks. As a result, the application scope of the robot expands, enabling it to perform more practical tasks.

[0015] 6. Creative Motion Generation: By combining a language model and a visual generation model, this system can generate new motion patterns. For example, using a visual generation model such as DALL-E, silhouettes can be generated based on keywords (e.g., "random bird", "random plant"), and their outlines can be extracted and drawn on a physical surface (e.g., a plate). This enables the execution of artistic motions (e.g., drawing, decoration) based on user instructions. Additionally, this technology can be extended to applications such as cake decoration and latte art. This effect is achieved by the creative motion generation ability of the language processing component and the image generation ability of the visual generation model. The language processing component extracts keywords from user instructions and inputs them into the visual generation model. The visual generation model generates an image based on the keywords and extracts a drawing trajectory from the image. The robot control system controls the robot's motion along the extracted trajectory. Thus, the system can generate and execute creative motions. As a specific example, when the user instructs "draw a picture of a cat on a plate", the system extracts the keyword "cat" and inputs it into DALL-E. DALL-E generates a silhouette image of a cat, and the system extracts the outline from the image. The extracted outline is converted according to the dimensions of the plate, and the robot uses a pen to draw along the outline. Also, for an instruction like "write 'Happy Birthday' on the cake and decorate it with flowers", the system can extract the keywords "Happy Birthday" and "flowers", generate a drawing combining the text and flower decoration, and execute it. Such creative motion generation capabilities significantly expand the expressiveness and application range of robots. Robots will be able to perform not only simple repetitive tasks but also artistic expressions and personalized creative activities. As a result, it is expected that the utilization of robots in new fields such as entertainment, education, and art will progress. 7. Improvement in Safety and Reliability: By incorporating safety constraints, this system can improve the safety and reliability of the system. Constraints such as maximum speed, force limits, and the boundaries of the working space are coded into the basic motion primitives, so language model errors cannot override these constraints. For example, the linear speed is limited within ±0.05 m / s, the angular speed is limited within ±60° / s, and the force of the end effector is also limited to 20 N. Also, the end effector is restricted within the boundaries of a pre-defined working space. This makes the robot's operations safe and highly reliable. This effect is achieved by implementing safety constraints in the robot control system. The robot control system implements safety constraints such as maximum speed, force limits, and the boundaries of the working space and constantly monitors these constraints. When an operation violating the constraints is detected, the system automatically stops or corrects the operation. This ensures the safety and reliability of the system. As a specific example, when a robot grasps and moves an object, even if the code generated by the language model erroneously instructs a high-speed operation, the operating speed is limited within a safe range due to the safety constraints of the robot control system. Also, when the robot approaches the boundary of the working space, the system automatically decelerates or stops the operation to prevent crossing the boundary. Furthermore, in the interaction with an object, when there is a possibility of excessive force being applied, the system limits the force to prevent damage to the object and the environment. Such improvements in safety and reliability are essential for the practical application and popularization of robots. Safe and highly reliable robots can be used with confidence in an environment where they coexist with humans. As a result, it is expected that the utilization of robots in fields closely related to humans, such as home, medical, and education, will progress. 8. Natural Interaction with Humans: This system can achieve natural interaction with users through the speech recognition and synthesis module. Users can give instructions in natural language, and the system can also respond in natural language. Moreover, during the execution of tasks, the interaction with users can be continued, and questions and confirmations can be made as needed. This makes the communication between users and robots smoother and improves the user experience. This effect is realized by the speech dialogue management ability of the speech recognition and synthesis module and the dialogue management ability of the language processing component. The speech recognition and synthesis module recognizes the user's speech instructions and outputs the system's response in speech. The language processing component manages the interaction with the user and makes questions and confirmations as needed. Thus, the system can achieve natural interaction with the user. As a specific example, when the user instructs "Make coffee", the system responds with "Understood. I will make coffee. Do you need milk or sugar?" and can confirm the user's preference. Also, when a problem occurs during the execution of the task (e.g., there is a shortage of coffee beans), the system asks "It seems there are few coffee beans. Should we continue like this or make another drink?" and can seek the user's instruction. Furthermore, after the completion of the task, it responds with "Coffee is ready. Is there anything else I can help with?" and can wait for the next instruction. Such natural interaction ability with humans greatly improves the usability and acceptance of robots. Without special training or knowledge, users can operate the robot through natural conversations, enabling various users, such as the elderly, disabled, and children, to easily use the robot. 9. Scalability and Flexibility: This system has a hardware-independent design and is applicable to various robot platforms. Also, since the RAG infrastructure can be flexibly selected, it is possible to use open-source tools such as Haystack and vebra, or the RAG process of OpenAI. Furthermore, the content of the knowledge base can be flexibly expanded, allowing the addition of knowledge for new tasks and environments. This enables the continuous expansion of the system's application scope. This effect is achieved by the modular design of the system and the standardized interface. Each component (language processing, vision system, force sensing module, etc.) operates independently and communicates through the standardized interface. Thus, even if a specific component is replaced with a different implementation, the entire system can continue to function. Also, since the knowledge base is organized in a standardized format, new knowledge can be easily added. As a specific example, when applying this system to different robot platforms (e.g., collaborative robots from Universal Robots, quadruped walking robots from Boston Dynamics), only the interface of the robot control system needs to be adjusted according to the target platform, and other components can be used as they are. Also, when changing the RAG infrastructure (e.g., changing from the RAG process of OpenAI to Haystack), only the RAG-related part of the language processing component needs to be adjusted, and other components can be used as they are. Furthermore, in order to handle new tasks (e.g., "folding laundry"), by simply adding relevant examples to the knowledge base, the system can be enabled to execute that task. Such scalability and flexibility enable the continuous development and wide application of the system. The system can be applied to various hardware platforms and can handle various tasks. As a result, the application scope of the system is significantly expanded, and it is expected that its utilization in diverse industrial fields will progress. 10. Energy conservation and reduction of environmental impact: This system can reduce energy consumption and environmental impact through efficient task execution. For example, the power consumption of the Kinova Gen3 robot arm is approximately 36W, and the power consumption of the NVIDIA RTX 2080 GPU is approximately 225W. The estimated carbon dioxide emissions during a 4-minute task execution are approximately 7g. This indicates high energy efficiency compared to the case when humans perform similar tasks. Also, through task optimization, it is possible to reduce wasted operations and further improve energy efficiency. This effect is achieved by the system's efficient task planning and execution capabilities. The system can execute tasks in an optimal order and minimize wasted operations. Also, by adapting to environmental changes and uncertainties, the need for retries and corrections can be reduced, improving energy efficiency. As a specific example, considering the task of tidying up multiple objects, the system can move the objects in an optimal order considering the positions and destinations of the objects. This can minimize the moving distance and time and reduce energy consumption. Also, if the grasping of an object fails, the system can adjust the grasping position and force using visual and force feedback to increase the success rate of retrial. This can reduce the number of retrials and improve energy efficiency. Such energy conservation and reduction of environmental impact contribute to the realization of a sustainable society. A highly energy-efficient robot system can efficiently utilize limited energy resources and minimize the load on the environment. Thus, it is expected that the spread of robot technology in an environmentally considerate form will progress. Due to these effects, the present invention can contribute to the realization of an "intelligent robot" that can collaborate with humans and perform complex tasks. In particular, applications in various fields are expected, such as support in unpredictable environments such as homes and medical sites, flexible production in manufacturing, and customer response in the service industry.

Modes for Carrying Out the Invention

[0016] Hereinafter, embodiments of the present invention will be described in detail. System Configuration This system is composed of the following hardware and software components: 1. Robot arm: In this embodiment, a 7-degree-of-freedom Kinova Gen3 robot arm is used. This robot arm has a high degree of freedom and precise control ability and is suitable for complex operation tasks. With seven rotational joints, it realizes an operating range and flexibility similar to that of a human arm. Each joint is equipped with a high-precision encoder that can accurately measure the joint angle. Also, each joint has a built-in torque sensor that can detect external forces. This realizes advanced functions such as force control and collision detection. The main specifications of the Kinova Gen3 robotic arm are as follows: - Degrees of freedom: 7 (3 for the shoulder, 1 for the elbow, 3 for the wrist) - Maximum reach: 902 mm - Portable weight: 4 kg - Repeatability accuracy: ±0.1 mm - Maximum speed: 0.5 m / s - Power consumption: 36 W (during normal operation) - Weight: 8.2 kg - Communication interface: Ethernet, USB The robotic arm is controlled through ROS (Robot Operating System). ROS is an open-source software framework for robot development and is suitable for cooperation with various hardware and the construction of complex robot systems. In this system, the Kinova ROS Kortex library is used to establish communication with the Kinova Gen3 robotic arm. 2. Gripper: A Robotiq 2F-140mm gripper is attached to the tip of the robotic arm. This gripper can hold a wide range of objects, and the maximum opening width is 140 mm. Also, since the gripping force can be adjusted, it can grip various objects, from delicate objects (e.g., eggs, glass products) to sturdy objects (e.g., tools, containers), with an appropriate force. The control of the gripper is performed through ROS, and the opening and closing speed and gripping force can be specified. The main specifications of the Robotiq 2F-140mm gripper are as follows: - Maximum opening width: 140 mm - Gripping force: 10 - 125 N (adjustable) - Closing speed: 20 - 150 mm / s (adjustable) - Weight: 1 kg - Power consumption: 5 W (during normal operation) - Communication interface: USB, RS-485 The gripper is controlled through the ROS action server. The action server provides functions such as opening and closing the gripper, adjusting the gripping force, and monitoring the status. The status of the gripper (open / closed position, gripping force, etc.) is updated at a frequency of 50 Hz. 3. Vision Sensor: To generate a three-dimensional representation of the environment, an Azure Kinect DK depth camera is used. This camera operates at a frame rate of 30 fps with a resolution of 640×576 pixels and can simultaneously acquire depth information and RGB images. The depth sensor adopts the Time-of-Flight method and enables high-precision depth measurement. In addition, since it is equipped with a wide-angle lens, a wide field of view (120 degrees horizontally and 120 degrees vertically) is ensured. For camera calibration, a 14 cm AprilTag is used to align the position between the camera and the base of the robot. This enables object position detection with an accuracy of less than 10^-6. The main specifications of the Azure Kinect DK depth camera are as follows: - RGB Resolution: 3840×2160 (4K), 2560×1440 (2K), 1920×1080 (FHD), 1280×720 (HD) - Depth Resolution: 640×576 (30 fps), 512×512 (30 fps), 1024×1024 (15 fps) - Depth Mode: NFOV Unbinned (near distance, high resolution), NFOV 2x2 Binned (near distance, low resolution), WFOV 2x2 Binned (wide angle, low resolution) - Depth Range: 0.5 - 5.46 m (NFOV Unbinned), 0.25 - 2.88 m (WFOV 2x2 Binned) - Field of View: 120 degrees horizontally and 120 degrees vertically - Communication Interface: USB 3.0 Image data from the camera is published as a ROS topic and processed by the vision system. The vision system uses Grounded-Segment-Anything to detect and segment objects. The pose (position and orientation) of the detected objects is updated at a frequency of approximately 1 / 3 Hz. 4. Force Sensor: To measure the force received by the robot's end effector, an ATI multi-axis force sensor is used. This sensor can measure 6 axes (3 axes of force and 3 axes of torque) and operates at a sampling rate of 100 Hz. The measurement accuracy is about 2% of the full scale, and it can also detect fine force changes. The calibration of the sensor is adjusted so that the sensor shows zero in the absence of external force to correct the influence of gravity. This allows the external force applied to the end effector to be accurately predicted. The calibration process is performed in the procedure of zeroing the sensor on one axis, rotating the sensor, and zeroing it on the next axis. The main specifications of the ATI multi-axis force sensor are as follows: - Measurement axes: 6 axes (Fx, Fy, Fz, Tx, Ty, Tz) - Measurement range: Fx, Fy: ±65 N, Fz: ±200 N, Tx, Ty, Tz: ±5 Nm - Resolution: Fx, Fy: 0.025 N, Fz: 0.05 N, Tx, Ty, Tz: 0.001 Nm - Sampling rate: 100 Hz - Communication interface: USB, Ethernet Data from the force sensor is published as a ROS topic and processed by the force sensing module. The force sensing module converts the measured force into the robot's base coordinate system and interprets its meaning. The conversion is performed by the following formula: Fglobal = Tend_effector_to_robot_base × Flocal Here, Fglobal is the force vector in the base coordinate system of the robot, Tend_effector_to_robot_base is the transformation matrix from the coordinate system of the end effector to the base coordinate system of the robot, and Flocal is the force vector in the local coordinate system of the end effector. 5. Voice Recognition and Synthesis Device: To enable voice interaction with the user, a device equipped with a microphone and a speaker (e.g., Amazon Echo, Google Home, or a dedicated microphone and speaker) is used. This device recognizes the user's voice commands and outputs the system's response as voice. Voice recognition and synthesis are performed using cloud-based services (e.g., Amazon Alexa, Google Assistant) or voice recognition and synthesis engines that run locally (e.g., Mozilla DeepSpeech, Festival). The main specifications of the voice recognition and synthesis device are as follows: - Microphone: An array microphone capable of long-distance voice recognition - Speaker: A high-quality speaker capable of clear voice output - Communication Interface: Wi-Fi, Bluetooth, USB - Supported Languages: Japanese, English, and other major languages - Voice Recognition Accuracy: 90% or higher (for general conversations) - Voice Synthesis Quality: Natural intonation and pronunciation Data from the voice recognition and synthesis device is published as a ROS topic and processed by the voice recognition and synthesis module. The voice recognition and synthesis module converts the user's voice commands into text and sends them to the language processing component. It also converts the text response from the language processing component into voice and outputs it to the user. 6. Computer: To perform the system's processing, a desktop computer equipped with an Intel Core i9 processor and an NVIDIA RTX 2080 GPU is used. This computer is connected to a robotic arm, a vision sensor, a force sensor, a speech recognition / synthesis device, and an Ethernet cable or a USB cable. The GPU is used for vision processing and running language models, enabling high-speed processing. It also has sufficient memory (more than 32GB of RAM) and can perform large-scale data processing and complex calculations. The main specifications of the computer are as follows: - Processor: Intel Core i9-9900K (8 cores, 16 threads, 3.6GHz) - GPU: NVIDIA RTX 2080 (8GB VRAM) - Memory: 32GB DDR4 RAM - Storage: 1TB NVMe SSD - OS: Ubuntu 20.04 LTS - Network: Gigabit Ethernet, Wi-Fi 6 On the computer, software such as ROS, vision processing libraries (OpenCV, PCL), machine learning frameworks (PyTorch, TensorFlow), and language model APIs (OpenAI API) is executed. These software are used to implement each component of the system (language processing, vision system, force sensing module, etc.). 7. Operating System: Ubuntu 20.04 is used as the software foundation of the system. This operating system has high compatibility with ROS and enables stable operation. It also supports many robot-related libraries and tools and is suitable as a development environment. The main features of Ubuntu 20.04 are as follows: - Kernel: Linux 5.4 - Desktop Environment: GNOME 3.36 - Package Manager: APT - Support Period: 5 years (until April 2025) - Security: Regular security updates - Hardware Compatibility: Supports a wide range of hardware Ubuntu is the official supported OS for ROS and is easy to install and set up ROS. Also, many robot-related libraries and tools are packaged for Ubuntu, making it easy to build a development environment.

[0017] 8. Robot Control Software: For robot control, ROS (Robot Operating System) is used. ROS is an open-source software framework for robot development and is suitable for cooperation with various hardware and the construction of complex robot systems. In this system, the Kinova ROS Kortex library is used to establish communication with the Kinova Gen3 robot arm. Utilizing the inter-node communication function of ROS, language processing, visual system, and feedback from the force sensing module are integrated to control the operation of the robot. The main functions of ROS are as follows: - Inter-node Communication: Communication between nodes through topics, services, and actions - Distributed Computing: Construction of robot systems spanning multiple computers - Hardware Abstraction: An interface for uniformly handling various hardware - Debugging Tools: Tools (RViz, rqt) to assist in debugging robot systems - Simulation: Cooperation with simulators such as Gazebo - Package Management: Management of robot-related software packages ROS provides three communication methods: topics, services, and actions. Topics are used for asynchronous one-way communication and are suitable for distributing sensor data. Services are used for synchronous request-response communication and are suitable for calculations and inquiries. Actions are used for controlling long-running tasks and are suitable for controlling robot movements. In this system, these communication methods are appropriately combined to realize communication between components. 9. Language Model: To process user queries and environmental data, large language models such as GPT-4 are used. GPT-4 has high natural language understanding and generation capabilities, can understand complex instructions, and can generate appropriate code. In this system, language processing is performed through the GPT-4 API. Also, to implement RAG, the OpenAI RAG process is used or open-source tools such as Haystack and vebra are used. The main features of GPT-4 are as follows: - Number of parameters: Not publicly available (GPT-3 has 175 billion) - Context window: 8192 tokens (about 32,000 words) - Language understanding ability: Understand complex instructions and context and respond appropriately - Code generation ability: Generate code in various programming languages - Reasoning ability: Perform logical reasoning, common sense reasoning, mathematical reasoning, etc. - Multilingual support: Support many languages such as English, Japanese, Chinese, Spanish, etc. The GPT-4 API can set the following parameters: - Model: gpt-4, gpt-4-0613, etc. - Maximum number of tokens: The maximum length of the text to be generated - Temperature: Control the diversity of generation (0 - 2, the lower the more deterministic) - Top P: Control the variations of generation (0 - 1, the lower the more deterministic) - Frequency penalty: Suppresses repetition (0 - 2, higher value means more suppression) - Presence penalty: Promotes the introduction of new topics (0 - 2, higher value means more promotion) This system uses the following settings: - Model: gpt-4-0613 - Maximum number of tokens: 8192 - Temperature: 0.7 (to balance creativity and consistency) - Top P: 0.95 (to generate diverse responses) - Frequency penalty: 0.0 (to suppress repetition) - Presence penalty: 0.0 (to promote the introduction of new topics) 10. Visual Processing Module: For object identification and segmentation, language-vision models such as Grounded-Segment-Anything are used. This model can identify objects based on language instructions and generate segmentation masks. Additionally, MobileSAM is used to create segmented masks and 3D voxels that enclose the detected objects. This allows for the extraction of object poses, which can be used for robot grasping and manipulation. The main features of Grounded-Segment-Anything are as follows: - Language-vision model: Identifies objects based on language instructions - Segmentation: Generates detailed segmentation masks of objects - Zero-shot transfer: Can identify objects not present in the training data - Real-time processing: Allows for fast processing (about 0.3 seconds / frame) - High accuracy: 52.5 AP (average precision) on the COCO zero-shot transfer benchmark The visual processing module performs the following processing: - Preprocessing of images: Noise removal, contrast adjustment, etc. - Object Detection: Detect objects based on language instructions - Segmentation: Generate a segmentation mask for the detected objects - 3D Reconstruction: Combine depth information and the segmentation mask to create a 3D model of the object - Pose Estimation: Estimate the position and orientation of the object - Tracking: Track the movement of the object The visual processing module is implemented as a ROS node, processes image data from the camera, and publishes the pose of the detected object as a topic. This allows the robot control system to grasp the position and orientation of the object and perform appropriate grasping and operations. 11. Knowledge Base: The system's knowledge base uses a curated database that includes examples of verified low-level and high-level actions. This database systematically organizes code examples, motion primitives, examples of task decomposition, etc. It also includes motion examples that incorporate known uncertainties and can handle various scenarios. The format of the knowledge base selects a format compatible with the RAG system, such as markdown files or structured JSON. The main components of the knowledge base are as follows: - Basic Motion Primitives: Basic actions such as linear motion, rotational motion, opening and closing of the gripper - Motion Primitives for Object Manipulation: Actions related to object manipulation such as grasping, placing, lifting, transporting - Special Motion Primitives: Special actions such as liquid injection, powder scooping, opening and closing of drawers, drawing - High-Level Task Examples: Complex tasks such as coffee making, dish decoration, object handover - Error Handling and Recovery Examples: Examples for dealing with and recovering from errors during operation - Environment Adaptation Examples: Examples for adapting to environmental changes and uncertainties Each example includes the following information: - Code: Executable Python code - Input parameters: Parameters required for code execution (e.g., target position, speed, force) - Output: Execution result of the code (e.g., success / failure, error message) - Prerequisites: Prerequisites required for code execution (e.g., the existence of a specific object, the robot being in a specific initial state) - Post - conditions: The expected state after code execution (e.g., the object being placed at a specific position, the robot being in a specific state) - Uncertainty: Uncertainty associated with code execution (e.g., uncertainty in the position of an object, accuracy of force control) - Coping methods: Methods for coping with uncertainty (e.g., utilization of visual feedback, utilization of force feedback) The knowledge base is searched by the RAG system, and relevant examples are provided to the language model. Thereby, the language model can generate responses suitable for the current task. Language processing component The language processing component is responsible for processing the user's query and environmental data and decomposing complex tasks into a series of executable steps. # Selection and configuration of the language model In this embodiment, GPT - 4 is used as the language model. GPT - 4 has high natural language understanding and generation capabilities, can understand complex instructions, and can generate appropriate code. Language processing is performed through the GPT - 4 API. The configuration parameters of the API are as follows: - Model: gpt - 4 - 0613 - Maximum number of tokens: 8192 - Temperature: 0.7 (to balance creativity and consistency) - Top - P: 0.95 (to generate diverse responses) - Frequency penalty: 0.0 (to suppress repetition) - Presence Penalty: 0.0 (to encourage the introduction of new topics) These parameters can be adjusted according to the nature of the task and the required level of creativity. For example, if a more decisive response is needed, the temperature and top P can be set low, and if a more creative response is needed, these parameters can be set high. The input (prompt) to GPT-4 consists of the following elements: - System message: Specifies the role of the model and the constraints on its operation - User query: Instructions or questions from the user - Environmental data: Images from visual sensors and the state of the environment - Relevant examples from the knowledge base: Relevant examples retrieved by RAG Examples of system messages: ``` You are an assistant for a robot control system. Based on the user's instructions and environmental data, analyze the tasks that the robot should perform and generate appropriate Python code. The code should be executed in a ROS environment and utilize visual and force feedback. Give top priority to safety and comply with constraints such as maximum speed and force limits, and the boundaries of the working space. ``` With such prompt design, GPT-4 can understand the appropriate role and generate safe and effective code. # The process of task decomposition The language processing component receives the user's high-level instructions (e.g., "Make coffee") and combines them with environmental data (images from visual sensors) to decompose the task into a series of subtasks. This process is carried out in the following steps: 1. Analyze the user's instructions and identify the main goal (e.g., "Make coffee"). 2. Analyze the environmental data to identify available resources (e.g., mug, coffee beans, kettle). 3. Identify the subtasks necessary to achieve the goal. For example, the task of "making coffee" can be decomposed into the following subtasks: - Find the mug - Find the coffee beans - Find the kettle - Place the mug in the appropriate position - Scoop up the coffee beans - Put the coffee beans into the mug - Pour hot water from the kettle 4. Identify the dependencies between subtasks. For example, "putting the coffee beans into the mug" depends on "placing the mug in the appropriate position". 5. Define the execution conditions and success criteria for each subtask. For example, the success criterion for the subtask of "finding the mug" is that "the position of the mug is identified and the robot is in a state where it can grasp the mug". This process of task decomposition is expressed in the form of conditional probabilities. For example, P(L2A, L2B|L1) specifies the probability of proceeding to task L2A or L2B after the successful execution of task L1. This allows the next subtask to be dynamically selected according to the progress of the task.

[0018] Specifically, the following conditional probabilities are defined: - P(finding the mug|initial state) = 0.9: The probability of finding the mug from the initial state is 90% - P(opening the drawer|not finding the mug) = 0.8: If the mug cannot be found, the probability of opening the drawer is 80% - P(finding the mug|opening the drawer) = 0.7: After opening the drawer, the probability of finding the mug is 70% These conditional probabilities are set based on examples in the knowledge base and past experiences. They can also be updated dynamically based on feedback obtained during the execution of the task. # Implementation of Retrieval-Augmented Generation (RAG) In this system, RAG is used to retrieve relevant examples from the knowledge base and improve the output of the language model. The following steps are adopted for the implementation of RAG: 1. Query Embedding: Convert the user's query and the current task state into vector embeddings. Use embedders such as OpenAI's text-embedding-ada-002 or HuggingFace's Sentence-BERT as the embedding model. 2. Chunking: Divide the knowledge base into meaningful units (chunks). The size of the chunks is determined considering the balance between context retention and retrieval efficiency. 3. Chunk Embedding: Convert each chunk into a vector embedding. 4. Similarity Calculation: Calculate the cosine similarity between the query embedding and each chunk embedding. 5. Selection of Top Chunks: Select the top k chunks based on similarity. The value of k is adjusted according to the complexity of the task and the size of the knowledge base. 6. Context Expansion: Add the selected chunks to the input to the language model. 7. Response Generation: Use the expanded context for the language model to generate a response (code). In this embodiment, the OpenAI RAG process is used to organize the curated knowledge base as a Markdown file. However, in the framework of this system, other RAG approaches using tools such as Haystack or vebra are also available. When using these tools, the following components can be selected: - Document Store: A way to store and organize the knowledge base (e.g., Markdown file, Elasticsearch) - Embedder: A model that converts text into vector representations (e.g., OpenAI Embeddings, Sentence - BERT) - Retriever: A method for retrieving documents relevant to a query (e.g., vector search, keyword search, hybrid search) - Chunking technique: A method for splitting documents into meaningful units (e.g., fixed - size, paragraph - based, semantic chunking) - Language model: A model that generates the final response (e.g., GPT - 4, GPT - 3.5 - turbo, Zephyr - 7B - beta) As an example implementation of RAG, Python code like the following can be considered: ```python from openai import OpenAI import numpy as np from sklearn.metrics.pairwise import cosine_similarity # Initialization of the OpenAI API client client = OpenAI(api_key="your - api - key") # Dictionary to store knowledge base chunks and their embeddings knowledge_base = {} # Initialization of the knowledge base (extract chunks from a markdown file and calculate embeddings) def initialize_knowledge_base(markdown_file): with open(markdown_file, 'r') as f: content = f.read() # Split the markdown into meaningful chunks chunks = split_into_chunks(content) # Calculate the embedding of each chunk for chunk in chunks: embedding = get_embedding(chunk) knowledge_base[chunk] = embedding # Get the embedding of the text def get_embedding(text): response = client.embeddings.create( model="text-embedding-ada-002", input=text ) return response.data[0].embedding # Search for chunks related to the query def retrieve_relevant_chunks(query, k=5): # Get the embedding of the query query_embedding = get_embedding(query) # Calculate the similarity between each chunk and the query similarities = {} for chunk, chunk_embedding in knowledge_base.items(): similarity = cosine_similarity([query_embedding], [chunk_embedding])[0][0] similarities[chunk] = similarity # Select the top k chunks based on similarity sorted_chunks = sorted(similarities.items(), key=lambda x: x[1], reverse=True) top_k_chunks = [chunk for chunk, _ in sorted_chunks[:k]] return top_k_chunks # Response generation using RAG def generate_response_with_rag(query, system_message, k=5): # Search for relevant chunks relevant_chunks = retrieve_relevant_chunks(query, k) # Expand the context context = "\n\n".join(relevant_chunks) # Send the input to the language model response = client.chat.completions.create( model="gpt-4-0613", messages= {"role": "system", "content": system_message}, {"role": "user", "content": f"Query: {query}\n\nRelevant Information:\n{context}"} , max_tokens=8192, temperature=0.7, top_p=0.95, frequency_penalty=0.0, presence_penalty=0.0 ) return response.choices[0].message.content ``` With such a RAG implementation, the system can efficiently retrieve relevant examples from the knowledge base and improve the output of the language model. # Code Generation and Execution The language processing component generates executable code based on the relevant examples obtained through RAG. The generated code is written in Python (v.3.8) that is compatible with ROS and includes the following elements: 1. Import necessary libraries and modules 2. Initialize the ROS node 3. Function calls to obtain visual and force feedback 4. Main function for task execution 5. Error handling and recovery mechanism 6. Implementation of safety constraints

[0019] Example of the generated code: ```python #! / usr / bin / env python3 # Import necessary libraries and modules import rospy import numpy as np from std_msgs.msg import String from geometry_msgs.msg import Pose, Twist from sensor_msgs.msg import Image, JointState from kortex_driver.srv import * from kortex_driver.msg import * from cv_bridge import CvBridge import tf2_ros import tf2_geometry_msgs from tf.transformations import quaternion_from_euler # Initialization of ROS node rospy.init_node('coffee_making_node', anonymous=True) # Global variables robot_name = "my_gen3" bridge = CvBridge() tf_buffer = tf2_ros.Buffer() tf_listener = tf2_ros.TransformListener(tf_buffer) # Configuration of service client base_service_name = ' / ' + robot_name + ' / base / base_service' base_client = rospy.ServiceProxy(base_service_name, Base_ClearFaults) # Function to obtain visual feedback def get_object_pose(object_name): try: # Obtain the detection result of the object object_pose = rospy.wait_for_message(' / vision / object_poses', Pose, timeout=5.0) return object_pose except rospy.ROSException: rospy.logerr(f"Failed to get pose for object: {object_name}") return None # Function to obtain force feedback def get_force_feedback(): try: # Obtain the reading value of the force sensor force_data = rospy.wait_for_message(' / force_torque_sensor / wrench', WrenchStamped, timeout=5.0) return force_data.wrench except rospy.ROSException: rospy.logerr("Failed to get force feedback") return None # Function to control the gripper def control_gripper(position, speed=0.1, force=0.8): try: # Call the gripper control service service_name = ' / ' + robot_name + ' / base / send_gripper_command' gripper_client = rospy.ServiceProxy(service_name, SendGripperCommand) # Set the gripper command gripper_command = SendGripperCommandRequest() gripper_command.input.gripper.finger[0].value = position gripper_command.input.mode = GripperMode.GRIPPER_POSITION # Send the gripper command gripper_client(gripper_command) return True except rospy.ServiceException as e: rospy.logerr(f"Failed to control gripper: {e}") return False # Function to move the robot def move_robot(target_pose, speed=0.1): try: # Set up the publisher for Cartesian velocity commands velocity_pub = rospy.Publisher(' / ' + robot_name + ' / in / cartesian_velocity', TwistCommand, queue_size=10) # Get the current position current_pose = rospy.wait_for_message(' / ' + robot_name + ' / base / tool_pose', Pose, timeout=5.0) # Calculate the direction vector to the target position direction = np.array( target_pose.position.x - current_pose.position.x, target_pose.position.y - current_pose.position.y, target_pose.position.z - current_pose.position.z ) # Normalize the direction vector distance = np.linalg.norm(direction) if distance > 0: direction = direction / distance # Setting the velocity command twist_cmd = TwistCommand() twist_cmd.reference_frame = CartesianReferenceFrame.CARTESIAN_REFERENCE_FRAME_BASE twist_cmd.twist.linear_x = direction[0] * speed twist_cmd.twist.linear_y = direction[1] * speed twist_cmd.twist.linear_z = direction[2] * speed # Send the velocity command until the target position is reached rate = rospy.Rate(10) # 10Hz while distance > 0.01: # Get the current position current_pose = rospy.wait_for_message(' / ' + robot_name + ' / base / tool_pose', Pose, timeout=5.0) # Recalculate the direction vector to the target position direction = np.array( target_pose.position.x - current_pose.position.x, target_pose.position.y - current_pose.position.y, target_pose.position.z - current_pose.position.z ) # Normalize the direction vector distance = np.linalg.norm(direction) if distance > 0: direction = direction / distance # Update velocity command twist_cmd.twist.linear_x = direction[0] * min(speed, distance) twist_cmd.twist.linear_y = direction[1] * min(speed, distance) twist_cmd.twist.linear_z = direction[2] * min(speed, distance) # Send velocity command velocity_pub.publish(twist_cmd) rate.sleep() # Send stop command twist_cmd.twist.linear_x = 0.0 twist_cmd.twist.linear_y = 0.0 twist_cmd.twist.linear_z = 0.0 velocity_pub.publish(twist_cmd) return True except rospy.ROSException as e: rospy.logerr(f"Failed to move robot: {e}") return False

[0020] # Function to find the mug def find_mug(): rospy.loginfo("Finding mug...") # Get the position of the mug mug_pose = get_object_pose("mug") if mug_pose is None: rospy.logwarn("Mug not found in the current view. Searching in drawers...") # Call the function to open the drawer open_drawer() # Get the position of the mug again mug_pose = get_object_pose("mug") if mug_pose is None: rospy.logerr("Failed to find mug even after opening drawer.") return None rospy.loginfo(f"Mug found at position: {mug_pose.position.x}, {mug_pose.position.y}, {mug_pose.position.z}") return mug_pose # Function to open the drawer def open_drawer(): rospy.loginfo("Opening drawer...") # Get the position of the drawer handle handle_pose = get_object_pose("drawer_handle") if handle_pose is None: rospy.logerr("Failed to find drawer handle.") return False # Approach the handle approach_pose = Pose() approach_pose.position.x = handle_pose.position.x - 0.1 approach_pose.position.y = handle_pose.position.y approach_pose.position.z = handle_pose.position.z approach_pose.orientation = handle_pose.orientation move_robot(approach_pose) # Grasp the handle move_robot(handle_pose) control_gripper(0.5) # Close the gripper # Obtain force feedback initial_force = get_force_feedback() if initial_force is None: rospy.logerr("Failed to get force feedback.") control_gripper(1.0) # Open the gripper return False # Pull the drawer pull_pose = Pose() pull_pose.position.x = handle_pose.position.x - 0.3 pull_pose.position.y = handle_pose.position.y pull_pose.position.z = handle_pose.position.z pull_pose.orientation = handle_pose.orientation # Pull while monitoring force feedback try: velocity_pub = rospy.Publisher(' / ' + robot_name + ' / in / cartesian_velocity', TwistCommand, queue_size=10) twist_cmd = TwistCommand() twist_cmd.reference_frame = CartesianReferenceFrame.CARTESIAN_REFERENCE_FRAME_BASE twist_cmd.twist.linear_x = -0.05 # Pull slowly rate = rospy.Rate(10) # 10Hz # Pull until the drawer opens or the force exceeds the threshold max_force = 10.0 # Maximum force (N) max_duration = rospy.Duration(5.0) # Maximum pulling time start_time = rospy.Time.now() while (rospy.Time.now() - start_time) < max_duration: # Get force feedback force = get_force_feedback() if force is None: break # Stop pulling if the force exceeds the threshold if abs(force.force.x) > max_force: rospy.logwarn("Force threshold exceeded. Stopping pull.") break # Sending velocity command velocity_pub.publish(twist_cmd) rate.sleep() # Sending stop command twist_cmd.twist.linear_x = 0.0 velocity_pub.publish(twist_cmd) except rospy.ROSException as e: rospy.logerr(f"Failed to pull drawer: {e}") control_gripper(1.0) # Open the gripper return False # Open the gripper control_gripper(1.0) rospy.loginfo("Drawer opened successfully.") return True # Function to scoop coffee def scoop_coffee(): rospy.loginfo("Scooping coffee...") # Get the position of the coffee container coffee_container_pose = get_object_pose("coffee_container") if coffee_container_pose is None: rospy.logerr("Failed to find coffee container.") return False # Get the position of the spoon spoon_pose = get_object_pose("spoon") if spoon_pose is None: rospy.logerr("Failed to find spoon.") return False # Grasp the spoon move_robot(spoon_pose) control_gripper(0.5) # Close the gripper # Lift the spoon lift_pose = Pose() lift_pose.position.x = spoon_pose.position.x lift_pose.position.y = spoon_pose.position.y lift_pose.position.z = spoon_pose.position.z + 0.1 lift_pose.orientation = spoon_pose.orientation move_robot(lift_pose) # Move above the coffee container above_container_pose = Pose() above_container_pose.position.x = coffee_container_pose.position.x above_container_pose.position.y = coffee_container_pose.position.y above_container_pose.position.z = coffee_container_pose.position.z + 0.1 above_container_pose.orientation = spoon_pose.orientation move_robot(above_container_pose) # Insert the spoon into the coffee container scoop_pose = Pose() scoop_pose.position.x = coffee_container_pose.position.x scoop_pose.position.y = coffee_container_pose.position.y scoop_pose.position.z = coffee_container_pose.position.z + 0.02 scoop_pose.orientation = spoon_pose.orientation move_robot(scoop_pose)

[0021] # Move the spoon horizontally to scoop the coffee scoop_motion_pose = Pose() scoop_motion_pose.position.x = coffee_container_pose.position.x + 0.05 scoop_motion_pose.position.y = coffee_container_pose.position.y scoop_motion_pose.position.z = coffee_container_pose.position.z + 0.02 scoop_motion_pose.orientation = spoon_pose.orientation move_robot(scoop_motion_pose) # Lift the spoon lift_after_scoop_pose = Pose() lift_after_scoop_pose.position.x = scoop_motion_pose.position.x lift_after_scoop_pose.position.y = scoop_motion_pose.position.y lift_after_scoop_pose.position.z = scoop_motion_pose.position.z + 0.1 lift_after_scoop_pose.orientation = spoon_pose.orientation move_robot(lift_after_scoop_pose) rospy.loginfo("Coffee scooped successfully.") return True # Function to pour water def pour_water(mug_pose): rospy.loginfo("Pouring water...") # Get the position of the kettle kettle_pose = get_object_pose("kettle") if kettle_pose is None: rospy.logerr("Failed to find kettle.") return False # Grasp the kettle move_robot(kettle_pose) control_gripper(0.5) # Close the gripper # Lift the kettle lift_pose = Pose() lift_pose.position.x = kettle_pose.position.x lift_pose.position.y = kettle_pose.position.y lift_pose.position.z = kettle_pose.position.z + 0.1 lift_pose.orientation = kettle_pose.orientation move_robot(lift_pose) # Move above the mug above_mug_pose = Pose() above_mug_pose.position.x = mug_pose.position.x above_mug_pose.position.y = mug_pose.position.y above_mug_pose.position.z = mug_pose.position.z + 0.2 above_mug_pose.orientation = kettle_pose.orientation move_robot(above_mug_pose) # Tilt the kettle pour_orientation = quaternion_from_euler(0, 0.5, 0) # Tilt by about 30 degrees pour_pose = Pose() pour_pose.position = above_mug_pose.position pour_pose.orientation.x = pour_orientation[0] pour_pose.orientation.y = pour_orientation[1] pour_pose.orientation.z = pour_orientation[2] pour_pose.orientation.w = pour_orientation[3] # Pour while monitoring force feedback try: # Get the initial force initial_force = get_force_feedback() if initial_force is None: rospy.logerr("Failed to get initial force feedback.") return False # Tilt the kettle move_robot(pour_pose) # Wait for a certain period (while the hot water is being poured) rospy.sleep(3.0) # Get the current force current_force = get_force_feedback() if current_force is None: rospy.logerr("Failed to get current force feedback.") return False # Estimate the amount poured from the change in force force_change = current_force.force.z - initial_force.force.z poured_amount = force_change / 9.8 # Convert the change in force to mass (approximate) rospy.loginfo(f"Estimated poured amount: {poured_amount} kg") # Return the kettle to its original orientation move_robot(above_mug_pose) except rospy.ROSException as e: rospy.logerr(f"Failed to pour water: {e}") return False # Return the kettle to its original position move_robot(lift_pose) move_robot(kettle_pose) # Open the gripper control_gripper(1.0) rospy.loginfo("Water poured successfully.") return True # Main function def make_coffee(): rospy.loginfo("Starting coffee making process...") try: # Find the mug mug_pose = find_mug() if mug_pose is None: rospy.logerr("Failed to find mug. Aborting coffee making.") return False # Scoop coffee if not scoop_coffee(): rospy.logerr("Failed to scoop coffee. Aborting coffee making.") return False # Pour coffee into the mug # (This part is omitted, but the process of moving the spoon above the mug and tilting it to pour coffee is required) # Pour hot water if not pour_water(mug_pose): rospy.logerr("Failed to pour water. Aborting coffee making.") return False rospy.loginfo("Coffee made successfully!") return True except Exception as e: rospy.logerr(f"An error occurred during coffee making: {e}") return False # Main execution if __name__ == '__main__': try: # Initialize the robot rospy.wait_for_service(base_service_name, timeout=10) base_client() # Make coffee make_coffee() except rospy.ROSInterruptException: pass ``` Such code is generated based on relevant examples obtained through RAG and is described in a format compatible with ROS. The code includes function calls for utilizing visual and force feedback and can handle changes and uncertainties in the environment. Safety constraints (such as speed limits and force limits) are also incorporated. The generated code is executed in a secure environment. This environment is restricted to accessing only pre-defined functions, ensuring the safety of the system. Also, before the code is executed, syntax checks and security checks are performed to detect potential problems.

[0022] Visual system The visual system is responsible for generating a three-dimensional representation of the environment and identifying the positions of objects. # Camera setup and calibration In this embodiment, an Azure Kinect DK depth camera is used. The camera settings are as follows: - Resolution: 640×576 pixels - Frame rate: 30fps - Depth mode: NFOV Unbinned (near distance, high resolution) - Color format: BGRA - Exposure time: Auto-adjust For camera calibration, a 14-cm AprilTag is used. The AprilTag is a marker with a known size and shape and is used for alignment between the camera and the base of the robot. The calibration process is performed in the following steps: 1. Place the AprilTag at a known position within the robot's workspace. 2. Photograph the AprilTag with a camera and detect its position and size. 3. Calculate the internal parameters (focal length, principal point) and external parameters (position, orientation) of the camera from the detected position and size. 4. Using the calculated parameters, obtain the transformation matrix from the camera coordinate system to the robot's base coordinate system. This calibration enables the detection of the object's position with an accuracy of less than 10^-6. The transformation matrix is represented by the following equation: PR = TAR × (TCA × PC) Here, PC is the point in the camera coordinate system, TCA is the transformation matrix from the camera coordinate system to the AprilTag coordinate system, TAR is the transformation matrix from the AprilTag coordinate system to the robot's base coordinate system, and PR is the point in the robot's base coordinate system. The calibration process is implemented using the ROS tf2 library. tf2 is a library for managing transformations between different coordinate systems and can handle transformations that change over time. # Object Detection and Segmentation For object detection and segmentation, Grounded-Segment-Anything is used. This model can identify objects based on language instructions and generate segmentation masks. The processing flow is as follows: 1. Input the image obtained from the camera into the Grounded-Segment-Anything model. 2. The model detects the corresponding object based on language instructions (e.g., "white mug", "black kettle"). 3. Generate a bounding box for the detected object. 4. Use MobileSAM to create a segmented mask. 5. Create 3D voxels that enclose the detected object in combination with depth information. 6. Extract the pose (position and orientation) of the object from the voxels. Through this process, the 3D position and orientation of the object can be determined and used for robot grasping and operation. The detection accuracy of the object varies depending on the type of object and environmental conditions. According to the experimental results, the detection accuracy of a white mug was approximately 70%, that of a black kettle was approximately 79%, and that of a hand was approximately 53%. The implementation of Grounded-Segment-Anything is carried out with the following Python code: ```python import torch import numpy as np import cv2 from PIL import Image from groundingdino.util.inference import load_model, load_image, predict from segment_anything import sam_model_registry, SamPredictor import rospy from sensor_msgs.msg import Image as RosImage from geometry_msgs.msg import Pose from cv_bridge import CvBridge class ObjectDetector: def __init__(self): # Load the Grounding DINO model self.grounding_dino_model = load_model("groundingdino / config / GroundingDINO_SwinT_OGC.py", "groundingdino / weights / groundingdino_swint_ogc.pth") # Loading the SAM model self.sam = sam_model_registry["vit_h"](checkpoint="sam_vit_h_4b8939.pth") self.sam_predictor = SamPredictor(self.sam) # ROS-related settings self.bridge = CvBridge() self.image_sub = rospy.Subscriber(" / camera / color / image_raw", RosImage, self.image_callback) self.depth_sub = rospy.Subscriber(" / camera / depth / image_raw", RosImage, self.depth_callback) self.pose_pub = rospy.Publisher(" / vision / object_poses", Pose, queue_size=10) # Variables to store the latest image and depth information self.latest_image = None self.latest_depth = None # Classes of objects to be detected self.classes = ["mug", "kettle", "spoon", "coffee_container", "drawer_handle"] def image_callback(self, msg): # Convert ROS Image type to OpenCV image self.latest_image = self.bridge.imgmsg_to_cv2(msg, "bgr8") def depth_callback(self, msg): # Convert ROS Image type to OpenCV depth image self.latest_depth = self.bridge.imgmsg_to_cv2(msg, "32FC1") def detect_objects(self, text_prompt): if self.latest_image is None or self.latest_depth is None: rospy.logwarn("No image or depth data available.") return None # Image preprocessing image_pil = Image.fromarray(self.latest_image) image_tensor, _ = load_image(image_pil) # Object detection with Grounding DINO boxes, logits, phrases = predict( model=self.grounding_dino_model, image=image_tensor, caption=text_prompt, box_threshold=0.3, text_threshold=0.25 ) # If no detection results if len(boxes) == 0: rospy.logwarn(f"No objects detected for prompt: {text_prompt}") return None # Segmentation with SAM self.sam_predictor.set_image(self.latest_image) result_poses = [] for box, logit, phrase in zip(boxes, logits, phrases): # Get the coordinates of the bounding box x0, y0, x1, y1 = box # Segmentation with SAM sam_mask, _, _ = self.sam_predictor.predict( box=np.array([x0, y0, x1, y1]), multimask_output=False ) # Calculate 3D position from depth information using the mask mask = sam_mask[0].astype(np.uint8) masked_depth = self.latest_depth.copy() masked_depth[mask == 0] = 0 # Extract pixels with valid depth values valid_depth = masked_depth[masked_depth > 0] if len(valid_depth) == 0: rospy.logwarn(f"No valid depth data for object: {phrase}") continue # Calculate the median depth (robust to noise) median_depth = np.median(valid_depth) # Calculate the center of the bounding box center_x = (x0 + x1) / 2 center_y = (y0 + y1) / 2 # Convert from image coordinate system to 3D coordinate system # (Camera intrinsic parameters are required) fx = 500 # Camera focal length x (replace with actual value) fy = 500 # Camera focal length y (replace with actual value) cx = 320 # Principal point x (replace with actual value)

[0023] # Calculate 3D coordinates z = median_depth x = (center_x - cx) * z / fx y = (center_y - cy) * z / fy # Create a Pose message pose = Pose() pose.position.x = x pose.position.y = y pose.position.z = z # The pose is represented as a unit quaternion for simplicity pose.orientation.x = 0.0 pose.orientation.y = 0.0 pose.orientation.z = 0.0 pose.orientation.w = 1.0 # Add the result result_poses.append((phrase, pose, logit.item())) # Sort by confidence result_poses.sort(key=lambda x: x[2], reverse=True) # Publish the results for phrase, pose, _ in result_poses: rospy.loginfo(f"Detected {phrase} at position: {pose.position.x}, {pose.position.y}, {pose.position.z}") self.pose_pub.publish(pose) return result_poses def detect_specific_object(self, object_name): # Create a prompt for detecting a specific object text_prompt = object_name # Execute object detection results = self.detect_objects(text_prompt) if results is None or len(results) == 0: return None # Return the most confident result return results[0][1] # Return the Pose ``` This code is implemented as a ROS node that processes images and depth information from a camera to detect the pose of an object and publish it as a ROS topic. The detected pose of the object is utilized by a robot control system for grasping or operating on the object. # Occlusion and Uncertainty Handling The vision system has mechanisms to handle occlusion (where part of an object is hidden by another object) and uncertainty. For handling occlusion, the following approaches are adopted: 1. Detection based on partial visibility of the object: Even when part of the object is visible, the object is detected based on its features. 2. Temporal tracking: Even when the object is temporarily hidden, tracking is continued based on past position information. 3. Observation from multiple viewpoints: If possible, observations from different angles are combined to reduce the impact of occlusion. For handling uncertainty, the following approaches are adopted: 1. Probabilistic representation: The position and orientation of the object are represented as a probability distribution, explicitly considering uncertainty. 2. Filtering: Techniques such as the Kalman filter are used to reduce the influence of noise and perform more stable estimation. 3. Active perception: When uncertainty is high, the robot takes actions such as actively changing the viewing point to collect more information. According to the experimental results, when the occlusion rate was 20% - 30%, the detection success rate of the white mug was about 90%. However, when the occlusion rate reached 80% - 90%, the detection success rate decreased to about 20%. In such situations, it is important to utilize other modalities such as force feedback. The handling of occlusion and uncertainty is implemented in the following Python code: ```python import numpy as np import rospy from geometry_msgs.msg import Pose from filterpy.kalman import KalmanFilter class ObjectTracker: def __init__(self, object_name): self.object_name = object_name # Initialization of the Kalman filter self.kf = KalmanFilter(dim_x=6, dim_z=3) # State: [x, y, z, vx, vy, vz], Observation: [x, y, z] # State transition matrix dt = 0.1 # Time step self.kf.F = np.array( [1, 0, 0, dt, 0, 0], [0, 1, 0, 0, dt, 0], [0, 0, 1, 0, 0, dt], [0, 0, 0, 1, 0, 0], [0, 0, 0, 0, 1, 0], [0, 0, 0, 0, 0, 1] ) # Observation matrix self.kf.H = np.array( [1, 0, 0, 0, 0, 0], [0, 1, 0, 0, 0, 0], [0, 0, 1, 0, 0, 0] ) # Measurement noise self.kf.R = np.eye(3) * 0.01 # Process noise self.kf.Q = np.eye(6) * 0.01 # Initial state self.kf.x = np.zeros(6) # Initial covariance self.kf.P = np.eye(6) * 1000 # Time when the object was last seen self.last_seen_time = None # Maximum time (seconds) to continue tracking after the object disappears self.max_tracking_duration = 3.0 # Visibility of the object self.is_visible = False # ROS-related settings self.pose_sub = rospy.Subscriber(f" / vision / object_poses / {object_name}", Pose, self.pose_callback) self.tracked_pose_pub = rospy.Publisher(f" / vision / tracked_object_poses / {object_name}", Pose, queue_size=10) # Update the state periodically with a timer self.timer = rospy.Timer(rospy.Duration(dt), self.update) def pose_callback(self, msg): # When the object is detected self.is_visible = True self.last_seen_time = rospy.Time.now() # Set the observation value z = np.array( msg.position.x, msg.position.y, msg.position.z ) # Update the Kalman filter self.kf.update(z) def update(self, event): # Predict the state self.kf.predict() # Calculate the elapsed time since the object disappeared if self.last_seen_time is not None: elapsed_time = (rospy.Time.now() - self.last_seen_time).to_sec() # Stop tracking if the object has not been seen for a certain period of time if elapsed_time > self.max_tracking_duration: self.is_visible = False # Publish the tracking result if self.last_seen_time is not None: pose = Pose() pose.position.x = self.kf.x[0] pose.position.y = self.kf.x[1] pose.position.z = self.kf.x[2] # The pose is represented as a unit quaternion for simplicity pose.orientation.x = 0.0 pose.orientation.y = 0.0 pose.orientation.z = 0.0 pose.orientation.w = 1.0

[0024] # Add uncertainty information (since it cannot be directly included in the ROS Pose message, it needs to be published on a separate topic or extended) uncertainty = np.sqrt(np.diag(self.kf.P)[:3]) # Position uncertainty (standard deviation) rospy.logdebug(f"Tracked {self.object_name} at position: {pose.position.x}, {pose.position.y}, {pose.position.z}") rospy.logdebug(f"Uncertainty: {uncertainty}") self.tracked_pose_pub.publish(pose) class ActivePerception: def __init__(self, object_detector, robot_controller): self.object_detector = object_detector self.robot_controller = robot_controller # List of observed positions (in the robot's base coordinate system) self.observation_poses = # Observation from the front {'position': [0.5, 0.0, 0.5], 'orientation': [0.0, 0.0, 0.0, 1.0]}, # Observation from the left {'position': [0.5, 0.3, 0.5], 'orientation': [0.0, 0.0, 0.3826834, 0.9238795]}, # Observation from the right {"position": [0.5, -0.3, 0.5], "orientation": [0.0, 0.0, -0.3826834, 0.9238795]}, # Observation from above {"position": [0.5, 0.0, 0.7], "orientation": [0.0, 0.3826834, 0.0, 0.9238795]} def find_object_with_active_perception(self, object_name, max_attempts=4): rospy.loginfo(f"Searching for {object_name} with active perception...") # Search for the object from each observation position for i, pose in enumerate(self.observation_poses[:max_attempts]): rospy.loginfo(f"Moving to observation pose {i+1} / {len(self.observation_poses)}") # Move the robot to the observation position target_pose = Pose() target_pose.position.x = pose['position'][0] target_pose.position.y = pose['position'][1] target_pose.position.z = pose['position'][2] target_pose.orientation.x = pose['orientation'][0] target_pose.orientation.y = pose['orientation'][1] ​target_pose.orientation.z = pose['orientation'][2] target_pose.orientation.w = pose['orientation'][3] self.robot_controller.move_to_pose(target_pose) # Detect the object object_pose = self.object_detector.detect_specific_object(object_name) if object_pose is not None: rospy.loginfo(f"Found {object_name} at position: {object_pose.position.x}, {object_pose.position.y}, {object_pose.position.z}") return object_pose rospy.logwarn(f"Failed to find {object_name} after {max_attempts} attempts") return None ``` This code implements object tracking using a Kalman filter and object exploration by active perception. In object tracking, even when the object is temporarily hidden, tracking continues based on past position and velocity information. Also, in active perception, when the object cannot be found, the robot makes observations from different viewpoints to explore the object. Force sensing module The force sensing module measures the force received by the robot's end effector and plays a role in improving the accuracy of object manipulation. # Calibration of the force sensor The calibration of the force sensor is adjusted so that the sensor indicates zero in the absence of an external force in order to correct the influence of gravity. This allows the external force applied to the end effector to be accurately predicted. The calibration process is performed in the following steps: 1. Zero the sensor on one axis. 2. Rotate the sensor. 3. Zero the sensor on the next axis. 4. Repeat the same process for all axes. After calibration, the local force is converted to the global plane to estimate the upward force applied to the end effector at different rotations. The conversion is performed using the following equation: Fglobal = Tend_effector_to_robot_base × Flocal Here, Fglobal is the force vector in the base coordinate system of the robot, Tend_effector_to_robot_base is the transformation matrix from the coordinate system of the end effector to the base coordinate system of the robot, and Flocal is the force vector in the local coordinate system of the end effector. The calibration of the force sensor is implemented using the following Python code: ```python import numpy as np import rospy import tf2_ros import tf2_geometry_msgs from geometry_msgs.msg import WrenchStamped, TransformStamped from sensor_msgs.msg import JointState class ForceSensorCalibrator: def __init__(self): # ROS-related settings self.force_sub = rospy.Subscriber(" / force_torque_sensor / raw", WrenchStamped, self.force_callback) self.joint_sub = rospy.Subscriber(" / joint_states", JointState, self.joint_callback) self.calibrated_force_pub = rospy.Publisher(" / force_torque_sensor / calibrated", WrenchStamped, queue_size=10) # TF-related settings self.tf_buffer = tf2_ros.Buffer() self.tf_listener = tf2_ros.TransformListener(self.tf_buffer) # Calibration parameters self.gravity_compensation = np.zeros(6) # [fx, fy, fz, tx, ty, tz] self.is_calibrated = False # Latest force sensor data self.latest_force_data = None # Latest joint state self.latest_joint_state = None def force_callback(self, msg): # Save the latest force sensor data self.latest_force_data = msg # If calibrated, publish the corrected force if self.is_calibrated and self.latest_joint_state is not None: calibrated_force = self.calibrate_force(msg) self.calibrated_force_pub.publish(calibrated_force) def joint_callback(self, msg): # Save the latest joint state self.latest_joint_state = msg def calibrate(self): rospy.loginfo("Starting force sensor calibration...") if self.latest_force_data is None or self.latest_joint_state is None: rospy.logwarn("No force data or joint state available. Cannot calibrate.") return False

[0025] # Calibrate each axis axes = ["x", "y", "z"] for axis_idx, axis in enumerate(axes): rospy.loginfo(f"Calibrating {axis} axis...") # Record the current force current_force = np.array( self.latest_force_data.wrench.force.x, self.latest_force_data.wrench.force.y, self.latest_force_data.wrench.force.z, self.latest_force_data.wrench.torque.x, self.latest_force_data.wrench.torque.y, self.latest_force_data.wrench.torque.z ) # Update the gravity compensation value self.gravity_compensation[axis_idx] = current_force[axis_idx] rospy.loginfo(f"{axis} axis calibrated. Offset: {self.gravity_compensation[axis_idx]}") # Prepare to rotate the robot to calibrate the next axis # (In the actual code, the process of rotating the robot appropriately is required) rospy.sleep(2.0) # Wait for stability after rotation rospy.loginfo("Force sensor calibration completed.") rospy.loginfo(f"Gravity compensation values: {self.gravity_compensation}") self.is_calibrated = True return True def calibrate_force(self, force_msg): # Convert the force sensor data to a NumPy array force_array = np.array( force_msg.wrench.force.x, force_msg.wrench.force.y, force_msg.wrench.force.z, force_msg.wrench.torque.x, force_msg.wrench.torque.y, force_msg.wrench.torque.z ) # Apply gravity compensation calibrated_force_array = force_array - self.gravity_compensation # Get the transformation from the end effector to the base coordinate system try: transform = self.tf_buffer.lookup_transform( "base_link", force_msg.header.frame_id, rospy.Time(0), rospy.Duration(1.0) ) # Extract the rotation matrix q = transform.transform.rotation rotation_matrix = self.quaternion_to_rotation_matrix(q) # Convert the force and torque force_local = calibrated_force_array[:3] torque_local = calibrated_force_array[3:] force_global = rotation_matrix @ force_local torque_global = rotation_matrix @ torque_local # Create a message with the transformed force and torque calibrated_msg = WrenchStamped() calibrated_msg.header = force_msg.header calibrated_msg.header.frame_id = "base_link" calibrated_msg.wrench.force.x = force_global[0] calibrated_msg.wrench.force.y = force_global[1] calibrated_msg.wrench.force.z = force_global[2] calibrated_msg.wrench.torque.x = torque_global[0] calibrated_msg.wrench.torque.y = torque_global[1] calibrated_msg.wrench.torque.z = torque_global[2] return calibrated_msg except (tf2_ros.LookupException, tf2_ros.ConnectivityException, tf2_ros.ExtrapolationException) as e: rospy.logwarn(f"Failed to transform force: {e}") # If the transformation fails, return a message with only gravity compensation applied calibrated_msg = WrenchStamped() calibrated_msg.header = force_msg.header calibrated_msg.wrench.force.x = calibrated_force_array[0] calibrated_msg.wrench.force.y = calibrated_force_array[1] calibrated_msg.wrench.force.z = calibrated_force_array[2] calibrated_msg.wrench.torque.x = calibrated_force_array[3] calibrated_msg.wrench.torque.y = calibrated_force_array[4] calibrated_msg.wrench.torque.z = calibrated_force_array[5] return calibrated_msg def quaternion_to_rotation_matrix(self, q): # Conversion from quaternion to rotation matrix x, y, z, w = q.x, q.y, q.z, q.w rotation_matrix = np.array( [1 - 2*y*y - 2*z*z, 2*x*y - 2*z*w, 2*x*z + 2*y*w], [2*x*y + 2*z*w, 1 - 2*x*x - 2*z*z, 2*y*z - 2*x*w], [2*x*z - 2*y*w, 2*y*z + 2*x*w, 1 - 2*x*x - 2*y*y] ) return rotation_matrix ``` This code implements the calibration of the force sensor and the conversion of force from the local coordinate system to the global coordinate system. In the calibration process, the gravity compensation value for each axis is calculated and used to correct the force sensor data. Also, the TF library is used to perform the conversion from the end effector coordinate system to the robot base coordinate system, expressing the force and torque in the global coordinate system.

[0026] # Utilization of Force Feedback Force feedback is utilized in various tasks as follows: 1. Control of liquid injection volume: When injecting liquid, force feedback is used to control the injection volume. Assuming static equilibrium conditions and maintaining low-speed operation, the relationship between force and mass is utilized to estimate the flow rate. Mathematically, it is expressed by the following formula: Fup ≒ mg ΔFup ≒ Δmg Here, Fup is the upward force, m is the mass, g is the acceleration due to gravity, ΔFup is the change in force, and Δm is the change in mass. Using this relationship, the injection volume can be estimated from the change in force. According to the experimental results, when the pitch speed was 4 m / s, the injection accuracy was approximately 5.4 g per 100 g. However, as the pitch speed increased, the accuracy decreased, and at 30 m / s, the error reached approximately 20 g / s. This is because the assumption of static equilibrium breaks down and the mass distribution of the injection medium and the container affects the measurement accuracy. 2. Opening and closing of drawers: When opening and closing a drawer, appropriate force is applied using force feedback. The magnitude and direction of the force are adjusted according to the characteristics of the drawer, such as its weight and friction. For example, when the drawer is heavy or stuck, a greater force needs to be applied. Also, the movement of the drawer is detected and the progress of opening and closing is monitored. 3. Placement of objects: When placing an object, use force feedback to detect contact and place the object with an appropriate force. For example, when placing a mug, use the peak of the upward force as an indicator of a successful placement. This allows the object to be placed stably and prevents dropping or tipping over. 4. Pen Pressure Control: In the drawing task, use force feedback to control the pressure of the pen. By maintaining a uniform pressure, a consistent line thickness and quality can be ensured. Also, the pressure can be adjusted according to the characteristics of the surface, such as hardness and friction. The system continuously manages the force vector along three axes and adjusts the applied force based on the criteria in the knowledge base. The LLM dynamically selects the magnitude and direction of the required force to meet the requirements of a specific downstream task. For example, the knowledge base may specify various force magnitudes to be applied according to the characteristics of the object and the requirements of the task. This approach allows the system to autonomously adjust its actions to a wide range of operating criteria. The utilization of force feedback is implemented with the following Python code: ```python import numpy as np import rospy from geometry_msgs.msg import WrenchStamped, Pose, Twist from std_msgs.msg import Float64 class ForceController: def __init__(self): # ROS-related settings self.force_sub = rospy.Subscriber(" / force_torque_sensor / calibrated", WrenchStamped, self.force_callback) self.velocity_pub = rospy.Publisher(" / robot / cartesian_velocity_command", Twist, queue_size=10) self.poured_amount_pub = rospy.Publisher(" / pouring / amount", Float64, queue_size=10) # Latest force sensor data self.latest_force = None # Parameters for estimating the pouring amount self.initial_force = None self.gravity = 9.8 # m / s^2 # Force control parameters self.force_threshold = 10.0 # N self.contact_threshold = 1.0 # N self.max_velocity = 0.05 # m / s def force_callback(self, msg): # Save the latest force sensor data self.latest_force = msg def start_pouring(self): rospy.loginfo("Starting pouring...") if self.latest_force is None: rospy.logwarn("No force data available. Cannot start pouring.") return False # Record the initial force self.initial_force = np.array( self.latest_force.wrench.force.x, self.latest_force.wrench.force.y, self.latest_force.wrench.force.z ) rospy.loginfo(f"Initial force: {self.initial_force}") return True def update_pouring(self, target_amount): if self.latest_force is None or self.initial_force is None: rospy.logwarn("No force data or initial force available. Cannot update pouring.") return 0.0, False # Get the current force current_force = np.array( self.latest_force.wrench.force.x, self.latest_force.wrench.force.y, self.latest_force.wrench.force.z ) # Estimate the pouring amount from the change in force force_change = current_force[2] - self.initial_force[2] # Change in force in the z-axis direction poured_amount = force_change / self.gravity # kg # Publish the pouring amount amount_msg = Float64() amount_msg.data = poured_amount self.poured_amount_pub.publish(amount_msg) rospy.logdebug(f"Estimated poured amount: {poured_amount} kg") # Determine whether the target amount has been reached is_complete = poured_amount >= target_amount return poured_amount, is_complete def open_drawer_with_force_control(self, direction, max_distance=0.3, max_duration=5.0): rospy.loginfo("Opening drawer with force control...") if self.latest_force is None: rospy.logwarn("No force data available. Cannot open drawer.") return False # Record the initial force initial_force = np.array( self.latest_force.wrench.force.x, self.latest_force.wrench.force.y, self.latest_force.wrench.force.z ) # Vector in the direction to pull the drawer (normalized) direction_norm = np.linalg.norm(direction) if direction_norm > 0: direction = direction / direction_norm # Setting the velocity command twist = Twist() twist.linear.x = direction[0] * self.max_velocity twist.linear.y = direction[1] * self.max_velocity twist.linear.z = direction[2] * self.max_velocity # Pull out the drawer start_time = rospy.Time.now() rate = rospy.Rate(10) # 10Hz moved_distance = 0.0 last_time = start_time while (rospy.Time.now() - start_time).to_sec() < max_duration and moved_distance < max_distance: if self.latest_force is None: rospy.logwarn("Lost force data during drawer opening.") break

[0027] # Get the current force current_force = np.array( self.latest_force.wrench.force.x, self.latest_force.wrench.force.y, self.latest_force.wrench.force.z ) # Calculate the magnitude of the force force_magnitude = np.linalg.norm(current_force - initial_force) # If the force exceeds the threshold, adjust the speed if force_magnitude > self.force_threshold: rospy.logwarn(f"Force threshold exceeded: {force_magnitude} N") # Decrease the speed scale_factor = self.force_threshold / force_magnitude twist.linear.x = direction[0] * self.max_velocity * scale_factor twist.linear.y = direction[1] * self.max_velocity * scale_factor twist.linear.z = direction[2] * self.max_velocity * scale_factor else: # Normal speed twist.linear.x = direction[0] * self.max_velocity twist.linear.y = direction[1] * self.max_velocity twist.linear.z = direction[2] * self.max_velocity # Send the velocity command self.velocity_pub.publish(twist) # Update the moving distance current_time = rospy.Time.now() dt = (current_time - last_time).to_sec() moved_distance += self.max_velocity * dt last_time = current_time rate.sleep() # Send the stop command twist.linear.x = 0.0 twist.linear.y = 0.0 twist.linear.z = 0.0 self.velocity_pub.publish(twist) rospy.loginfo(f"Drawer opening completed. Moved distance: {moved_distance} m") return moved_distance > 0.1 # Consider it successful if it has moved more than 10 cm def place_object_with_force_control(self, target_pose, approach_distance=0.1, contact_force=2.0): rospy.loginfo("Placing object with force control...") if self.latest_force is None: rospy.logwarn("No force data available. Cannot place object.") return False # Calculate approach position approach_pose = Pose() approach_pose.position.x = target_pose.position.x approach_pose.position.y = target_pose.position.y approach_pose.position.z = target_pose.position.z + approach_distance approach_pose.orientation = target_pose.orientation # Move to the approach position # (In the actual code, the process of moving the robot is required) rospy.sleep(1.0) # Wait for the movement to complete # Record the initial force initial_force = np.array( self.latest_force.wrench.force.x, self.latest_force.wrench.force.y, self.latest_force.wrench.force.z ) # Lower the object twist = Twist() twist.linear.z = -0.01 # Lower slowly rate = rospy.Rate(10) # 10Hz contact_detected = False while not contact_detected: if self.latest_force is None: rospy.logwarn("Lost force data during object placement.") break # Get the current force current_force = np.array( self.latest_force.wrench.force.x, self.latest_force.wrench.force.y, self.latest_force.wrench.force.z ) # Calculate the change in force force_change = np.linalg.norm(current_force - initial_force) # Detect contact if force_change > self.contact_threshold: rospy.loginfo(f"Contact detected: {force_change} N") contact_detected = True break # Send velocity command self.velocity_pub.publish(twist) rate.sleep() # Send stop command twist.linear.z = 0.0 self.velocity_pub.publish(twist) if contact_detected: # Press until the target contact force twist.linear.z = -0.005 # Lower more slowly target_force_reached = False while not target_force_reached: if self.latest_force is None: rospy.logwarn("Lost force data during force application.") break # Obtain the current force current_force = np.array( self.latest_force.wrench.force.x, self.latest_force.wrench.force.y, self.latest_force.wrench.force.z ) # Calculate the change in force force_change = np.linalg.norm(current_force - initial_force) # Check if the target force has been reached if force_change >= contact_force: rospy.loginfo(f"Target force reached: {force_change} N") target_force_reached = True break # Send the velocity command self.velocity_pub.publish(twist)

[0028] This code implements various tasks that utilize force feedback (such as controlling the injection volume of liquid, opening and closing drawers, placing objects, and controlling the pressure of a pen). In each task, force feedback is used to apply appropriate forces or estimate the state from force changes. For example, in controlling the injection volume of liquid, the injection volume is estimated from the change in force, and injection is stopped when the target volume is reached. Also, in opening and closing a drawer, the speed is adjusted when the force exceeds a threshold to safely open the drawer. Voice Recognition and Synthesis Module The voice recognition and synthesis module is responsible for recognizing the user's voice instructions and outputting the system's response in voice. # Voice Recognition Voice recognition is the process of converting the user's voice instructions into text. In this system, a cloud-based voice recognition service (e.g., Google Cloud Speech-to-Text) or a locally executed voice recognition engine (e.g., Mozilla DeepSpeech) is used. The voice recognition process is carried out in the following steps: 1. Obtain the voice signal from the microphone. 2. Digitize the voice signal and perform preprocessing (such as noise removal, normalization, etc.). 3. Send the voice data to the voice recognition engine to convert it into text. 4. Send the recognition result to the language processing component. The accuracy of voice recognition is affected by factors such as environmental noise, the speaker's pronunciation, dialect, and the quality of the language model. In this system, noise cancellation technology and beamforming technology are utilized to handle the operating sounds of the robot and environmental noise. # Voice Synthesis Speech synthesis is the process of converting the system's response from text to speech. In this system, a cloud-based speech synthesis service (e.g., Google Cloud Text-to-Speech) or a locally executed speech synthesis engine (e.g., Festival, eSpeak) is used. The speech synthesis process is carried out in the following steps: 1. Receive a text response from the language processing component. 2. Send the text to the speech synthesis engine and convert it into speech data. 3. Output the speech data from the speaker. The quality of speech synthesis is affected by factors such as the performance of the synthesis engine, voice selection, and control of prosody (intonation, accent, rhythm). In this system, the latest speech synthesis technology is utilized to generate speech with natural intonation and pronunciation. # Speech Dialogue Management Speech dialogue management is the process of managing the speech dialogue with the user and realizing a natural conversation. In this system, the following functions are provided: 1. Speaker identification: Identify the voices of multiple users and provide personalized responses. 2. Emotion recognition: Recognize the emotion from the user's speech and provide appropriate responses. 3. Ambient sound recognition: Recognize ambient sounds (e.g., warning sounds, door opening and closing sounds) to understand the situation. 4. Dialogue state management: Manage the context of the dialogue and generate appropriate responses. 5. Interruption handling: Detect the user's interruption and respond appropriately. Speech dialogue management cooperates with the language processing component to realize a natural interaction with the user. The speech recognition and synthesis module is implemented with the following Python code: ```python import speech_recognition as sr import pyttsx3 import numpy as np import rospy from std_msgs.msg import String from google.cloud import speech from google.cloud import texttospeech class SpeechModule: def __init__(self, use_cloud=True): # ROS-related settings self.speech_pub = rospy.Publisher(" / speech / recognized", String, queue_size=10) self.text_sub = rospy.Subscriber(" / speech / synthesize", String, self.synthesize_callback) # Speech recognition-related settings self.recognizer = sr.Recognizer() self.microphone = sr.Microphone() # Ambient noise adjustment with self.microphone as source: self.recognizer.adjust_for_ambient_noise(source) # Whether to use cloud services self.use_cloud = use_cloud if use_cloud: # Initialize Google Cloud Speech-to-Text client It should be noted that there is an undefined variable `sr` in the original code. You may need to correct it according to the actual situation.self.speech_client = speech.SpeechClient() # Initialization of Google Cloud Text-to-Speech client self.tts_client = texttospeech.TextToSpeechClient() else: # Initialization of local text-to-speech engine self.engine = pyttsx3.init() # Voice settings voices = self.engine.getProperty('voices') self.engine.setProperty('voice', voices[0].id) # 0: male, 1: female self.engine.setProperty('rate', 150) # Speaking rate self.engine.setProperty('volume', 0.8) # Volume # Timer for speech recognition self.timer = rospy.Timer(rospy.Duration(0.1), self.recognize_speech) # Speaker profiles for speaker identification self.speaker_profiles = {} # Dialogue state self.dialogue_state = { "context": [], "last_query": "", "last_response": "", "is_listening": True } def recognize_speech(self, event=None): if not self.dialogue_state["is_listening"]: return try: with self.microphone as source: audio = self.recognizer.listen(source, timeout=1.0, phrase_time_limit=5.0) if self.use_cloud: # Google Cloud Speech-to-Text を使用 audio_data = audio.get_raw_data() audio_config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_code="ja-JP", enable_automatic_punctuation=True, model="command_and_search" ) response = self.speech_client.recognize( config=audio_config, audio=speech.RecognitionAudio(content=audio_data) ) if response.results: text = response.results[0].alternatives[0].transcript confidence = response.results[0].alternatives[0].confidence rospy.loginfo(f"Recognized: {text} (confidence: {confidence})") if confidence > 0.7: # Recognized result is published msg = String() msg.data = text self.speech_pub.publish(msg)

[0029] # Update the dialogue state self.dialogue_state["is_listening"] = False self.dialogue_state["last_response"] = text self.dialogue_state["context"].append({"role": "assistant", "content": text}) if self.use_cloud: # Use Google Cloud Text-to-Speech synthesis_input = texttospeech.SynthesisInput(text=text) voice = texttospeech.VoiceSelectionParams( language_code="ja-JP", ssml_gender=texttospeech.SsmlVoiceGender.NEUTRAL ) audio_config = texttospeech.AudioConfig( audio_encoding=texttospeech.AudioEncoding.LINEAR16 ) response = self.tts_client.synthesize_speech( input=synthesis_input, voice=voice, audio_config=audio_config ) # Save the audio data to a temporary file with open("output.wav", "wb") as out: out.write(response.audio_content) # Play the audio import subprocess subprocess.call(["aplay", "output.wav"]) else: # Use the local text-to-speech engine self.engine.say(text) self.engine.runAndWait() # Update the dialogue state self.dialogue_state["is_listening"] = True def identify_speaker(self, audio): # Implementation of speaker identification # (In the actual code, an algorithm for speaker identification is required) return "unknown" def recognize_emotion(self, audio): # Implementation of Emotional Recognition # (In the actual code, an algorithm for emotional recognition is required) return "neutral" def recognize_environmental_sound(self, audio): # Implementation of Environmental Sound Recognition # (In the actual code, an algorithm for environmental sound recognition is required) return [] ``` This code implements the basic functions of the speech recognition and synthesis module. In speech recognition, the SpeechRecognition library or Google Cloud Speech-to-Text is used to convert the user's speech into text. In speech synthesis, the pyttsx3 library or Google Cloud Text-to-Speech is used to convert text into speech. It also provides functions for managing the dialogue state and realizing natural conversations with the user. Robot Control System The robot control system integrates feedback from the language processing, vision system, force sensing module, and speech recognition and synthesis module, and plays a role in controlling the robot's operations. # ROS Settings and Operations In this embodiment, ROS is used to control the robot. The settings and operations of ROS are performed in the following steps: 1. Start the Kinova ROS Kortex driver to establish communication between the ROS network and the Kinova Gen3 robot. This node enables communication within the ROS network, publishes several topics that subscribers can access, and provides services that can be called to change the robot's configuration. The base joint is updated at a frequency of 40Hz. 2. Start the Robotiq 2F-140mm gripper node at 50 Hz. This node sets up a communication link with the gripper via a USB connection and starts an action server that enables precise control of the gripper and the exchange of operation data. 3. Start the vision module node. Use the "classes" variable to identify the target poses of selected objects in the environment. This variable can be updated dynamically and adapt to changes in the scene. The pose coordinates of the objects are published at approximately 1 / 3 Hz. This is mainly due to the processing time for Grounding DINO to detect the objects and establish the bounding boxes. 4. Start the force sensing node at 100 Hz and provide readings of multi-axis force and torque localized to the ATI force sensor transducer. The readings are transformed to match the robot's global base frame using a quaternion-based 3x3 rotation matrix and provide raw and average values over the past five time steps for the fixed degrees of freedom. 5. Start the speech recognition and synthesis node to recognize the user's voice commands and output the system's responses as voice. 6. ROS continuously processes multi-modal feedback data from the language processing, vision system, force sensing module, and speech recognition and synthesis module. The configuration and operation of ROS are implemented in a launch file as follows: ```xml <launch> <!-- Startup of Kinova Gen3 robot driver --> <include file="$(find kortex_driver) / launch / kortex_driver.launch"> <arg name="robot_name" value="my_gen3" / > <arg name="ip_address" value="192.168.1.10" / > <arg name="cyclic_data_publish_rate" value="40" / > < / include> <!-- Startup of Robotiq gripper driver --> <node name="robotiq_2f_140_driver" pkg="robotiq_2f_gripper_control" type="Robotiq2FGripperRtuNode.py" output="screen"> <param name="comport" value=" / dev / ttyUSB0"> <param name="baud" value="115200"> < / node> <!-- Startup of Azure Kinect driver --> <include file="$(find azure_kinect_ros_driver) / launch / driver.launch"> <arg name="color_resolution" value="1080P" / > <arg name="depth_mode" value="NFOV_UNBINNED" / > <arg name="fps" value="30" / > < / include> <!-- Startup of vision system --> <node name="object_detector" pkg="my_robot_vision" type="object_detector.py" output="screen"> <param name="model_path" value="$(find my_robot_vision) / models"> <param name="classes" value="mug,kettle,spoon,coffee_container,drawer_handle"> < / node> <!-- Startup of force sensor driver --> <node name="force_sensor_driver" pkg="ati_force_sensor" type="force_sensor_driver" output="screen"> <param name="ip_address" value="192.168.1.11"> <param name="port" value="49152"> <param name="frame_id" value="force_sensor_link"> <param name="sample_rate" value="100"> < / node> <!-- Startup of force module --> <node name="force_controller" pkg="my_robot_force" type="force_controller.py" output="screen" / > <!-- Startup of speech recognition and synthesis module --> <node name="speech_module" pkg="my_robot_speech" type="speech_module.py" output="screen"> <param name="use_cloud" value="true"> < / node> <!-- Startup of language processing component --> <node name="language_processor" pkg="my_robot_language" type="language_processor.py" output="screen"> <param name="api_key" value="your-api-key"> <param name="model" value="gpt-4-0613"> < / node>

[0030] <!-- Startup of multimodal integration module --> <node name="multimodal_integrator" pkg="my_robot_integration" type="multimodal_integrator.py" output="screen" / > <!-- Startup of robot control system --> <node name="robot_controller" pkg="my_robot_control" type="robot_controller.py" output="screen" / > < / launch> ``` This launch file starts each component of the system (Kinova Gen3 robot driver, Robotiq gripper driver, Azure Kinect driver, vision system, force sensing module, speech recognition and synthesis module, language processing component, multimodal integration module, robot control system) and establishes communication between them. # Implementation of Safety Constraints The robot's motion is based on a 6 - degree - of - freedom twist command for controlling speed and a gripper procedure with variable speed and force for opening and closing. This allows for integrating hard - coded safety constraints such as maximum speed and force limits, and the boundaries of the working space. The specific constraints are as follows: 1. The linear speed is limited within ±0.05 m / s. 2. The angular speed is limited within ±60° / s. 3. The force of the end - effector is limited to 20N. 4. The end - effector is restricted within the boundaries of a pre - defined working space (x = [0.0, 1.1], y = [-0.3, 0.3], z = [0, 1.0]). Since these constraints are coded into the basic motion primitives, errors in the language model cannot override these constraints. Also, the position of the end - effector is checked by a publisher at a frequency of 10Hz for future time steps. The implementation of safety constraints is done in Python code as follows: ```python import numpy as np import rospy from geometry_msgs.msg import Twist, Pose from std_msgs.msg import Bool class SafetyConstraints: def __init__(self): # Parameters for safety constraints self.max_linear_velocity = 0.05 # m / s self.max_angular_velocity = 1.047 # rad / s (60 deg / s) self.max_force = 20.0 # N # Bounds of the workspace self.workspace_bounds = { 'x': [0.0, 1.1], 'y': [-0.3, 0.3], 'z': [0.0, 1.0] } # ROS-related settings self.velocity_sub = rospy.Subscriber(" / robot / cartesian_velocity_command_unsafe", Twist, self.velocity_callback) self.pose_sub = rospy.Subscriber(" / robot / cartesian_pose", Pose, self.pose_callback) self.force_sub = rospy.Subscriber(" / force_torque_sensor / calibrated", WrenchStamped, self.force_callback) self.velocity_pub = rospy.Publisher(" / robot / cartesian_velocity_command", Twist, queue_size=10) self.safety_violation_pub = rospy.Publisher(" / robot / safety_violation", Bool, queue_size=10) # Latest state self.latest_pose = None self.latest_force = None # Safety check timer self.timer = rospy.Timer(rospy.Duration(0.1), self.check_safety) def velocity_callback(self, msg): # Limit the velocity command according to safety constraints safe_velocity = self.apply_velocity_constraints(msg) # Publish the safe velocity command self.velocity_pub.publish(safe_velocity) def pose_callback(self, msg): # Save the latest pose self.latest_pose = msg def force_callback(self, msg): # Save the latest force self.latest_force = msg def apply_velocity_constraints(self, velocity): # Create a copy of the velocity command safe_velocity = Twist() # Limit the linear velocity safe_velocity.linear.x = np.clip(velocity.linear.x, -self.max_linear_velocity, self.max_linear_velocity) safe_velocity.linear.y = np.clip(velocity.linear.y, -self.max_linear_velocity, self.max_linear_velocity) safe_velocity.linear.z = np.clip(velocity.linear.z, -self.max_linear_velocity, self.max_linear_velocity) # Limit of angular velocity safe_velocity.angular.x = np.clip(velocity.angular.x, -self.max_angular_velocity, self.max_angular_velocity) safe_velocity.angular.y = np.clip(velocity.angular.y, -self.max_angular_velocity, self.max_angular_velocity) safe_velocity.angular.z = np.clip(velocity.angular.z, -self.max_angular_velocity, self.max_angular_velocity) return safe_velocity def check_workspace_bounds(self, pose): # Check the boundaries of the workspace x_in_bounds = self.workspace_bounds['x'][0] <= pose.position.x <= self.workspace_bounds['x'][1] y_in_bounds = self.workspace_bounds['y'][0] <= pose.position.y <= self.workspace_bounds['y'][1] z_in_bounds = self.workspace_bounds['z'][0] <= pose.position.z <= self.workspace_bounds['z'][1] return x_in_bounds and y_in_bounds and z_in_bounds def check_force_limits(self, force): # Check force limits force_magnitude = np.sqrt( force.wrench.force.x**2 + force.wrench.force.y**2 + force.wrench.force.z**2 ) return force_magnitude <= self.max_force def check_safety(self, event=None): if self.latest_pose is None or self.latest_force is None: return # Check workspace bounds workspace_safe = self.check_workspace_bounds(self.latest_pose) # Check force limits force_safe = self.check_force_limits(self.latest_force) # Publish whether there is a safety violation safety_violation = not (workspace_safe and force_safe) msg = Bool() msg.data = safety_violation self.safety_violation_pub.publish(msg)

[0031] if safety_violation: # If there is a safety violation, stop the robot stop_velocity = Twist() self.velocity_pub.publish(stop_velocity) rospy.logwarn("Safety violation detected! Robot stopped.") if not workspace_safe: rospy.logwarn("Workspace bounds violated.") if not force_safe: rospy.logwarn("Force limits exceeded.") ``` This code implements the safety constraints of the robot. It limits the velocity commands according to the safety constraints, checks the workspace bounds and force limits. When a safety violation is detected, it stops the robot and displays a warning message. This makes the robot's operation safe and reliable. # Implementation of the feedback loop The robot control system utilizes visual and force feedback to adjust the operation in real time. The implementation of the feedback loop is as follows: 1. Visual feedback: Using the information from the vision system, continuously update the position and orientation of the object. When the object moves, identify the new position and adjust the motion plan. Also, when an unexpected obstacle appears, plan an avoidance behavior. Visual feedback is important in tasks where position accuracy is crucial, such as object grasping and placement. 2. Force feedback: Using the information from the force sensor, control the interaction with the object. For example, when pouring a liquid, estimate the pouring amount from the change in force and stop pouring when the target amount is reached. Also, when opening a drawer, adjust the magnitude and direction of the force to achieve smooth opening and closing. Force feedback is important especially in situations where visual information is limited (e.g., when the field of view is blocked). 3. Audio feedback: Using the information from the speech recognition / synthesis module, continue the interaction with the user. Update the task plan or provide additional information according to the user's instructions. Audio feedback is important especially for realizing a natural interaction with the user. 4. Multimodal integration: Integrate information from different modalities such as vision, force, and audio to achieve a richer environmental understanding and generate adaptive behaviors. For example, identify the position of an object from visual information, estimate the weight of the object from force information, and integrate these information to determine the appropriate grasping force and position. The feedback loop includes updating the position (p) and orientation (q) of the end effector at 40 Hz, enabling the robot to respond to disturbances (e.g., movement of the cup by the user). The implementation of the feedback loop is performed using the following Python code: ```python import numpy as np import rospy import tf2_ros from geometry_msgs.msg import Pose, Twist, WrenchStamped from std_msgs.msg import String from sensor_msgs.msg import JointState class FeedbackController: def __init__(self): # ROS related settings self.pose_sub = rospy.Subscriber(" / robot / cartesian_pose", Pose, self.pose_callback) self.joint_sub = rospy.Subscriber(" / joint_states", JointState, self.joint_callback) self.force_sub = rospy.Subscriber(" / force_torque_sensor / calibrated", WrenchStamped, self.force_callback) self.object_pose_sub = rospy.Subscriber(" / vision / tracked_object_poses / target", Pose, self.object_pose_callback) self.speech_sub = rospy.Subscriber(" / speech / recognized", String, self.speech_callback) self.velocity_pub = rospy.Publisher(" / robot / cartesian_velocity_command_unsafe", Twist, queue_size=10) self.speech_pub = rospy.Publisher(" / speech / synthesize", String, queue_size=10) # TF related settings self.tf_buffer = tf2_ros.Buffer() self.tf_listener = tf2_ros.TransformListener(self.tf_buffer) # Latest state self.latest_pose = None self.latest_joints = None self.latest_force = None self.target_object_pose = None self.latest_speech = None # Task state self.task_state = { "current_task": None, "target_pose": None, "is_tracking": False, "is_force_controlled": False, "is_speech_controlled": False } # Feedback loop timer self.timer = rospy.Timer(rospy.Duration(0.025), self.feedback_loop) # 40Hz def pose_callback(self, msg): # Save the latest pose self.latest_pose = msg def joint_callback(self, msg): # Save the latest joint state self.latest_joints = msg def force_callback(self, msg): # Save the latest power self.latest_force = msg def object_pose_callback(self, msg): # Save the latest target object pose self.target_object_pose = msg def speech_callback(self, msg): # Save the latest voice recognition result self.latest_speech = msg.data # Process voice commands self.process_speech_command(msg.data) def process_speech_command(self, command): # Process voice commands if "stop" in command or "halt" in command or "stop" in command: # Stop the robot self.stop_robot() self.speak("Stop") elif "continue" in command or "resume" in command or "continue" in command: # Resume the task self.task_state["is_tracking"] = True self.speak("Resume task") elif "a little more" in command and "pour" in command: # Add more self.speak("Pour more") # (In the actual code, additional pouring processing is required) elif "enough" in command or "enough" in command: # Stop pouring self.speak("Stop pouring") # (In the actual code, you need to take action to stop the pouring.) def speak(self, text): # Publish voice synthesis messages msg = String() msg.data = text self.speech_pub.publish(msg) def stop_robot(self): # Stop the robot self.task_state["is_tracking"] = False self.task_state["is_force_controlled"] = False # Send stop command twist = Twist() self.velocity_pub.publish(twist) def track_object(self): if self.latest_pose is None or self.target_object_pose is None: return

[0032] This code implements a feedback loop that integrates visual, force, and audio feedback. In visual feedback, the position of the target object is tracked and the movement is directed towards it. In force feedback, the speed is adjusted according to the magnitude and direction of the force. In audio feedback, the task is controlled according to the user's voice commands. By integrating these feedbacks, the task can be continuously executed while adapting to changes and uncertainties in the environment. Multimodal Integration Module The multimodal integration module is responsible for integrating information from different modalities such as vision, force, and audio, and realizing richer environmental understanding and adaptive behavior generation. # Information Integration between Modalities The multimodal integration module provides a mechanism to appropriately weight and integrate information from each modality. For example, it identifies the position of an object from visual information, estimates the weight of the object from force information, and integrates this information to determine the appropriate grasping force and position. Also, it understands the user's intention from audio information and combines it with visual information to identify the appropriate target object. The information integration between modalities is implemented with the following Python code: ```python import numpy as np import rospy from geometry_msgs.msg import Pose, WrenchStamped from std_msgs.msg import String, Float64MultiArray from sensor_msgs.msg import Image from cv_bridge import CvBridge class MultimodalIntegrator: def __init__(self): # ROS-related settings self.vision_sub = rospy.Subscriber(" / vision / object_poses", Pose, self.vision_callback) self.force_sub = rospy.Subscriber(" / force_torque_sensor / calibrated", WrenchStamped, self.force_callback) self.speech_sub = rospy.Subscriber(" / speech / recognized", String, self.speech_callback) self.integrated_info_pub = rospy.Publisher(" / multimodal / integrated_info", Float64MultiArray, queue_size=10) # Latest modal information self.latest_vision = None self.latest_force = None self.latest_speech = None # Integrated information self.integrated_info = { "object_position": None, "object_weight": None, "user_intent": None, "confidence": { "vision": 0.0, "force": 0.0, "speech": 0.0 } } # Modal weights self.modality_weights = { "vision": 0.5, "force": 0.3, "speech": 0.2 } # Integrated timer self.timer = rospy.Timer(rospy.Duration(0.1), self.integrate_modalities) def vision_callback(self, msg): # Save the latest vision information self.latest_vision = msg # Extract the object's position from the vision information object_position = np.array( msg.position.x, msg.position.y, msg.position.z ) # Update the integrated information self.integrated_info["object_position"] = object_position self.integrated_info["confidence"]["vision"] = 0.8 # Confidence of vision information (example) def force_callback(self, msg): # Save the latest force sense information self.latest_force = msg # Estimate the object's weight from the force sense information force_magnitude = np.sqrt( msg.wrench.force.x**2 + msg.wrench.force.y**2 + msg.wrench.force.z**2 ) # Estimate weight considering gravitational acceleration estimated_weight = force_magnitude / 9.8 # kg # Update integrated information self.integrated_info["object_weight"] = estimated_weight self.integrated_info["confidence"]["force"] = 0.7 # Confidence of force sensing information (example) def speech_callback(self, msg): # Save the latest speech information self.latest_speech = msg.data # Extract the user's intent from the speech information user_intent = self.extract_user_intent(msg.data) # Update integrated information self.integrated_info["user_intent"] = user_intent self.integrated_info["confidence"]["speech"] = 0.6 # Confidence of speech information (example) def extract_user_intent(self, speech_text): # Extract the user's intent from the speech text intent = { "action": None, "object": None, "location": None, "modifier": None } # Simple intent extraction (in actual code, more advanced natural language processing is required) if "take" in speech_text or "hold" in speech_text: intent["action"] = "pick" elif "put" in speech_text: intent["action"] = "place" elif "pour" in speech_text: intent["action"] = "pour" if "cup" in speech_text or "mug" in speech_text: intent["object"] = "mug" elif "kettle" in speech_text: intent["object"] = "kettle" elif "spoon" in speech_text: intent["object"] = "spoon" if "table" in speech_text: intent["location"] = "table" elif "drawer" in speech_text: intent["location"] = "drawer" if "slowly" in speech_text: intent["modifier"] = "slow" elif "quickly" in speech_text: intent["modifier"] = "fast" return intent def integrate_modalities(self, event=None): # Integrate information from each modality # Publish the integrated information integrated_data = [] if self.integrated_info["object_position"] is not None: integrated_data.extend(self.integrated_info["object_position"]) else: integrated_data.extend([0.0, 0.0, 0.0]) if self.integrated_info["object_weight"] is not None: integrated_data.append(self.integrated_info["object_weight"]) else: integrated_data.append(0.0)

[0033] # Quantify the user's intention if self.integrated_info["user_intent"] is not None: action_code = 0.0 if self.integrated_info["user_intent"]["action"] == "pick": action_code = 1.0 elif self.integrated_info["user_intent"]["action"] == "place": action_code = 2.0 elif self.integrated_info["user_intent"]["action"] == "pour": action_code = 3.0 integrated_data.append(action_code) else: integrated_data.append(0.0) # Add reliability information integrated_data.append(self.integrated_info["confidence"]["vision"]) integrated_data.append(self.integrated_info["confidence"]["force"]) integrated_data.append(self.integrated_info["confidence"]["speech"]) # Publish integrated information msg = Float64MultiArray() msg.data = integrated_data self.integrated_info_pub.publish(msg) def update_modality_weights(self, vision_weight, force_weight, speech_weight): # Update the weights of modalities total_weight = vision_weight + force_weight + speech_weight if total_weight > 0: self.modality_weights["vision"] = vision_weight / total_weight self.modality_weights["force"] = force_weight / total_weight self.modality_weights["speech"] = speech_weight / total_weight else: rospy.logwarn("Total weight is zero or negative. Weights not updated.") ``` This code implements a multimodal integration module that integrates information from vision, force, and speech modalities to achieve a richer understanding of the environment. It extracts the position of an object from visual information, estimates the weight of the object from force information, and extracts the user's intention from speech information. These pieces of information are integrated to achieve a more accurate understanding of the environment, taking into account the reliability of each modality. # Cross-modal Learning Cross-modal learning is a process of learning the relationships between different modalities to achieve information complementarity. For example, by learning the relationship between visual information and force information, it is possible to estimate the weight of an object from visual information or estimate the shape of an object from force information. Also, by learning the relationship between speech information and visual information, it is possible to identify the target object from a speech instruction. Cross-modal learning is implemented with the following Python code: ```python import numpy as np import rospy from geometry_msgs.msg import Pose, WrenchStamped from std_msgs.msg import String, Float64MultiArray from sklearn.linear_model import LinearRegression from sklearn.ensemble import RandomForestRegressor from sklearn.svm import SVR class CrossModalLearner: def __init__(self): # ROS-related settings self.vision_sub = rospy.Subscriber(" / vision / object_poses", Pose, self.vision_callback) self.force_sub = rospy.Subscriber(" / force_torque_sensor / calibrated", WrenchStamped, self.force_callback) self.speech_sub = rospy.Subscriber(" / speech / recognized", String, self.speech_callback) self.predicted_weight_pub = rospy.Publisher(" / multimodal / predicted_weight", Float64MultiArray, queue_size=10) # Training data self.vision_data = [] self.force_data = [] self.speech_data = [] # Training models self.vision_to_force_model = RandomForestRegressor() self.force_to_vision_model = RandomForestRegressor() self.speech_to_vision_model = None # Another approach is needed for natural language processing # Trained flag self.is_trained = False # Training timer self.timer = rospy.Timer(rospy.Duration(60.0), self.train_models) # Train every 60 seconds def vision_callback(self, msg): # Add visual information to the dataset vision_features = msg.position.x, msg.position.y, msg.position.z, msg.orientation.x, msg.orientation.y, msg.orientation.z, msg.orientation.w self.vision_data.append(vision_features) # If trained, predict force from visual information if self.is_trained: predicted_force = self.predict_force_from_vision(vision_features) self.publish_predicted_weight(predicted_force) def force_callback(self, msg): # Add force sensor information to the dataset force_features = ​msg.wrench.force.x, msg.wrench.force.y, msg.wrench.force.z, msg.wrench.torque.x, msg.wrench.torque.y, msg.wrench.torque.z self.force_data.append(force_features) def speech_callback(self, msg): # Add voice information to the dataset # (In the actual code, the process of converting text to feature vectors is required) self.speech_data.append(msg.data) def train_models(self, event=None): # If there is enough data, train the model if len(self.vision_data) > 10 and len(self.force_data) > 10: rospy.loginfo("Training cross-modal models...") # Use the latest data recent_vision_data = np.array(self.vision_data[-100:]) recent_force_data = np.array(self.force_data[-100:])

[0034] # Train a model to predict force / torque information from visual information self.vision_to_force_model.fit(recent_vision_data, recent_force_data) ​ # Train a model to predict visual information from force information self.force_to_vision_model.fit(recent_force_data, recent_vision_data) self.is_trained = True rospy.loginfo("Cross-modal models trained successfully.") else: rospy.logwarn("Not enough data to train cross-modal models.") def predict_force_from_vision(self, vision_features): # Predict force information from visual information vision_features_array = np.array([vision_features]) predicted_force = self.vision_to_force_model.predict(vision_features_array)[0] return predicted_force def predict_vision_from_force(self, force_features): # Predict visual information from force information force_features_array = np.array([force_features]) predicted_vision = self.force_to_vision_model.predict(force_features_array)[0] return predicted_vision def publish_predicted_weight(self, predicted_force): # Calculate weight from the predicted force force_magnitude = np.sqrt( predicted_force[0]**2 + predicted_force[1]**2 + predicted_force[2]**2 ) # Estimate weight considering gravitational acceleration estimated_weight = force_magnitude / 9.8 # kg # Publish the predicted weight msg = Float64MultiArray() msg.data = [estimated_weight] self.predicted_weight_pub.publish(msg) ``` This code implements cross-modal learning that learns the relationship between visual information and force-sensing information and predicts one from the other. It learns a model that predicts force-sensing information from visual information and a model that predicts visual information from force-sensing information. This allows for compensation from the other modality even when information in one modality is missing. # Multimodal Inference Multimodal inference is a process of making inferences about the environment and situation based on integrated information. For example, integrating visual information and force-sensing information to estimate the material or contents of an object, or integrating visual information and audio information to estimate the user's intention. Multimodal inference is implemented in Python code as follows: ```python import numpy as np import rospy from geometry_msgs.msg import Pose, WrenchStamped from std_msgs.msg import String, Float64MultiArray from sklearn.ensemble import RandomForestClassifier class MultimodalReasoner: def __init__(self): # ROS-related settings self.integrated_info_sub = rospy.Subscriber(" / multimodal / integrated_info", Float64MultiArray, self.integrated_info_callback) self.object_property_pub = rospy.Publisher(" / multimodal / object_property", String, queue_size=10) self.user_intent_pub = rospy.Publisher(" / multimodal / user_intent", String, queue_size=10) # Object property classifier self.material_classifier = RandomForestClassifier() self.content_classifier = RandomForestClassifier() # Object property database self.material_database = { "plastic": {"weight_range": [0.05, 0.2], "hardness": "medium"}, "ceramic": {"weight_range": [0.2, 0.5], "hardness": "high"}, "metal": {"weight_range": [0.3, 1.0], "hardness": "high"}, "glass": {"weight_range": [0.2, 0.6], "hardness": "high"}, "wood": {"weight_range": [0.1, 0.4], "hardness": "medium"} } self.content_database = { "empty": {"weight_change": [0.0, 0.1]}, "water": {"weight_change": [0.1, 1.0]}, "coffee": {"weight_change": [0.1, 0.3]}, "sugar": {"weight_change": [0.1, 0.2]} } # Interpretation rules for user intentions self.intent_rules = { "pick": { "mug": "pick up the mug", "kettle": "pick up the kettle", "spoon": "pick up the spoon" }, "place": { "mug": "place the mug", "kettle": "place the kettle", "spoon": "place the spoon" }, "pour": { "kettle": "pour water from the kettle" } } def integrated_info_callback(self, msg): # Analyze integrated information integrated_data = msg.data # Object position object_position = integrated_data[0:3] # Object weight object_weight = integrated_data[3] # User's intention action_code = integrated_data[4] # Confidence vision_confidence = integrated_data[5] force_confidence = integrated_data[6] speech_confidence = integrated_data[7] # Estimate the material of the object material = self.infer_material(object_weight) # Estimate the content of the object content = self.infer_content(object_weight) # Interpret the user's intention user_intent = self.interpret_user_intent(action_code) # Publish the inference result self.publish_object_property(material, content) self.publish_user_intent(user_intent)

[0035] def infer_material(self, weight): # Infer the material of the object from the weight for material, properties in self.material_database.items(): if properties["weight_range"][0] <= weight <= properties["weight_range"][1]: return material return "unknown" def infer_content(self, weight): # Infer the content of the object from the change in weight # (In the actual code, it is necessary to track the change in weight) weight_change = 0.2 # Example for content, properties in self.content_database.items(): if properties["weight_change"][0] <= weight_change <= properties["weight_change"][1]: return content return "unknown" def interpret_user_intent(self, action_code): # Interpret the user's intent from the action code action = "unknown" if action_code == 1.0: action = "pick" elif action_code == 2.0: action = "place" elif action_code == 3.0: action = "pour" object_type = "mug" # Example (in actual code, need to estimate the type of object) if action in self.intent_rules and object_type in self.intent_rules[action]: return self.intent_rules[action][object_type] else: return "unknown intent" def publish_object_property(self, material, content): # Publish the properties of the object msg = String() msg.data = f"Material: {material}, Content: {content}" self.object_property_pub.publish(msg) def publish_user_intent(self, user_intent): # Publish the user's intent msg = String() msg.data = user_intent self.user_intent_pub.publish(msg) ``` This code implements multimodal inference that infers the material, contents, user intent, etc. of an object based on integrated information. It estimates the material from the weight of the object and the contents from changes in weight. Also, it interprets the user's intent from the action code and the type of object. This enables richer environmental understanding and adaptive behavior generation.

Example

[0036] The following describes specific examples of the present invention. The following experiments were conducted using Categorical AI of New York General Group. Categorical AI partially uses the Claude-3.7-Sonnet model operated by Anthropic, and can perform high-precision calculations in numerical analysis, efficient solution of optimization problems, automatic program generation, bug detection and correction, etc., and can be used from the following URL: https: / / www.newyorkgeneralgroup.com / ouraimodels In this example, in order to verify in detail the effectiveness of the proposed multimodal feedback integrated knowledge search enhanced robot control system, an advanced simulation experiment was conducted using Gymnasium (formerly OpenAI Gym). Gymnasium is a Python library that provides a standard interface for the development of reinforcement learning algorithms and benchmarks, and can simulate robot control tasks in various environments. In this experiment, the functions of Gymnasium were greatly extended to construct an advanced custom environment that integrates multimodal feedback and retrieval augmented generation (RAG).

[0037] Construction of Simulation Environment The simulation environment precisely mimics an actual kitchen environment, with various objects (such as mugs, coffee beans, kettles, spoons, drawers, refrigerators, cupboards, etc.) placed for performing the task of "making coffee". The objects within the environment have positions, postures, and physical properties (weight, friction coefficient, elastic coefficient, thermal conductivity, etc.), and interact based on advanced physical simulations. Additionally, simulations of non-rigid substances such as liquids (water, coffee) and powders (coffee beans) are also implemented to enable more realistic task execution. For constructing the environment, a custom environment class "KitchenEnvironment" that inherits from the basic classes of Gymnasium was implemented. This class has the following characteristics: 1. Observation Space: - Visual information: RGB image (1920×1080 pixels, 24-bit color depth) and depth map (1920×1080 pixels, 16-bit depth) - Force sense information: 6-dimensional vector (3-dimensional force and 3-dimensional torque) - Audio information: Speech recognition result in text form and acoustic features (frequency spectrum) - Environment state: Positions, postures, and physical states (temperature, water content, etc.) of objects 2. Action Space: - Joint angles of the robot (7 degrees of freedom) - Opening and closing state of the gripper (continuous value) - Position and posture of the end effector (6-dimensional) - Force control parameters (maximum force, compliance, etc.) - High-level actions ("grasp", "lift", "pour", etc.) 3. Physical Simulation: - Rigid body dynamics: High-precision rigid body simulation using PyBullet - Fluid Mechanics: Liquid simulation using the SPH (Smoothed Particle Hydrodynamics) method - Thermodynamics: Simulation of heat conduction and convection (such as the temperature change of coffee) - Contact Mechanics: Simulation of contact characteristics such as friction, elasticity, and viscosity 4. Environmental Uncertainty: - Randomization of the initial position of the object (standard deviation 5 cm) - Variation in physical properties (weight ±10%, friction coefficient ±20%, etc.) - Sensor noise (Vision: Gaussian noise σ = 0.01, Force sense: Gaussian noise σ = 0.1 N) - Actuator delay (10 - 50 ms) and uncertainty (±2% of the command value) - Dynamic changes in the environment (such as the movement of an object by another agent) 5. Reward Function: - Task completion reward: Reward for each subtask (+10) and reward for the completion of the entire task (+100) - Efficiency reward: Reward inversely proportional to the task completion time (maximum +50) - Safety reward: Penalty for excessive force or speed (maximum -50) - Precision reward: Reward proportional to the accuracy of the liquid injection volume (maximum +30) The environment is implemented on Python 3.9 and uses major libraries such as PyBullet 3.2.1, NumPy 1.22.3, and OpenCV 4.6.0. The simulation is run on a workstation equipped with an Intel Core i9-12900K CPU, an NVIDIA RTX 3090 GPU, and 64GB of RAM, and the simulation progresses at a real-time factor of 0.8 (80% of real time). Detailed Implementation of the System The proposed system is implemented with a modular architecture. Each component is developed and tested independently and then integrated. The overall architecture of the system is based on the Publish-Subscribe pattern similar to ROS2 (Robot Operating System 2), and the communication between each component is performed through asynchronous messaging. The detailed implementation of the main components is as follows: # 1. Language Processing Component The language processing component provides functions for natural language understanding, task decomposition, code generation, and Retrieval-Augmented Generation (RAG). The implementation details are as follows: - Language Model: GPT-4 (via OpenAI API) is used as the main language model. As a backup, the locally executable Llama 2 (70B) model is also implemented. - Embedding Model: OpenAI's text-embedding-ada-002 (1536 dimensions) is used for the vector representation of the text. - RAG Implementation: Implemented using the Langchain framework. Document chunking (chunk size 512 tokens, overlap 50 tokens), vectorization, and approximate nearest neighbor search using Faiss are implemented. - Prompt Engineering: Dedicated prompt templates for task decomposition, code generation, and error handling are developed. The prompts include the state of the system, descriptions of the environment, past action histories, error information, etc. - Task Decomposition Algorithm: Hierarchical task decomposition is implemented. A three-layer structure that decomposes high-level tasks into medium-level tasks and medium-level tasks into low-level tasks. The dependencies between each level are represented by a directed acyclic graph (DAG). - Conditional Probability Model: The dependencies between subtasks are represented as conditional probability P(T_j|T_i). The probability values are learned from examples in the knowledge base and implemented as a Bayesian network. - Code Generation: Safe code generation and verification using Python AST (Abstract Syntax Tree). The generated code undergoes static analysis and dynamic checks before execution in a sandbox environment. - Error Handling: A dedicated module for error detection, diagnosis, and recovery. Utilizes a database of error patterns and past error recovery cases. The language processing component monitors the environmental state at a 40Hz cycle and updates the task plan as needed. Task decomposition and code generation are executed asynchronously, and the results are stored in a queue. The system selects and executes the code most suitable for the current environmental state. # 2. Visual System The visual system generates a 3D representation of the environment and performs object detection, segmentation, and pose estimation. The implementation details are as follows: - Object Detection: A real-time object detector (mAP 0.56@0.5IoU) based on YOLOv7, fine-tuned for 30 types of kitchen objects. The detection speed is 30FPS. - Instance Segmentation: A hybrid approach combining Mask R-CNN (ResNet-101 backbone) and the Segment Anything model. The segmentation accuracy is mIoU 0.78. - Language-Visual Module: Object detection based on language instructions using the CLIP (Contrastive Language-Image Pretraining) model, supporting attribute-based instructions such as "red mug" and "big kettle". - 3D Reconstruction: Generation of 3D point clouds by combining depth maps and segmentation masks. Plane and surface fitting using the RANSAC (Random Sample Consensus) algorithm. - Pose Estimation: 6DoF pose estimation using the ICP (Iterative Closest Point) algorithm and a pre-trained shape model. The estimation accuracy is on average position ±5 mm and orientation ±3 degrees. - Temporal Tracking: Multi-object tracking using a Kalman filter and the Hungarian algorithm. Tracking is maintained for up to 5 seconds against occlusion. - Uncertainty Estimation: Estimation of probability distributions (multivariate Gaussian distributions) for each detection and pose estimation. When uncertainty is high, active perception (active viewpoint change) is triggered. - Visual Attention Mechanism: Dynamic control of visual attention based on task relevance. Priority attention is given to objects related to the current subtask. The vision system performs image processing at a cycle of 20 Hz and publishes the detection results and pose estimations. Computationally intensive processes (such as 3D reconstruction) are executed at a low frequency (5 Hz) as needed. # 3. Force Sensing Module The force sensing module measures the forces and torques received by the robot's end effector and improves the accuracy of object manipulation. The implementation details are as follows: - Force Sensing Filtering: Hybrid filtering combining a Kalman filter and a moving average filter. It achieves both removal of high-frequency noise and detection of abrupt changes. - Calibration Algorithm: An automatic calibration process including gravity compensation and temperature drift correction. Each axis of the 6-axis force sensor is calibrated independently to minimize the cross-talk effect. - Contact Detection: Detection of contact events based on an adaptive threshold. An SVM model classifies the types of contact (collision, sliding, grasping). - Object Property Estimation: A Bayesian model that estimates the weight, hardness, friction coefficient, etc. of an object from force feedback. The estimation accuracy is weight ±5%, hardness ±10%, and friction ±15%. - Force Control Algorithm: A meta-controller that switches between impedance control, hybrid position / force control, and adaptive control according to the situation. The control frequency is 1 kHz. - Liquid operation model: A model that estimates the liquid injection volume from the change in force. A two-stage estimation algorithm considering static equilibrium conditions and dynamic effects. The estimation accuracy is ±5 ml (when injecting 100 ml). - Anomaly detection: An anomaly detector that detects deviations from the expected force pattern. Calculates an anomaly score based on the Mahalanobis distance. - Force sense memory: An associative memory that stores past force sense patterns and recalls appropriate force control parameters in similar situations. The force sense module performs data acquisition and basic filtering at a high frequency of 1 kHz, and executes high-order processing (such as characteristic estimation and anomaly detection) at a frequency of 100 Hz. The force sense data is stored in a ring buffer, and the data for the past 5 seconds can be accessed.

[0038] # 4. Speech Recognition and Synthesis Module The speech recognition and synthesis module enables natural interaction with the user. The implementation details are as follows: - Speech recognition engine: A speech recognition system based on the Whisper Large v3 model. Supports bilingual recognition of Japanese and English. The word error rate (WER) is 5.2% for Japanese and 3.8% for English. - Speech synthesis engine: A high-quality speech synthesis system combining the FastSpeech 2 model and the HiFi-GAN vocoder. The naturalness MOS (Mean Opinion Score) is 4.3 / 5.0. - Speaker identification: A speaker identification system using x-vector and PLDA. Can identify 20 speakers with an accuracy of 98%. Has a function to register new speakers. - Emotion recognition: Recognition of 7 types of emotions (happiness, sadness, anger, fear, disgust, surprise, neutral) using the prosodic and spectral features of speech. The recognition accuracy is 72%. - Ambient sound recognition: An ambient sound recognition system based on YAMNet. Can recognize 50 types of kitchen-related sounds (boiling sound, timer sound, cutting sound, etc.). The recognition accuracy is 85%. - Dialogue Management: A hybrid dialogue management system that combines a hierarchical finite state machine and the RASA dialogue management framework. Implements functions such as context maintenance, repair strategies, and confirmation requests. - Adaptive Speech Processing: An adaptive algorithm that adjusts recognition parameters according to environmental noise. Achieves stable recognition even in an environment with a low SNR (Signal-to-Noise Ratio). - Multimodal Dialogue: A dialogue system that integrates visual information and audio information. Performs an interpretation considering the visual context for utterances containing directive words such as "Get that". The speech recognition and synthesis module acquires speech at a sampling rate of 16 kHz and performs real-time recognition processing. The recognition result is obtained within 100 ms and sent to the dialogue management system. Speech synthesis generates synthesized speech within 200 ms from text input. # 5. Multimodal Integration Module The multimodal integration module integrates information from different modalities such as vision, force sense, and speech, and realizes richer environmental understanding and adaptive behavior generation. The implementation details are as follows: - Probabilistic Integration Framework: A probabilistic integration framework that combines a Bayesian network and a particle filter. Explicitly considers the uncertainty of each modality. - Dynamic Weighting Algorithm: Dynamic weighting based on the reliability of each modality. The reliability is calculated based on the state of the sensor, environmental conditions, and task requirements. - Cross-modal Learning: Cross-modal representation learning using a variational autoencoder (VAE). Trains a model to predict one modality from another modality. - Multimodal Attention Mechanism: A cross-modal attention mechanism based on the Transformer architecture. Learns the relevance between modalities and optimizes information integration. - Missing Modality Completion: A mechanism that predicts missing information from other modalities when some modalities are unavailable (sensor failure, occlusion, etc.). - Temporal integration: A temporal alignment algorithm for integrating modalities at different time scales (such as high-speed force perception and low-speed vision). - Uncertainty propagation: Propagate the uncertainty of each modality throughout the integration process to quantify the uncertainty of the final state estimation. - Multimodal memory: An episodic memory system that stores past multimodal experiences and recalls appropriate integration strategies in similar situations. The multimodal integration module runs at a cycle of 50 Hz, integrates the latest information from each modality, and updates the state estimation of the environment. The integration result is sent to the robot control system and the language processing component. # 6. Robot Control System The robot control system integrates feedback from the language processing, vision system, force perception module, speech recognition / synthesis module, and multimodal integration module to control the robot's movements. The implementation details are as follows: - Hierarchical control architecture: A four-layer structure consisting of a task planning layer, motion planning layer, trajectory generation layer, and low-level control layer. Each layer operates at a different time scale (from 0.1 Hz to 1 kHz). - Motion primitive library: Implements 40 basic motions (such as linear movement, rotation, grasping, lifting, pouring, etc.). Each primitive is parameterized and adjustable according to the situation. - Trajectory optimization: Trajectory optimization based on model predictive control (MPC) and optimal control theory. Multi-objective optimization considering energy efficiency, smoothness, and safety. - Adaptive control: An adaptive control algorithm that adjusts control parameters according to the environment and tasks. Model identification and control parameter adjustment are executed in parallel. - Safety constraint implementation: Implements safety constraints such as speed limits (±0.05 m / s), force limits (±20 N), and workspace limits (x = [0.0, 1.1], y = [-0.3, 0.3], z = [0, 1.0]). Includes a mechanism for detecting and avoiding constraint violations. - Obstacle Avoidance: A real-time obstacle avoidance algorithm that takes into account dynamic obstacles. A hybrid approach combining the Potential Field method and RRT* (Rapidly-exploring Random Tree Star). - Failure Detection and Recovery: A mechanism that detects failures during operation execution and executes appropriate recovery strategies. Implemented a database of failure patterns and a library of recovery strategies. - Learning-based Control: A learning-based controller that combines imitation learning and reinforcement learning. Implemented initial policy learning from demonstrations and policy improvement from execution experience. The robot control system consists of multiple loops operating at different frequencies. The low-level control loop is executed at 1kHz and is responsible for trajectory tracking and safety monitoring. Motion planning and trajectory generation are executed at 50Hz and update the plan according to changes in the environment. The task plan is updated as needed and is usually executed at a frequency of 0.1 - 1Hz. # 7. Knowledge Base The knowledge base is a curated database containing examples of verified low-level and high-level actions. The implementation details are as follows: - Knowledge Representation: Knowledge is represented in a structured JSON format. Each entry includes task description, preconditions, postconditions, execution code, success / failure cases, uncertainty information, etc. - Hierarchical Organization: Knowledge is organized in a three-layer hierarchical structure (high-level tasks, middle-level tasks, low-level tasks). Links between the layers facilitate task decomposition and synthesis. - Indexing: Indexing from multiple perspectives (task types, object types, environmental conditions, etc.). Implemented fast vector search using Faiss. - Uncertainty Modeling: Record uncertainty information (success probability, expected error, etc.) for each action example. Quantification of uncertainty using Bayesian models. - Automatic Expansion Function: A function that automatically adds successful and failed cases of task execution to the knowledge base. Includes duplicate detection and quality evaluation mechanisms. - Version Management: A version management system that manages the change history of the knowledge base. Implement functions such as change tracking, rollback, and branch management. - Distributed Access: A distributed knowledge base that can be accessed by multiple agents simultaneously. Implement a consistency maintenance and conflict resolution mechanism. - Semantic Search: A semantic search function based on natural language queries. Hybrid search combining embedding-based similarity calculation and structured query language. The knowledge base is implemented as a hybrid database combining MongoDB (document store) and Faiss (vector index). Currently, the knowledge base contains the following entries: - Basic Operation Primitives: 100 types (linear movement, rotation, grasping, etc.) - Object Manipulation Primitives: 80 types (lifting, placing, pouring, etc.) - Special Operation Primitives: 50 types (liquid injection, powder scooping, drawer opening / closing, etc.) - Medium-Level Tasks: 60 types (coffee preparation, toast making, etc.) - High-Level Tasks: 30 types (breakfast preparation, cake making, etc.) - Error Handling and Recovery: 40 types (object dropping, liquid spilling, etc.) - Environmental Adaptation Examples: 35 types (object position change, obstacle appearance, etc.) Experimental Setup In the experiment, the task of "making coffee" was targeted, and the performance of the system was evaluated under the following detailed conditions: # 1. Level of Environmental Uncertainty The environmental uncertainty was set at the following 3 levels, and the performance of the system was evaluated at each level: - Low Uncertainty (Level 1): - Randomization of Object Position: Standard Deviation 2 cm - Variation in Physical Properties: Weight ±5%, Friction Coefficient ±10% - Sensor noise: vision (σ = 0.005), force sense (σ = 0.05 N) - All objects are placed within the field of view - Medium uncertainty (Level 2): - Randomization of object position: standard deviation 5 cm - Variation in physical properties: weight ±10%, coefficient of friction ±20% - Sensor noise: vision (σ = 0.01), force sense (σ = 0.1 N) - Some objects (mug or coffee beans) are placed in the drawer - High uncertainty (Level 3): - Randomization of object position: standard deviation 10 cm - Variation in physical properties: weight ±20%, coefficient of friction ±30% - Sensor noise: vision (σ = 0.02), force sense (σ = 0.2 N) - Multiple objects (mug, coffee beans) are placed in the drawer or cabinet - Introduction of dynamic obstacles (moving objects) - Temporary sensor failure (temporary loss of vision or force sense information) # 2. Abstraction level of instructions The abstraction level of instructions from the user was set at the following three levels, and the understanding and execution capabilities of the system were evaluated at each level: - High abstraction (Level A): - "Make coffee" - "Prepare a morning drink" - "Need caffeine" - Medium abstraction (Level B): - "Put coffee in the mug" - "Grind the coffee beans and pour hot water" - "Prepare drip coffee" - Low abstraction (Level C): - "Pick up the mug, scoop the coffee beans, and pour hot water from the kettle." - "Take out the red mug from the right shelf, put in 2 spoons of coffee beans, and pour 150 ml of 80-degree hot water." - "Make coffee according to the following procedure: 1. Prepare a mug, 2. Put in coffee beans, 3. Pour hot water." # 3. Feedback Conditions To evaluate the system's performance under different feedback conditions, the following conditions were set: - Single Modality: - Vision only: Disable force and audio feedback - Force only: Disable vision and audio feedback - Audio only: Disable vision and force feedback - Dual Modality: - Vision + Force: Disable audio feedback - Vision + Audio: Disable force feedback - Force + Audio: Disable vision feedback - Full Modality: - Vision + Force + Audio: Enable all feedback - Adaptive Integration: - Dynamically adjust the weight of each modality according to the situation - Optimal modality selection based on uncertainty - Information collection by active perception

[0039] # 4. RAG Configuration To evaluate the effect of Retrieval-Augmented Generation (RAG), the following configurations were compared: - Without RAG: - Do not use the knowledge base and rely only on the general knowledge of the language model - Basic RAG: - Search based on simple cosine similarity - Use the top 5 search results - Fixed chunk size (512 tokens) - Extended RAG: - Hybrid search (cosine similarity + BM25) - Dynamic number of searches (3 - 10 depending on task complexity) - Adaptive chunking (semantic chunking) - Re - ranking mechanism - Adaptive RAG: - Context - aware query expansion - Multi - hop search (additional search based on initial search results) - Search considering feedback information - Knowledge base update based on execution history # 5. Evaluation Metrics To comprehensively evaluate the system's performance, the following metrics were measured: - Task completion rate: Completion rate of the overall task and subtasks - Execution time: Execution time of the overall task and subtasks - Accuracy metrics: - Liquid injection accuracy: Error from the target volume (ml) - Object placement accuracy: Error from the target position (mm) - Force control accuracy: Error from the target force (N) - Efficiency metrics: - Energy consumption: Sum of the squares of the robot's joint torques - Motion smoothness: Sum of the squares of accelerations - Computational efficiency: CPU / GPU usage rate and processing time - Safety metrics: - Maximum force: Maximum contact force during task execution - Speed violations: Number of times exceeding the safety speed limit - Number of collisions: Number of unintended contacts with the environment - Adaptability metrics: - Recovery rate: The success rate of recovery from errors - Exploration efficiency: The discovery rate of hidden objects - Adaptation to environmental changes: The success rate of adaptation to dynamic changes The experiment was executed 20 times for each combination of conditions (uncertainty level × instruction abstraction level × feedback condition × RAG configuration) to ensure the statistical significance of the results. Also, the order of the experimental conditions was randomized to minimize bias due to learning effects. Experimental Results and Detailed Analysis # Comprehensive Analysis of Task Completion Rate The results of a detailed analysis of the completion rate of the "make coffee" task under different experimental conditions are shown below: 1. Interaction between uncertainty level and feedback condition: Task completion rate by uncertainty level of the proposed system (full modality + adaptive RAG): - Low uncertainty: 95% (19 successes out of 20) - Medium uncertainty: 85% (17 successes out of 20) - High uncertainty: 70% (14 successes out of 20) Task completion rate in a high-uncertainty environment by feedback condition: - Visual only: 45% (9 successes out of 20) - Haptic only: 30% (6 successes out of 20) - Auditory only: 15% (3 successes out of 20) - Visual + haptic: 65% (13 successes out of 20) - Visual + auditory: 50% (10 successes out of 20) - Haptic + auditory: 35% (7 successes out of 20) - Full modality: 70% (14 successes out of 20) - Adaptive integration: 75% (15 successes out of 20) From these results, it became clear that as environmental uncertainty increased, the task completion rate decreased, but the impact was mitigated by the integration of multimodal feedback. In particular, the combination of vision and force sense was effective, and it was confirmed that the adaptive integration approach further improved performance. As a result of statistical analysis (two-way ANOVA), the main effect of the uncertainty level (F(2,456)=78.3, p<0.001), the main effect of the feedback condition (F(7,456)=62.5, p<0.001), and the interaction between the two (F(14,456)=12.7, p<0.001) were all significant. This indicates that the effect of the feedback condition varies depending on the uncertainty level. 2. Interaction between Instruction Abstraction Level and RAG Configuration: Task Completion Rate by Instruction Abstraction Level (Full Modality Condition): - High Abstraction Level ("Make coffee"): - Without RAG: 50% (10 successes out of 20 attempts) - Basic RAG: 70% (14 successes out of 20 attempts) - Extended RAG: 80% (16 successes out of 20 attempts) - Adaptive RAG: 85% (17 successes out of 20 attempts) - Medium Abstraction Level ("Put coffee in a mug"): - Without RAG: 65% (13 successes out of 20 attempts) - Basic RAG: 75% (15 successes out of 20 attempts) - Extended RAG: 85% (17 successes out of 20 attempts) - Adaptive RAG: 90% (18 successes out of 20 attempts) - Low Abstraction Level (Detailed Procedure): - Without RAG: 80% (16 successes out of 20 attempts) - Basic RAG: 85% (17 successes out of 20 attempts) - Extended RAG: 90% (18 successes out of 20 attempts) - Adaptive RAG: 95% (19 successes out of 20 attempts) From these results, it became clear that the higher the level of abstraction of the instructions, the more significant the effect of RAG. In particular, for instructions with a high level of abstraction, a difference of as much as 35% occurred between the case without RAG and the case of adaptive RAG. This indicates that the structured knowledge provided by RAG plays an important role when decomposing abstract instructions into specific subtasks. As a result of statistical analysis (two-way ANOVA), the main effect of instruction abstraction level (F(2,228)=45.2, p<0.001), the main effect of RAG configuration (F(3,228)=38.7, p<0.001), and the interaction effect between the two (F(6,228)=8.3, p<0.001) were all significant. 3. Detailed analysis of failure modes: As a result of a detailed analysis of the causes of task failure, the following patterns were revealed: - Object detection failure: 28% (among all failures) - Main causes: Visual similarity, partial occlusion, lighting conditions - Improvement measures: Multi-view integration, active perception, strengthening of temporal consistency - Grasping failure: 22% - Main causes: Slippery surface, unstable posture, weight misestimation - Improvement measures: Adaptive grasping force control, multi-finger grasping, pre-grasp posture optimization - Liquid manipulation failure: 18% - Main causes: Flow rate estimation error, poor control of container tilt, temperature change - Improvement measures: Closed-loop flow control, enhanced visual-force integration, refinement of prediction model - Planning failure: 15% - Main causes: Incomplete task decomposition, misrecognition of dependencies, non-adaptation to environmental changes - Improvement measures: Strengthening of hierarchical planning, dynamic replanning, context-aware RAG - System integration failure: 10% - Main causes: Communication delay between modules, problems with asynchronous processing, resource contention - Improvement measures: Optimization of the messaging system, priority-based scheduling, distributed processing - Others: 7% - Issues specific to simulation, random failures, etc. From this analysis, it became clear that visual recognition and object manipulation (especially liquid manipulation) are the main causes of failure. To address these problems, it was suggested that the integration of multimodal feedback and adaptive control is particularly important. # Detailed analysis of task decomposition and execution The results of a detailed analysis of the decomposition and execution process of the "make coffee" task are shown below: 1. Qualitative evaluation of task decomposition: Examples of task decomposition with different RAG configurations for the high-level instruction "make coffee": - Without RAG: 1. Find a mug 2. Find coffee beans 3. Prepare hot water 4. Put coffee in (Incomplete decomposition, lack of specific subtasks) - Basic RAG: 1. Find a mug 2. Find coffee beans 3. Find a kettle 4. Place the mug in the appropriate position 5. Put coffee beans into the mug 6. Pour hot water (Includes basic subtasks but lacks detail) - Adaptive RAG: 1. Find a mug (search the drawer or cupboard if necessary) 2. Find coffee beans (open the container if necessary) 3. Find a spoon 4. Find a kettle (boil water if necessary) 5. Place the mug in a stable location 6. Scoop an appropriate amount of coffee beans with a spoon 7. Put the coffee beans into the mug 8. Pour an appropriate amount of hot water from the kettle (monitor the temperature and amount) 9. Stir if necessary (Comprehensive decomposition including detailed and conditional actions) When using Adaptive RAG, conditional actions ("if necessary...") according to the state of the environment were included, and it was confirmed that the ability to handle uncertainty was high. Also included were important subtasks that are often overlooked in other approaches, such as "find a spoon". 2. Success rate of subtask execution: Success rate of each subtask (Adaptive RAG, full modality conditions): - Find a mug: 90% - Normal location: 95% - In drawer: 85% - In pantry: 80% - Find coffee beans: 92% - Find a spoon: 95% - Find a kettle: 98% - Place the mug: 97% - Scoop coffee beans: 88% - Put coffee beans in: 93% - Pour hot water: 85% - Accuracy (error from target amount): ±8 ml From these results, it became clear that "find a mug" (especially when it is hidden) and "pour hot water" are relatively difficult subtasks. In these subtasks, the integration of multimodal feedback was particularly important. 3. Dependency processing between subtasks: Example of dependency between subtasks using a conditional probability model: - P(searching for a mug in the drawer | the mug is not in its normal position) = 0.85 - P(searching for a mug in the cupboard | the mug is not in the drawer) = 0.90 - P(boiling water | the water in the kettle is cold) = 0.95 - P(opening the coffee container | the coffee beans are not visible) = 0.88 These conditional probabilities were set based on the knowledge obtained through RAG and the patterns learned from execution experience. The system was able to dynamically adjust the plan based on these probabilities and adapt to the changes in the environmental state. 4. Execution time analysis: Average execution time (seconds) for each subtask: - Finding a mug: 12.3 (standard position), 28.7 (when hidden) - Finding coffee beans: 10.5 - Finding a spoon: 8.2 - Finding a kettle: 9.1 - Placing the mug: 5.4 - Scooping coffee beans: 15.8 - Putting coffee beans in: 7.3 - Pouring hot water: 22.6 Average execution time for all tasks: 95.7 seconds (without parallel execution), 78.3 seconds (with parallel execution) Example of parallel execution: By executing independent subtasks in parallel, such as "preparing the mug and coffee beans while boiling water from the kettle", the overall execution time could be shortened.

[0040] # Detailed analysis of multimodal feedback The results of a detailed analysis of the integration and effects of multimodal feedback are shown below: 1. Analysis of modality integration patterns: Contribution degree (weight) of each modality in different subtasks: - Finding the mug: - Vision: 0.75 - Force sense: 0.05 - Voice: 0.20 (Vision is the main, supplemented by voice-based position information) - Opening the drawer: - Vision: 0.40 - Force sense: 0.55 - Voice: 0.05 (Force sense is the main, assisted by vision for position adjustment) - Scooping coffee beans: - Vision: 0.60 - Force sense: 0.35 - Voice: 0.05 (Cooperation between vision and force sense) - Pouring hot water: - Vision: 0.30 - Force sense: 0.65 - Voice: 0.05 (Force sense is the main, assisted by vision for position adjustment) These weights were dynamically adjusted by an adaptive integration algorithm. The optimal weights were determined based on the reliability of each modality (sensor state, environmental conditions, etc.) and the requirements of the task. 2. Accuracy of cross-modal prediction: Accuracy of predicting one modality from another modality: - Vision → Force sense prediction: - Object weight prediction: ±15% (estimation from visual features) - Surface friction prediction: ±25% (estimation from visual texture) - Force sense → Visual prediction: - Object shape prediction: 60% accuracy (estimated from tactile exploration) - Liquid level prediction: ±12% (estimated from weight change) - Voice → Visual prediction: - Object position prediction: 45% accuracy (estimated from sound source localization) - Event detection: 75% accuracy (event estimation from characteristic sounds) These predictions were generated by a cross-modal learning model (based on variational autoencoders). The model learned the relationships between different modalities from past experiences and was able to make predictions from one modality when the other was not available. 3. Performance during modality loss: Task completion rate when some modalities are not available (medium uncertainty environment): - All modalities available: 85% - Visual loss (temporary): - Without complementation: 40% - With cross-modal complementation: 65% - Force sense loss (temporary): - Without complementation: 55% - With cross-modal complementation: 70% - Voice loss (temporary): - Without complementation: 75% - With cross-modal complementation: 80% From these results, it was confirmed that the cross-modal complementation mechanism significantly improves the robustness against sensor failures and temporary information loss. In particular, complementation using information from force sense and voice was effective when visual information was missing. 4. Effect of temporal integration: Effect of integrating modalities at different time scales: - High-speed feedback (1 kHz): Improved stability of force control - Force control error: ±0.5 N (with temporal integration) vs ±1.2 N (without) - Contact detection delay: 5 ms (with) vs 25 ms (without) - Medium-speed feedback (30 Hz): Improved accuracy of visual tracking - Position estimation error: ±3 mm (with temporal integration) vs ±8 mm (without) - Success rate of dynamic object tracking: 85% (with) vs 60% (without) - Low-speed feedback (5 Hz): Improved efficiency of plan update - Adaptation time to environmental changes: 1.2 s (with temporal integration) vs 3.5 s (without) - Reprogramming success rate: 90% (with) vs 75% (without) The temporal integration algorithm was able to appropriately integrate feedback at different frequencies, achieving both high-speed response and long-term consistency. In particular, by combining high-speed force feedback and medium-speed visual feedback, the accuracy and stability of object manipulation were improved. # Detailed analysis of RAG The results of a detailed analysis of the effects and operating mechanisms of Retrieval-Augmented Generation (RAG) are shown below: 1. Analysis of retrieval accuracy and relevance: Retrieval accuracy (relevance score, 0 - 1 scale) with different RAG configurations: - Basic RAG (cosine similarity only): - Average relevance score: 0.68 - Number of highly relevant documents (score > 0.8) among the top 5: 1.8 - Extended RAG (hybrid search): - Average relevance score: 0.79 - Number of highly relevant documents among the top 5: 3.2 - Adaptive RAG (Context-Aware): - Average Relevance Score: 0.85 - Number of Highly Relevant Documents among the Top 5: 4.1 The relevance score was calculated based on the degree of agreement with criteria pre-evaluated by human evaluators. Adaptive RAG was able to retrieve more relevant documents through query expansion and utilization of context information. 2. Relationship between Knowledge Base Size and RAG Performance: Task Completion Rate (High Uncertainty Environment) when Varying the Size of the Knowledge Base: - 50 Entries: 55% - 100 Entries: 62% - 200 Entries: 68% - 400 Entries: 73% - 800 Entries: 75% - 1600 Entries: 76% From these results, it was confirmed that as the size of the knowledge base increases, the task completion rate improves, but the improvement in performance becomes gradual at around 800 entries or more. This suggests that not only the quantity of knowledge but also its quality and diversity are important. 3. Computational Efficiency of RAG: Computation Time (Milliseconds) for Different RAG Configurations: - Basic RAG: - Embedding Computation: 45ms - Vector Search: 15ms - Context Construction: 10ms - Total: 70ms - Extended RAG: - Embedding Computation: 45ms - Hybrid Search: 25ms - Re-ranking: 20ms - Context Construction: 15ms - Total: 105ms - Adaptive RAG: - Query Expansion: 15ms - Embedding Calculation: 45ms - Hybrid Search: 25ms - Multi-hop Search: 40ms - Re-ranking: 20ms - Context Construction: 20ms - Total: 165ms These calculation times are the measured values on a workstation equipped with an Intel Core i9-12900K CPU, 32GB RAM, and an NVIDIA RTX 3090 GPU. The embedding calculation was processed in parallel on the GPU, and the search was executed on the CPU. Although Adaptive RAG has the highest computational cost, by introducing a caching mechanism (reusing the results of similar queries), the average calculation time could be reduced to approximately 85ms. 4. Specific Contribution Examples of RAG: Specific Cases where RAG was Particularly Effective: - Search Strategy for Hidden Objects: In the situation of "the mug cannot be found", RAG provides a strategy of "look in the drawer or pantry" from past similar cases. The search success rate has improved from 40% to 85%. - Improvement in Liquid Injection Accuracy: In the task of "pouring hot water", RAG provides an appropriate pouring angle and speed from similar liquid injection examples. The injection accuracy has improved from ±15ml to ±8ml. - Error Recovery Strategy: In the situation of "spilling coffee beans", RAG provides a strategy of "collecting with a spoon" from similar error recovery examples. The recovery success rate has improved from 50% to 85%. - Operation of Unknown Objects: For the "French press" encountered for the first time, RAG provides an appropriate operation method from similar coffee maker operation examples. The task completion rate has improved from 30% to 70%. These examples show that RAG has the ability to adapt knowledge learned from past experiences to new situations, rather than just information retrieval. 5. Analysis of knowledge transfer: Effect of knowledge transfer between similar tasks: - "Making coffee" → "Putting in tea": - Without transfer (without RAG): 60% success rate - With transfer (adaptive RAG): 85% success rate - "Taking a mug" → "Taking a glass": - Without transfer: 75% success rate - With transfer: 95% success rate - "Opening a drawer" → "Opening a refrigerator": - Without transfer: 65% success rate - With transfer: 90% success rate From these results, it was confirmed that knowledge transfer through RAG significantly improves the learning and execution of new tasks. In particular, the transfer effect was remarkable between tasks with different operation targets but similar basic operation patterns. # Detailed analysis of safety and efficiency The results of a detailed analysis of the safety and efficiency of the system are shown below: 1. Effectiveness of safety constraints: Comparison of incident occurrence rates with and without safety constraints: - Without safety constraints: - Excessive force (>20N): Occurs in 15.3% of task executions - High-speed operation (>0.5m / s): Occurs in 22.7% of task executions - Movement outside the working space: Occurred in 8.5% of task executions - Collision: Occurred in 18.9% of task executions - With safety constraints: - Excessive force: Occurred in 0.5% of task executions (97% reduction) - High-speed operation: Occurred in 0.8% of task executions (96% reduction) - Movement outside the working space: Occurred in 0% of task executions (100% reduction) - Collision: Occurred in 3.2% of task executions (83% reduction) It was confirmed that most incidents were prevented by the implementation of safety constraints. The remaining incidents were mainly due to dynamic environmental changes (such as unexpected object movement). 2. Analysis of energy efficiency: Energy consumption in different approaches (sum of squares of joint torques, relative units): - Baseline (without RAG, single modality): 1.00 (reference) - With RAG, single modality: 0.85 (15% reduction) - Without RAG, multimodality: 0.80 (20% reduction) - With RAG, multimodality: 0.65 (35% reduction) Main factors for improving energy efficiency: - Smoother trajectory planning (minimizing acceleration changes) - Selection of appropriate gripping force (avoiding excessive force) - Efficient operation sequence (reducing wasted movements) - Predictive control (avoiding sudden corrections by look-ahead) The combination of multimodal feedback and RAG achieved the highest energy efficiency. This indicates that the appropriate use of feedback and knowledge leads to more efficient motion planning and execution. 3. Analysis of computational efficiency: Computational load per system component (CPU / GPU usage rate): - Language processing component: - CPU: 15 - 25% - GPU: 60 - 80% (during language model inference) - Memory: 4 - 6GB - Vision system: - CPU: 10 - 15% - GPU: 30 - 50% (during object detection and segmentation) - Memory: 2 - 3GB - Force sensing module: - CPU: 5 - 8% - GPU: < 5% - Memory: 0.5 - 1GB - Multimodal integration module: - CPU: 8 - 12% - GPU: 10 - 20% - Memory: 1 - 2GB - Robot control system: - CPU: 10 - 15% - GPU: < 5% - Memory: 1 - 2GB Overall peak usage rate: - CPU: 45 - 60% - GPU: 70 - 90% - Memory: 12 - 18GB From these measurement values, it was confirmed that language model inference and vision processing are the components with the highest computational load. As areas for optimization, model quantization, introduction of batch processing, dynamic allocation of computational resources, etc. can be considered.

[0041] 4. Latency analysis: End - to - end latency (time from instruction input to the start of robot movement): - High-level abstraction instruction ("Make coffee"): - Without RAG: 3.2 seconds - Adaptive RAG: 3.8 seconds - Medium-level abstraction instruction ("Put coffee in a mug"): - Without RAG: 2.1 seconds - Adaptive RAG: 2.5 seconds - Low-level abstraction instruction (detailed procedure): - Without RAG: 1.3 seconds - Adaptive RAG: 1.6 seconds Breakdown of latency by component (in the case of adaptive RAG and high-level abstraction instruction): - Speech recognition: 0.3 seconds - Language understanding: 0.5 seconds - RAG processing: 0.6 seconds - Task decomposition: 1.2 seconds - Code generation: 0.8 seconds - Preparation for execution: 0.4 seconds Although using RAG results in a slight increase in latency, it is offset by the improvement in task completion rate and the reduction in execution time. In addition, through the caching mechanism and prefetch optimization, the latency of RAG could be reduced by approximately 40%. 5. Analysis of real-time performance: Response time to environmental changes: - Detection of object movement and plan update: - Vision only: 450 ms - Multimodal: 320 ms - Response to unexpected obstacles: - From detection to start of avoidance behavior: 180 ms - Complete trajectory replanning: 520 ms - Response to user instruction changes: - From speech recognition to plan update: 850 ms - Until the start of execution of the new plan: 1200 ms These response times are fast enough compared to human reaction times (about 250 ms for visual stimuli and about 170 ms for auditory stimuli), and it has been confirmed that natural interaction and safe operation can be achieved. Conclusion and Future Outlook As a result of detailed simulation experiments using the Gymnasium, the proposed multi-modal feedback integrated knowledge search enhanced robot control system has shown high effectiveness in the following aspects: 1. High-level instruction understanding and execution ability: The proposed system was able to understand abstract instructions such as "make coffee" and decompose them into appropriate subtasks for execution. In particular, by utilizing the knowledge base with RAG, the understanding and execution ability of abstract instructions were significantly improved (from 50% without RAG to 85% with adaptive RAG). 2. Integration effect of multi-modal feedback: By integrating vision, force sense, and voice, a significantly higher task completion rate was achieved compared to single modality (from a maximum of 45% for single modality in a high-uncertainty environment to 75% for multi-modal). In particular, the combination of vision and force sense was effective, and the adaptive integration approach further improved the performance. 3. Adaptability to environmental uncertainty: The proposed system showed high adaptability to uncertainties such as randomization of object positions, variations in physical properties, and sensor noise. In particular, a task completion rate of 70% was achieved even in a high-uncertainty environment, and the combination of multi-modal feedback and RAG contributed significantly to the improvement of adaptability. 4. Efficient utilization of the knowledge base and knowledge transfer: By efficiently utilizing the knowledge base through RAG, the success rate of tasks was improved. Also, knowledge transfer between similar tasks was confirmed, and the learning and execution of new tasks were made more efficient (from 60 - 75% without transfer to 85 - 95% with transfer). 5. Improvement in Safety and Efficiency: By implementing safety constraints, incidents such as excessive force, speed, movement outside the working space, and collisions were significantly reduced (by 83 - 100%). Also, the combination of multimodal feedback and RAG improved energy efficiency by 35%. From these results, it was confirmed that the proposed system has high effectiveness in executing complex tasks in unpredictable environments. In particular, it was revealed that the combination of the integration of multimodal feedback and RAG greatly contributes to the improvement of environmental understanding and adaptation ability. Future research issues and prospects include the following: 1. Verification on a real robot: Verify the effectiveness confirmed in the simulation environment using a real robot. It is necessary to develop technologies to address the simulation-to-real gap. 2. Expansion to more complex tasks: Consider the application to complex tasks that manipulate multiple objects (e.g., cooking, assembling objects). In particular, the evaluation of long-term planning ability and adaptation ability is important. 3. Enhancement of multimodal learning: Develop methods for more effectively learning the relationships between different modalities. In particular, efficient learning from a small amount of data and understanding of causal relationships between modalities are issues. 4. Automatic expansion and optimization of the knowledge base: Develop automatic knowledge acquisition from execution experience and optimization methods for the knowledge base. In particular, resolving conflicting knowledge and improving the generalization ability of knowledge are important. 5. Enhancement of natural interaction with users: Develop interaction technologies to achieve more natural dialogue and collaborative work. In particular, understanding user intentions and providing appropriate feedback are issues. 6. Decentralized Processing and Improvement of Computational Efficiency: Technological development for efficiently executing computationally intensive processes, including inference of large language models, in a decentralized system. In particular, low-latency processing on edge devices is important. 7. Examination of Ethical and Legal Considerations: Examination of the ethical and legal framework regarding autonomous judgment and actions of robots. In particular, clarification of safety assurance and responsibility sharing is necessary. By addressing these issues, the realization of a robot system with higher intelligence and adaptability is expected. In particular, the development of robots that can collaborate with humans and execute complex tasks safely and efficiently expands the application possibilities in various fields such as industry, healthcare, and households.

Industrial Applicability

[0042] The present invention is expected to be used in the following industrial fields:

[0043] 1. Home Service Robots The home environment is unpredictable and diverse, and high adaptability and intelligence are required for robots to execute various tasks. This system is suitable for applications of home service robots such as the following: 1. Housework Support: - Cooking assistance: Preparation of ingredients, support for the cooking process, setting of tableware, etc. - Cleaning: Floor cleaning, wiping of surfaces, tidying up of items, etc. - Laundry: Sorting of laundry, loading into the washing machine, folding of dried clothes, etc. - Dishwashing: Collection, cleaning, storage of tableware, etc. 2. Support for the Elderly and Disabled: - Support for daily life: Support for meals, changing clothes, taking a bath, etc. - Medicine management: Provision of medicine at appropriate times, management of remaining amounts, etc. - Mobility support: Acquisition of items, opening and closing of doors, assistance with movement, etc. - Safety monitoring: detecting falls, abnormal behavior, and making emergency reports, etc. 3. Child-rearing support: - Toy tidying: collecting and storing scattered toys - Educational support: interactive learning support, operating educational toys, etc. - Safety monitoring: detecting dangerous behavior, removing dangerous objects, etc. The high-level instruction understanding ability and environmental adaptation ability of this system are suitable for coping with the diversity and unpredictability of the home environment. Also, through the integration of visual and force feedback, delicate operations (e.g., handling tableware, folding clothes) are also possible. Furthermore, with the voice dialogue function, natural interaction with the user is realized, improving usability. For example, in the home of an elderly person, it is possible to say, "It's time to take your medicine. Shall I get the medicine for you?" and after obtaining consent, prepare the medicine and water and assist with taking the medicine. Also, it is possible to ask, "What would you like for dinner today?" and help with the preparation of cooking according to the wishes. Such natural dialogue and adaptable support are important in supporting the independent life of the elderly.

[0044] 2. Manufacturing In the manufacturing industry, the demand for multi-variety small-batch production and flexible production lines is increasing, and a robot system with high adaptability is required. This system is suitable for the following applications in the manufacturing industry: 1. Assembly work: - Assembling complex parts: assembling parts with different shapes and arrangements - Precision assembly: aligning and assembling fine parts - Flexible assembly: adapting the assembly process according to product variations 2. Quality inspection: - Visual inspection: inspecting the appearance of products, detecting defects, etc. - Function inspection: checking the operation of products, performance testing, etc. - Dimension measurement: confirming the dimensional accuracy of products, etc. 3. Flexible production line: - Product switching: Quick switching to different products - Process adaptation: Adaptation to changes in the manufacturing process - Collaborative work: Execution of complex tasks through collaboration with humans By efficiently utilizing the knowledge base using the RAG of this system, it is possible to handle various products and manufacturing processes. Also, through the integration of visual and force feedback, precise assembly and inspection become possible. Furthermore, with the voice dialogue function, smooth communication with workers is realized, improving the efficiency of collaborative work. For example, on an assembly line for automotive parts, it is possible to confirm with the worker "The next product is the assembly of Model B. Are you ready?" and prepare the necessary parts and tools. Also, in response to the worker's question "Please teach me how to install this part," it is possible to explain the procedure or perform a demonstration. Such flexible responses and collaborative work are important in multi-variety small-batch production and complex assembly work.

[0045] 3. Medical and Healthcare In the medical and healthcare field, high accuracy and safety are required, and it is necessary to adapt to different situations for each patient. This system is suitable for the following medical and healthcare applications: 1. Surgical assistance: - Instrument delivery: Provide the appropriate instrument to the surgeon at the appropriate timing - Visual field securing: Secure the visual field by adjusting the position of the endoscope or camera - Tissue holding: Stably hold the tissue during surgery 2. Rehabilitation assistance: - Movement support: Provide movement support according to the patient's condition - Progress monitoring: Record and analyze the progress of rehabilitation - Feedback provision: Provide real-time feedback to the patient 3. Patient care: - Meal support: Providing meals according to the patient's condition - Medication management: Providing medications at appropriate times - Mobility support: Assisting with getting up from bed, walking, etc. The high safety and adaptability of this system are suitable for handling different situations for each patient. Also, through the integration of visual and force feedback, delicate operations (e.g., instrument handover, physical interaction with patients) are made possible. Furthermore, with the voice dialogue function, natural communication with patients is realized, improving the quality of care. For example, in a rehabilitation facility, the patient can be asked "How are you feeling today?" and the rehabilitation program can be adjusted according to the condition. Also, while giving encouragement such as "Try a little harder" or "That's the way, wonderful", the patient's movements can be supported. Such individualized care and encouragement are important for enhancing the effectiveness of rehabilitation.

[0046] 4. Retail and Service Industries In the retail and service industries, various tasks such as interacting with customers and handling products are required. This system is suitable for the following applications in the retail and service industries: 1. Product management: - Product display: Appropriate arrangement and organization of products - Inventory check: Stocktaking, checking inventory status, etc. - Product replenishment: Replenishing out-of-stock products, etc. 2. Customer service: - Guidance: Guiding within the store, guiding the location of products, etc. - Information provision: Providing product information, answering questions, etc. - Order reception: Receiving and processing customer orders, etc. 3. Food service: - Cooking assistance: Preparing ingredients, supporting the cooking process, etc. - Plating: Plating dishes, decoration, etc. - Meal preparation: transporting food, arranging it on the table, etc. The natural language understanding ability and creative motion generation ability of this system are suitable for interactions with customers and artistic presentations. Also, through the integration of visual and force feedback, it becomes possible to handle delicate products and plate food. Furthermore, with the voice dialogue function, natural communication with customers is realized, improving the quality of the service. For example, in a café, the system can greet customers with "Welcome. What can I help you with?" and explain the menu or make recommendations according to their preferences. Also, in response to a question like "How is this coffee brewed?", it can explain the brewing method and characteristics of the coffee. Additionally, it can provide drinks with creative decorations such as latte art. Such individualized services and creative expressions are important for enhancing customer satisfaction.

[0047] 5. Education and Research In the field of education and research, flexible responses and creative problem-solving are required. This system is suitable for the following applications in education and research: 1. Experiment support: - Experiment preparation: preparing reagents, setting up equipment, etc. - Data collection: recording experimental data, collecting samples, etc. - Experiment execution: automating routine experimental procedures, etc. 2. Educational robots: - Interactive learning: answering questions, explaining learning content, etc. - Demonstration: demonstrating experiments or operations, explaining procedures, etc. - Evaluation: evaluating learners' understanding, providing feedback, etc. 3. Research and development support: - Prototyping: quickly implementing and validating new ideas - Data analysis: processing and analyzing experimental data - Literature research: searching for and summarizing related studies, etc. By efficiently utilizing the knowledge base using RAG in this system, it is possible to handle various experiments and research themes. Also, the creative motion generation ability is useful for the development of new experimental methods and the creation of educational content. Furthermore, the voice dialogue function enables natural communication with learners and improves the learning effect. For example, in a chemistry laboratory, in response to a student's question such as "Please teach me the procedure of this experiment," it is possible to demonstrate while explaining the procedure. Also, in response to a question such as "Why does this reaction occur?" it is possible to explain the mechanism of the chemical reaction. Furthermore, in response to a request such as "Please analyze the results of this experiment," it is possible to collect data and perform statistical analysis. Such interactive learning support and practical demonstrations are important for deepening the understanding of learners.

[0048] 6. Disaster Response and Hazardous Environments At disaster sites and in hazardous environments, work is required in places where it is difficult for humans to enter. This system is suitable for the following applications in disaster response and hazardous environments: 1. Disaster Relief: - Exploration: Search for victims, confirmation of the situation, etc. - Rescue: Removal of rubble, rescue of victims, etc. - Material Transportation: Transportation of food, water, pharmaceuticals, etc. 2. Hazardous Substance Handling: - Inspection: Detection of hazardous substances, confirmation of the state, etc. - Recovery: Safe recovery and storage of hazardous substances - Treatment: Detoxification, disposal of hazardous substances, etc. 3. Work in Extreme Environments: - Deep Sea Work: Underwater surveys, installation of equipment, etc. - Space Work: Experiments, repairs, etc. on the space station - Nuclear Facilities: Inspections, repairs, etc. in a radiation environment The high adaptability and autonomy of this system are suitable for working in disaster scenes and dangerous environments where it is difficult to predict. Also, through the integration of visual and force feedback, operations in complex environments and precise work become possible. Furthermore, with the voice dialogue function, smooth communication with remote operators is achieved, improving the efficiency and safety of work. For example, at the accident site of a nuclear facility, in response to the instruction from a remote operator such as "Please check the condition of this pipe", the system can visually inspect the pipe and report the presence and degree of damage. Also, in response to the instruction "Please carefully close this valve", the system can operate the valve with appropriate force. Such remote operations and situation reports are important in work in environments where it is dangerous for humans to enter.

[0049] 7. Entertainment and Art In the field of entertainment and art, creativity and expressiveness are required. This system is suitable for the following applications in entertainment and art: 1. Performances: - Dance: Generation and execution of movements synchronized with music - Theater: Emotional expression, dialogue, progression of the story, etc. - Music: Instrument performance, singing, etc. 2. Creative activities: - Drawing: Painting, illustration, graphic design, etc. - Sculpture: Three-dimensional modeling, solid art, etc. - Crafts: Pottery, textiles, woodworking, etc. 3. Interactive art: - Audience participation works: Works that respond to the movements and voices of the audience - Media art: Artistic expressions that utilize digital technologies - Environmental art: Artistic experiences that utilize the entire space The creative motion generation ability and natural language understanding ability of this system are suitable for artistic expression and interaction with the audience. In addition, the integration of visual and haptic feedback enables delicate artistic operations (e.g., painting, sculpture). Furthermore, the voice dialogue function realizes natural communication with the audience and improves the immersion of the performance. For example, in an interactive art exhibition, the audience can be asked "What's your favorite color?" and an improvised painting can be drawn based on the answer. Also, in response to the question "How are you feeling?", music and movements can be generated according to the audience's emotions. Such interaction with the audience and improvised creation are important for enhancing the individuation and immersion of the artistic experience. In these industrial fields, this system enables the execution of complex tasks that were difficult for conventional robot systems, and is expected to contribute to improving productivity, ensuring safety, and creating new services. In particular, the adaptability in unpredictable environments and the ability to understand high-level abstract instructions can generate value in various application scenarios. Also, through the integration and utilization of multimodal feedback, richer environmental understanding and adaptive behavior generation are realized, enabling natural interaction with humans. As a result, the robot can function not just as a tool, but as a collaborative partner of humans.

[0050] 8. Agriculture and Food Industry In the agriculture and food industry, it is necessary to respond to changes in environmental conditions and individual differences in crops and food materials. This system is suitable for the following applications in the agriculture and food industry: 1. Precision agriculture: - Crop management: Monitoring the state of crops, irrigation, fertilization, etc. - Harvesting: Judging ripeness, harvesting with appropriate force, sorting, etc. - Environmental control: Optimizing temperature, humidity, light quantity, etc. 2. Food processing: - Raw material processing: Washing, peeling, cutting, separating, etc. - Cooking: Mixing, heating, cooling, fermentation management, etc. - Packaging: Filling appropriate amounts, sealing, labeling, etc. 3. Quality control: - Visual inspection: Detecting shape, color, scratches, etc. - Physical property measurement: Hardness, viscosity, moisture content, etc. - Component analysis: Detecting nutrients, additives, foreign substances, etc. Through the integration of visual and force feedback in this system, appropriate processing according to the state of crops and food materials becomes possible. Also, by efficiently utilizing the knowledge base using RAG, it is possible to handle a variety of crops and food materials. Furthermore, due to the ability to adapt to environmental changes and uncertainties, stable operations are possible even under outdoor environments and changing conditions. For example, in a tomato farm, in response to an instruction such as "Please check the ripeness of these tomatoes and select those that are ready for harvest," the ripeness can be judged from visual information and harvested with appropriate force. Also, in response to an instruction such as "Please judge whether watering is required in this section," it is possible to make a judgment considering the soil condition and weather conditions. Such precise management and judgment contribute to improving agricultural productivity and quality. In a food processing factory, in response to an instruction such as "Please cut this vegetable into uniform sizes," the shape of the vegetable can be recognized from visual information and the appropriate cutting position and force can be adjusted. Also, in response to an instruction such as "Please check the viscosity of this sauce and adjust it if necessary," the viscosity can be estimated from force feedback and adjusted. Such precise processing and quality control are important for ensuring food consistency and safety.

[0051] 9. Logistics and warehouse management In the field of logistics and warehouse management, handling goods of various shapes and weights and efficient storage and shipping are required. This system is suitable for the following applications in logistics and warehouse management: 1. Item picking: - Item identification: Identifying items by barcode, QR code, appearance - Adaptive grasping: Adjustment of the grasping method according to the shape, weight, and hardness of the product - Packaging: Appropriate arrangement of products, insertion of cushioning materials, sealing of boxes, etc. 2. Inventory management: - Inventory counting: Automatic counting, position confirmation, and status check of products - Product placement: Planning and execution of an efficient storage layout - Inventory replenishment: Detection of out-of-stock situations and generation of replenishment orders 3. Logistics optimization: - Route planning: Planning and execution of an efficient movement route - Loading optimization: Efficient placement of products onto trucks and delivery boxes - Delivery preparation: Sorting and preparation of products according to the delivery order The integration of visual and force feedback in this system enables the proper handling of a variety of products. Also, by efficiently utilizing the knowledge base using RAG, it is possible to handle new products and packaging methods. Furthermore, the ability to adapt to environmental changes and uncertainties allows it to handle congested warehouse environments and fluctuating order patterns. For example, in an e-commerce warehouse, in response to an instruction such as "Please pick and package the products on this order list," each product can be identified, grasped, and packaged in an appropriate manner. Also, in response to an instruction such as "Please check the inventory status of this shelf," the products can be scanned and the inventory quantity and location reported. Such efficient picking and inventory management contribute to reducing logistics costs and improving customer satisfaction.

[0052] 10. Smart City - Public Services In the fields of smart cities and public services, it is necessary to respond to the diverse needs of citizens and maintain a safe and efficient urban environment. This system is suitable for the following applications of smart city - public services: 1. Public space management: - Cleaning: Cleaning of roads, parks, public facilities, etc. - Inspection: Inspection of the status of infrastructure facilities and public facilities - Repair: Minor repair work and maintenance 2. Citizen services: - Guidance: Tourist guidance, facility guidance, road guidance, etc. - Information provision: Event information, traffic information, weather information, etc. - Support: Support for the movement of the elderly and disabled, luggage handling, etc. 3. Safety management: - Monitoring: Detection of abnormal behavior and recognition of dangerous situations - Reporting: Reporting of emergencies and communication with relevant agencies - Initial response: Initial fire extinguishing, evacuation guidance, emergency treatment, etc. The high-level instruction understanding ability and environmental adaptation ability of this system are suitable for responding to various public spaces and citizen needs. In addition, through the integration of visual and force feedback, delicate work (e.g., garbage sorting, facility inspection) is also possible. Furthermore, with the voice dialogue function, natural communication with citizens is realized, and the quality of the service is improved. For example, in public facilities, in response to the request of a visitor such as "Please guide me through this building", the overview of the facility can be explained and the visitor can be guided to the destination. Also, in response to the instruction of a manager such as "Please inspect the status of this facility", the facility can be visually inspected and the presence or absence of abnormalities can be reported. Such interactive guidance and inspection contribute to the efficiency and quality improvement of public services. In parks and squares, in response to an instruction such as "Please clean this area", garbage can be identified and collected in an appropriate manner. Also, in response to an instruction such as "Please check the water pressure of this fountain", the status of the fountain can be observed and any abnormalities can be reported. Such daily management and inspection are important for maintaining the aesthetics and safety of public spaces.

[0053] 11. Space exploration · Extreme environment In space exploration and operations in extreme environments, high autonomy and adaptability are required. This system is suitable for applications in space exploration and extreme environments as follows: 1. Space exploration: - Planetary surface exploration: terrain survey, sample collection, experiment implementation - Space station operations: equipment installation, repair, maintenance - Space construction: structure assembly, construction of resource utilization facilities 2. Deep-sea exploration: - Seafloor survey: terrain mapping, ecosystem observation, resource exploration - Seafloor operations: cable laying, equipment installation, sample collection - Inspection of underwater structures: inspection of oilfield platforms, bridge piers, pipelines 3. Polar exploration: - Ice sheet survey: ice layer structure analysis, core sampling, weather observation - Polar base support: material transportation, equipment maintenance, research support - Environmental monitoring: ecosystem observation, pollutant detection, climate change data collection The high autonomy and adaptability of this system are suitable for operations in environments with communication delays and restricted conditions. Also, through the integration of visual and force feedback, safe operation in unknown environments becomes possible. Furthermore, by efficiently utilizing the knowledge base using RAG, it is possible to respond to unexpected situations. For example, in Mars exploration, in response to an instruction from Earth such as "Please collect a sample of this rock," a sample can be collected in an appropriate method according to the hardness and shape of the rock. Also, in response to an instruction such as "Survey this terrain and find a safe route," the surrounding terrain can be analyzed and the vehicle can move safely while avoiding obstacles. Such autonomous judgment and adaptation are essential for exploration in remote environments with communication delays. On the space station, in response to the astronauts' request of "Please support the repair of this equipment," it is possible to prepare the necessary tools and assist in the work while explaining the repair procedures. Also, in response to the instruction of "Please set up this experimental device," it is possible to appropriately arrange and connect the various parts of the device. Such precise work and support are important in the space environment with limited resources and strict safety requirements.

[0054] 12. Education and Training Simulation In the field of education and training simulation, it is necessary to reproduce real environments and situations to provide an effective learning experience. This system is suitable for the following applications of education and training simulation: 1. Medical Training: - Surgical Simulation: Practice of surgical procedures and techniques - Emergency Response Training: Practice of emergency medical treatment and triage - Patient Care Training: Acquisition of examination, nursing, and rehabilitation techniques 2. Technical Training: - Industrial Technology: Training in machine operation, assembly, repair, and inspection - Construction Technology: Acquisition of heavy machinery operation, surveying, and construction techniques - Scientific Experiments: Practice of chemical, physical, and biological experiments 3. Disaster Response Training: - Evacuation Guidance: Training in evacuation plans and guidance techniques during disasters - Relief Activities: Practice of searching for, rescuing, and providing emergency treatment to disaster victims - Crisis Management: Training in command systems, information collection, and decision-making The integration of the visual and force feedback of this system can provide a real tactile and visual experience. Also, the natural language understanding ability enables interactive guidance with learners. Furthermore, the ability to adapt to environmental changes and uncertainties allows it to handle various scenarios and difficulties. For example, in medical education, in response to an instruction such as "Please examine this patient," it is possible to represent the symptoms of a virtual patient and respond to the students' questions. Also, in response to a request such as "Please demonstrate this surgical procedure," it is possible to demonstrate while explaining each step of the surgery and provide feedback when the students practice. Such interactive simulations and immediate feedback are effective for acquiring medical skills. In disaster response training, in response to a scenario such as "A fire has broken out in this building. Please conduct an evacuation guide," it is possible to reproduce changes in the situation such as the spread of the fire and the filling of smoke, and change the situation according to the judgment and actions of the trainees. Also, in a scenario such as "There are victims under this rubble. Please rescue them," it is possible to train appropriate rescue methods according to the stability of the rubble and the condition of the victims. Such dynamic simulations and situation responses contribute to the improvement of practical disaster response capabilities.

Claims

1. A multimodal feedback integrated knowledge retrieval enhanced robot control system, comprising a language processing component, a vision system, a force sensing module, a speech recognition / synthesis module, a robot control system, a multimodal integration module, and a curated knowledge base, wherein the language processing component dynamically selects and adapts relevant examples from the curated knowledge base using retrieval augmented generation (RAG) technology. A multimodal feedback integrated knowledge retrieval enhanced robot control system characterized by this.

2. The vision system generates a 3D representation of the environment using a depth camera, the force sensing module measures the force received by the robot's end effector using a multi-axis force sensor, the multimodal integration module appropriately weights and integrates information from each modality, and the robot control system incorporates safety constraints such as maximum speed and force limits, and boundaries of the working space. The multimodal feedback integrated knowledge retrieval enhanced robot control system according to Claim 1, characterized by this.

3. A multimodal feedback integrated knowledge retrieval enhanced robot control method, comprising steps of a language processing component processing a user's query and environmental data to decompose complex tasks, selecting relevant examples from a knowledge base using retrieval augmented generation technology, generating executable code, a vision system generating a 3D representation of the environment, a force sensing module measuring the force of the end effector, a speech recognition / synthesis module interacting with the user, a multimodal integration module integrating information from different modalities, and a robot control system integrating feedback to control the robot's operation. A multimodal feedback integrated knowledge retrieval enhanced robot control method including these steps.

Citation Information

Cited By

  • Assessment method and device for robot operation and electronic equipment

    CN118941164A

  • Mechanical arm control method and device based on multi-mode AI intelligent interaction

    CN120620238A

  • Mechanical arm control method and device based on multi-mode AI intelligent interaction

    CN120620238B

  • Retrieval generation method and system based on multi-agent collaboration, terminal and medium

    CN120723880A

  • Operation intention recognition method, system and equipment based on multi-modal fusion and medium

    CN120873982A