Robot material identification system based on multi-modal interaction and large language model
Through multimodal interaction and a robot material recognition system with large language models, natural language, tactile perception and augmented reality vision are integrated, efficient material recognition and human-machine collaboration are achieved, and recognition accuracy and operation transparency are improved in complex environments.
Patent Information
- Application Number
- CN202510635877.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-29
AI Technical Summary
The existing robot systems lack material recognition capabilities in complex environments, lack multimodal interaction and high-level semantic understanding, making it difficult to achieve accurate judgment and human-computer collaboration.
The multimodal input module is used to integrate natural language, tactile perception and augmented reality vision, combine large language models for semantic reasoning and coordination, integrate material recognition, path planning and multimodal coordination strategies, and support system self-learning and real-time visualization.
It improves the robot's material recognition accuracy and human-machine collaboration capabilities in complex environments, enhances operation transparency and safety, and is suitable for various scenarios such as industry, service and medical care.
Smart Images

Figure CN120552044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence augmented reality technology, and in particular to a robot material recognition system based on multimodal interaction and a large language model. Specifically, the system is an intelligent control system that integrates tactile perception, language understanding, AR visualization, human-machine collaboration, and database scheduling. The system is suitable for robot perception, material recognition, decision-making, and control in complex scenarios. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, robots are increasingly being used in complex environments such as manufacturing, healthcare, and services. Traditional robotic systems rely primarily on structured programming languages, fixed task processes, and predefined perception and judgment models, which significantly limits their ability to handle unknown tasks in open environments.
[0003] While some integrated perception and control systems have emerged, these systems often rely on single-modal sensory input (such as vision or mechanics), making it difficult to accurately judge complex materials and multiple scenarios, and lacking material recognition capabilities. Furthermore, most of these systems employ fixed logic-driven strategies and lack high-level semantic understanding and language interaction capabilities, making them difficult to effectively collaborate with non-expert users.
[0004] In recent years, large language models (LLMs) have made breakthrough progress in natural language processing, demonstrating powerful capabilities in understanding semantics, task intent, and knowledge retrieval. However, LLMs are currently primarily used in scenarios such as conversation, question-answering, and text generation, and have yet to be effectively integrated with low-level sensory data (such as force, thermal, and electromagnetic sensor signals) in robotic systems, limiting their application in intelligent control and real-world operations.
[0005] In addition, although augmented reality (AR) technology can provide an intuitive way of interaction, in existing systems AR is mostly used to assist visualization and does not have the ability to display and interactively adjust perception results and control paths in real time. Users often lack an intuitive understanding of the robot's current status and next action, which reduces the transparency and safety of operations.
[0006] Therefore, there is an urgent need for an intelligent control system that can integrate multimodal input (natural language, tactile perception, visual interaction), realize semantic reasoning and coordinated response through a large language model, and support result visualization and operation supervision, so as to achieve true human-machine collaboration, semantic understanding-driven task execution and feedback regulation.
[0007] Current existing technologies for material identification include: Patent application publication number CN119888634A discloses an image processing method and system, and its application in unmanned vehicle inspections of closed converter stations, which uses image processing technology to identify the material of equipment; Patent application publication number CN119880846A discloses a plastic material identification method based on the slopes on both sides of the near-infrared spectrum trough, which determines the material by analyzing the spectral characteristics. However, both of these inventions suffer from the characteristics of single-modal input, relying solely on a single image or spectral data, making it difficult to meet the needs of material identification in complex environments, and lacking high-level language interaction and human-computer collaboration capabilities. Summary of the Invention
[0008] In order to overcome the defects of the above-mentioned prior art, the purpose of the present invention is to provide a robot material recognition system based on multimodal interaction and a large language model, which integrates natural language understanding, tactile perception, and augmented reality vision, and integrates multimodal input, a large language model, robot path planning, augmented reality interaction and multimodal coordination strategy to achieve more flexible, understandable and visual material recognition and robot control, and solve the problems of the prior art such as single interaction mode, fixed logic strategy, lack of adjustment mechanism, and lack of human-computer collaboration.
[0009] To achieve the above object, the technical solution of the present invention is:
[0010] A robot material recognition system based on multimodal interaction and a large language model, comprising:
[0011] The multimodal input module includes a natural language input interface, a tactile sensor module and a user augmented reality interface interaction module; the natural language input interface realizes the transmission of natural language commands, the tactile sensor obtains the physical properties of the object, and the AR augmented reality interface interaction module realizes human-computer operation.
[0012] The natural language parsing module uses a large language model to perform semantic analysis and identify task objectives of the language instructions input by users through the natural language input interface, including location descriptions, object attributes, and operation verb information, and combines perception data to generate contextual understanding.
[0013] The material recognition module includes a preprocessing unit that filters, normalizes and segments the raw data transmitted from the tactile sensor. It uses a neural network classifier (CNN+Transformer) to extract features, record data, output probabilities and determine materials on the preprocessed data, and dynamically adjusts parameters based on the cross-entropy loss function.
[0014] The multimodal coordination strategy unit comprehensively processes the natural language instructions in the natural language parsing module, the material recognition results in the material recognition module, and the augmented reality interaction instructions in the augmented reality (AR) rendering module, and can mediate conflicts between different modalities. At the same time, this module has self-learning capabilities and can dynamically adjust the multimodal fusion rules through the local task database in the database module to store historical task results.
[0015] The database module includes a material property database, a local task database and a connectable network database. The material property database and the connectable network database jointly assist the material identification module in determining the material category; the local task database guides the weight adjustment of the multimodal coordination strategy unit, realizes the real-time storage of previous task results and identification data, and facilitates self-learning of the multimodal coordination strategy unit and the material identification module.
[0016] The augmented reality (AR) rendering module, based on the augmented reality engine, presents material recognition results, grasping paths, safety warnings, etc. to the user augmented reality interface interaction module in the multimodal input module, supporting user visual adjustments and operations.
[0017] The robot control module can realize path planning, motion trajectory control and real-time status feedback based on the decisions made by the multimodal coordination strategy unit.
[0018] The material recognition module uses a neural network (CNN+Transformer) to perform feature extraction, data recording, probability output and material determination on the preprocessed data, and dynamically adjusts the parameters based on the cross entropy loss function, which is specifically expressed as:
[0019]
[0020] in:
[0021] y i is the probability that the object actually belongs to the i-th material category;
[0022] is the probability that the object is predicted to belong to the i-th material category;
[0023] y and is the true probability distribution and the predicted probability distribution, represented by a one-hot encoding vector;
[0024] By dynamically adjusting parameters, the cross entropy loss function is minimized, thereby improving the accuracy of the material recognition module.
[0025] The multimodal coordination strategy unit mediates conflicts between different modalities, which means:
[0026] When the natural language instructions in the natural language parsing module conflict with the tactile recognition results in the material recognition module, the system mediates the conflict based on the preset priority or user interaction instructions. The specific design of the multimodal coordination strategy unit is as follows:
[0027] S1: Receive natural language instructions and extract the first semantics;
[0028] S2: Receive material recognition results;
[0029] S3: Perform multimodal conflict detection and compare the first semantics with the material recognition result. If the two are consistent, proceed to S4; if the two conflict, proceed to S5.
[0030] S4: Generates robotic arm motion control commands based on natural language instructions and material recognition results;
[0031] S5: The user sends intervention instructions in real time through the AR augmented reality interface interaction module, and makes manual adjustments. This instruction is given the highest priority response processing;
[0032] S6: Record the task results to the local task database, including whether the task is successful, user intervention instructions, and material recognition confidence;
[0033] S7: Adjust the weights based on historical task results, including the natural language instruction weight α, the material recognition module weight β, and the augmented reality interaction weight γ;
[0034] If the task is successful and there is no user intervention, increase α and β, and keep γ unchanged; if the task fails due to material recognition errors and frequent user intervention, decrease β, increase γ, and keep α unchanged; if the task fails due to natural language instruction parsing errors and frequent user intervention, decrease α, increase γ, and keep β unchanged;
[0035] S8: Dynamically adjust weights and optimize rules.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] 1. The multimodal input module of the present invention supports multiple input methods such as natural language, tactile perception, and visual interface, which can significantly improve the system's understanding and execution capabilities and realize multimodal human-computer interaction.
[0038] 2. The present invention introduces a large language model to participate in decision-making through the material recognition module and the natural language analysis module, and combines it with the robot control module, which has the advantages of high-level integration of material recognition, semantic understanding and control planning.
[0039] 3. The multimodal coordination strategy unit of the present invention has self-learning capabilities and can dynamically adjust the multimodal fusion rules based on historical task results to improve system accuracy and reliability.
[0040] 4. The augmented reality (AR) rendering module of the present invention makes the decision path transparent and controllable, improves the human operator's trust in the system and the control efficiency, and enhances reality interaction.
[0041] In summary, the present invention integrates natural language understanding, tactile perception, and augmented reality vision, and integrates multimodal input, a large language model, robot path planning, augmented reality interaction, and multimodal coordination strategies. The system architecture has strong versatility and effectively improves intelligent interactivity, execution reliability, and human-computer collaboration in complex environments. It has broad application prospects and can be expanded to various scenarios such as industry, services, and medical care. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION
[0043] The present invention will be described in detail below with reference to the accompanying drawings.
[0044] Reference Figure 1 , a robot material recognition system based on multimodal interaction and large language model, including:
[0045] The multimodal input module includes a natural language input interface, a tactile sensor module and a user augmented reality interface interaction module; the natural language input interface realizes the transmission of natural language commands, the tactile sensor obtains the physical properties of the object, and the AR augmented reality interface interaction module realizes human-computer operation.
[0046] Natural language input interfaces (such as speech recognition or text-command interfaces);
[0047] Tactile sensor module, used to obtain the mechanical, thermal, and electromagnetic properties of the target object;
[0048] AR augmented reality interface interaction module supports gesture input, line of sight locking, graphic annotation and other interaction methods.
[0049] The natural language parsing module uses a large language model to perform semantic analysis and identify task objectives of the language instructions input by users through the natural language input interface, including location descriptions, object attributes, and operation verb information, and combines perception data to generate contextual understanding.
[0050] The material recognition module includes a preprocessing unit that filters, normalizes and segments the raw data transmitted from the tactile sensor. It uses a neural network classifier (CNN+Transformer) to extract features, record data, output probabilities and determine materials on the preprocessed data, and dynamically adjusts parameters based on the cross-entropy loss function.
[0051] The multimodal coordination strategy unit comprehensively processes natural language instructions from the natural language parsing module, material recognition results from the material recognition module, and augmented reality interaction instructions from the augmented reality (AR) rendering module, mediating conflicts between different modalities. Furthermore, this module possesses self-learning capabilities, dynamically adjusting multimodal fusion rules by storing historical task results in the local task database within the database module.
[0052] The database module includes a material property database, a local task database and a connectable network database. The material property database and the connectable network database jointly assist the material identification module in determining the material category; the local task database guides the weight adjustment of the multimodal coordination strategy unit, realizes the real-time storage of previous task results and identification data, and facilitates self-learning of the multimodal coordination strategy unit and the material identification module.
[0053] Material property database: records the physical property data obtained by the tactile sensor module during each task, including the thermal conductivity, electrical constant, friction characteristics, etc. of the material;
[0054] Local task database: records the system coordination plan and final results for each task based on the multi-coordination strategy module;
[0055] Connectable network database: Established databases shared by the network (such as network material property database) to obtain new object information.
[0056] The augmented reality (AR) rendering module, based on the augmented reality engine, presents material recognition results, grasping paths, safety warnings, etc. to the user augmented reality interface interaction module in the multimodal input module, supporting user visual adjustments and operations.
[0057] The augmented reality (AR) rendering module can be used as a user AR interaction within multimodal input to influence system decision-making. Specifically, the AR engine projects the following information onto the user interface: current recognition results and material type; grasping path and destination; robot motion status and warning information. Users can visually adjust paths and operations through the AR interface.
[0058] The robot control module implements path planning, motion trajectory control, and real-time status feedback based on the decisions made by the multimodal coordination strategy unit. Specifically, it generates robot control commands based on the decisions made by the large language model. This includes: path planning based on the control strategy output by the large language model; calling the motion controller to issue the trajectory; and synchronizing the real-time status to the feedback module.
[0059] The material recognition module uses a neural network (CNN+Transformer) to perform feature extraction, data recording, probability output and material determination on the preprocessed data, and dynamically adjusts the parameters based on the cross entropy loss function, which is specifically expressed as:
[0060]
[0061] in:
[0062] y i is the probability that the object actually belongs to the i-th material category;
[0063] is the probability that the object is predicted to belong to the i-th material category;
[0064] y and is the true probability distribution and the predicted probability distribution, represented by a one-hot encoding vector.
[0065] In practical applications, it is usually assumed that an object belongs to only one material category, so only one element in y is 1 and the rest are 0; It is the material probability distribution output by the neural network, and the value of each element is determined by the sensor data after feature extraction.
[0066] The cross-entropy loss function measures the difference between the model's predicted probability distribution and the actual material. The smaller the value, the greater the probability that the material recognition result is consistent with the actual situation. By dynamically adjusting the parameters, the cross-entropy loss function is minimized, thereby improving the accuracy of the material recognition module.
[0067] The multimodal coordination strategy unit mediates conflicts between different modalities, which means:
[0068] When the natural language instructions in the natural language parsing module conflict with the tactile recognition results in the material recognition module, the system mediates the conflict based on a preset priority (e.g., safety priority) or user interaction instructions. The specific design of the multimodal coordination strategy unit is as follows:
[0069] S1: Receive natural language instructions and extract the first semantics;
[0070] S2: Receive material recognition results;
[0071] S3: Perform multimodal conflict detection and compare the first semantics with the material recognition results. If the two are consistent, proceed to S4; if the two conflict, proceed to S5.
[0072] S4: Generates robotic arm motion control commands based on natural language instructions and material recognition results;
[0073] S5: The user sends intervention commands in real time through the AR interface interaction module to make manual adjustments, such as "re-command", "re-identify", "stop identification", "re-plan", etc. This command is given the highest priority response processing;
[0074] S6: Record the task results to the local task database, including whether the task is successful, user intervention instructions, and material recognition confidence;
[0075] S7: Adjust the weights based on historical task results, including the natural language instruction weight α, the material recognition module weight β, and the augmented reality interaction weight γ;
[0076] If the task is successful and there is no user intervention, increase α and β, and keep γ unchanged; if the task fails due to material recognition errors and frequent user intervention, decrease β, increase γ, and keep α unchanged; if the task fails due to natural language instruction parsing errors and frequent user intervention, decrease α, increase γ, and keep β unchanged;
[0077] S8: Dynamically adjust weights and optimize rules to improve overall system performance.
[0078] In summary, the workflow of the present invention is as follows: user inputs the task → the system parses the language instructions → the material perception module identifies the object attributes → the LLM integrates the semantics and perception results → calls the database to assist in judgment → outputs the control command → the robot executes the task → AR feedback status → the user can monitor and adjust in real time.
[0079] Through this system, users can intuitively control the robot to complete intelligent operation tasks on objects of different materials through "natural language + perception collaboration".
[0080] Example: Applying the present invention to intelligent sorting
[0081] In smart production lines and waste sorting, it's necessary to sort objects according to user requirements and material, placing different materials in different areas. The traditional method is manual sorting, which is difficult, time-consuming, and costly. This system can significantly improve sorting efficiency, cost, and accuracy.
[0082] In this embodiment, the robot material recognition system based on multimodal interaction and large language model is deployed on a service robot platform equipped with a UR5 industrial robot arm and a Robotiq gripper, as described in detail as follows:
[0083] The tactile sensor of the present invention is installed inside the gripper at the end of a robotic arm and integrates mechanical, electromagnetic, and thermal sensing capabilities. It is a material recognition sensor that integrates a large language model and a deterministic classifier. Patent application number is 2024114704757.
[0084] The natural language input interface, natural language parsing module, material recognition module, database module, augmented reality (AR) rendering module and multimodal coordination strategy unit of the present invention are integrated and run in a computer's Windows or Linux system;
[0085] Among them, the natural language input interface includes speech recognition and text input interfaces; the natural language parsing module and the multi-module coordination strategy unit are large language models deployed on the HiAgent platform; the material recognition module includes data preprocessing algorithms and material recognition algorithms based on CNN+Transformer neural networks; the database module includes a local database and a network database interface; the augmented reality (AR) rendering module uses Unity as the development platform and runs on Windows or Linux systems.
[0086] The robot control module of the present invention is built into the system and connected to the service robot platform equipped with a UR5 industrial robot arm and a Robotiq gripper.
[0087] The mission scenario is as follows:
[0088] When the user issues the command "Please place the aluminum object on the table in the basket on the left" through natural language, the system's natural language input interface receives the information, calls the natural language parsing module's large language model to perform structured analysis of the semantics, and identifies "aluminum object" as the target object and "place it in the basket on the left" as the operation task.
[0089] Based on the recognition task, the robotic arm drives the gripper to contact the target object. Tactile sensors collect electromagnetic, thermal, and mechanical information in real time and store it in a material property database. The material recognition module then preprocesses the collected data sequence and classifies and identifies the material of the grasped object using a neural network and database module.
[0090] The multimodal coordination strategy unit determines conflicts based on the instructions parsed by the natural language parsing module, the identification results of the material recognition module, and the operation records of "aluminum objects" in the local task database, confirming whether the object can be grasped. During this process, users can send the highest priority instructions at any time through the augmented reality interactive interface.
[0091] After receiving the graspable decision instruction, the robot control module generates a grasping path and a transport path that avoids table corners and obstacle areas, controls the robotic arm to perform tasks according to the planned path, and smoothly grasps and moves the object to the designated area.
[0092] While using the system, users can see the following in real time through the augmented reality (AR) interface: a label above the grasped object indicates its material is "aluminum metal"; the planned path is displayed as a virtual track line, with the user being able to click "Adjust Path" or "Confirm Execution." If the user finds that the path may be too close to an obstacle, they can drag the path node with a gesture, and the system will immediately re-plan the control trajectory.
[0093] After the task is completed, information such as the robot status, path deviation, and material judgment confidence level are recorded in the local task database for subsequent learning and optimization of the system.
Claims
1. A robot material recognition system based on multimodal interaction and large language model, characterized by: include: Multimodal input module, including natural language input interface, tactile sensor module and user augmented reality interface interaction module; the natural language input interface realizes the transmission of natural language instructions, the tactile sensor obtains the physical properties of the object, and the AR augmented reality interface interaction module realizes human-computer operation; The natural language parsing module uses a large language model to perform semantic analysis and identify task objectives of user input language instructions obtained through the natural language input interface, including location descriptions, object attributes, and operation verb information, and combines it with perception data to generate contextual understanding; The material recognition module includes a preprocessing unit that filters, normalizes, and segments the raw data transmitted from the tactile sensor. It uses a neural network classifier (CNN+Transformer) to extract features, record data, output probabilities, and determine materials from the preprocessed data. It also dynamically adjusts parameters based on a cross-entropy loss function. The multimodal coordination strategy unit comprehensively processes natural language instructions in the natural language parsing module, material recognition results in the material recognition module, and augmented reality (AR) interaction instructions in the augmented reality (AR) rendering module, and can mediate conflicts between different modalities. This module also has self-learning capabilities and can dynamically adjust multimodal fusion rules by storing historical task results in the local task database in the database module. The database module includes a material property database, a local task database, and a connectable network database. The material property database and the connectable network database jointly assist the material identification module in determining the material category. The local task database guides the weight adjustment of the multimodal coordination strategy unit, realizes the real-time storage of previous task results and recognition data, and facilitates the self-learning of the multimodal coordination strategy unit and the material identification module. The augmented reality (AR) rendering module uses the AR engine to present material recognition results, grabbing paths, safety warnings, etc. to the user AR interface interaction module in the multimodal input module, supporting user visual adjustments and operations; The robot control module can realize path planning, motion trajectory control and real-time status feedback based on the decisions made by the multimodal coordination strategy unit.
2. The robot material recognition system based on multimodal interaction and large language model according to claim 1 is characterized in that: The material recognition module uses a neural network (CNN+Transformer) to perform feature extraction, data recording, probability output and material determination on the preprocessed data, and dynamically adjusts the parameters based on the cross entropy loss function, which is specifically expressed as: in: y i is the probability that the object actually belongs to the i-th material category; is the probability that the object is predicted to belong to the i-th material category; y and is the true probability distribution and the predicted probability distribution, represented by a one-hot encoding vector; By dynamically adjusting parameters, the cross entropy loss function is minimized, thereby improving the accuracy of the material recognition module.
3. The robot material recognition system based on multimodal interaction and large language model according to claim 1 is characterized in that: The multimodal coordination strategy unit mediates conflicts between different modalities, which means: When the natural language instructions in the natural language parsing module conflict with the tactile recognition results in the material recognition module, the system mediates the conflict based on the preset priority or user interaction instructions. The specific design of the multimodal coordination strategy unit is as follows: S1: Receive natural language instructions and extract the first semantics; S2: Receive material recognition results; S3: Perform multimodal conflict detection and compare the first semantics with the material recognition results. If the two are consistent, proceed to S4; if the two conflict, proceed to S5. S4: Generates robotic arm motion control commands based on natural language instructions and material recognition results; S5: The user sends intervention instructions in real time through the AR augmented reality interface interaction module, and makes manual adjustments. This instruction is given the highest priority response processing; S6: Record the task results to the local task database, including whether the task is successful, user intervention instructions, and material recognition confidence; S7: Adjust the weights based on historical task results, including the natural language instruction weight α, the material recognition module weight β, and the augmented reality interaction weight γ; If the task is successful and there is no user intervention, increase α and β, and keep γ unchanged; if the task fails due to material recognition errors and frequent user intervention, decrease β, increase γ, and keep α unchanged; if the task fails due to natural language instruction parsing errors and frequent user intervention, decrease α, increase γ, and keep β unchanged; S8: Dynamically adjust weights and optimize rules.
Citation Information
Patent Citations
Plastic material identification method based on slopes of two sides of wave trough of near infrared spectrum
CN119880846A
Image processing method and system, and application of image processing method and system in closed converter station unmanned vehicle inspection
CN119888634A