Desktop cleaning method and system
By employing a "brain-brain collaboration" design and a loosely coupled ROS architecture, the problem of accurate identification and multi-brand compatibility for cleaning robots in complex environments has been solved, achieving efficient and low-latency desktop cleaning and improving user experience and cleaning efficiency.
Patent Information
- Application Number
- CN202511161770.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, cleaning robots cannot effectively handle complex dynamic environments in catering services and leisure and entertainment scenarios. They have difficulty accurately identifying heterogeneous targets such as garbage and tableware. Furthermore, the hardware control module and perception algorithm are tightly coupled, making it difficult to be compatible with multiple brands of robotic arms or sensors, resulting in response delays and high upgrade and maintenance costs.
Adopting a "brain-brain collaboration" design, it achieves efficient collaboration between voice interaction, visual perception, and motion planning through a combination of speech recognition module, LLM model, VLM model, and VLA model. It utilizes ROS to build a loosely coupled distributed architecture, supports compatibility with multiple brands of robotic arms and sensors, and reduces maintenance costs.
It achieves efficient cleaning in complex environments, rapid response, and low-latency desktop cleaning, improving cleaning efficiency and user experience in service scenarios, and reducing system expansion and maintenance costs.
Smart Images

Figure CN121015072A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of robots, in particular to a table cleaning method and system. BACKGROUND
[0002] In the scenarios of catering services, leisure and entertainment, the efficient cleaning of bars and tables directly affects the user experience and operational efficiency. The existing technology mainly has the following defects: Traditional cleaning relies on real-time response of service personnel, and response delay is easy to occur during peak hours, resulting in prolonged waiting time for customers and reduced turnover rate. Existing cleaning robots are mostly based on preset rules, which cannot effectively handle complex dynamic environments and accurately identify heterogeneous targets such as garbage, tableware, and recyclable materials. Some systems that support voice control rely on simple instruction templates and lack deep understanding of natural language intent, resulting in high error rate. The planning of mechanical arm movements and the grasping of objects often fail due to environmental interference, especially for small garbage or fragile items. The hardware control module and perception algorithm are tightly coupled, making it difficult to be compatible with multiple brands of mechanical arms or sensors, and the upgrade and maintenance cost is high. SUMMARY
[0003] The main purpose of the present application is to solve the technical problems of low interaction intelligence and limited execution accuracy in the prior art. The present application proposes a table cleaning method, comprising the following steps: In the instruction acquisition step, the user's voice instruction is acquired, and the voice recognition module is used to convert it into raw text; In the semantic analysis step, the LLM model is started, the raw text is analyzed semantically, converted into a standardized structured task instruction, and a reply to the customer is generated; In the scene perception and task planning step, the VLM model is started, the standardized structured task instruction is converted into a task sequence including multiple sub-tasks, and the multiple sub-tasks are sequentially transmitted to the VLA model, In the action generation step, the scene perception and task planning step, the VLA model converts the sub-tasks into continuous action instructions executable by the robot in real time; In the robot execution and control step, the generated action instructions are issued to the mechanical arm controller to drive the mechanical arm to complete the table cleaning action.
[0004] The raw text is converted by the voice recognition module, which includes: The voice instruction of the user is captured by an audio acquisition device, the raw audio signal is preprocessed, and a pre-trained automatic speech recognition model is used to convert the processed voice signal into a text string, which is used as the initial load of the user's intent.
[0005] The starting LLM model performs semantic analysis on the original text, converts it into a standardized structured task instruction, and generates a reply sentence for the customer, including: A pre-configured LLM model is called to understand the intention of the obtained text string; the LLM model outputs a standardized structured task instruction, which is a pre-defined JSON object, and the JSON object includes task type and records user original instruction; the LLM model also generates a personified and concise reply sentence.
[0006] The starting VLM model converts the standardized structured task instruction into a task sequence including multiple sub-tasks, including: The standardized structured task instruction generated by the semantic analysis step is received; real-time multi-view RGB images are obtained from multiple cameras deployed at different positions; the VLM model understands the association between the instruction and the visual scene, identifies all the objects on the desktop and their spatial positions, and generates a task sequence including multiple sub-tasks according to the task type.
[0007] The multiple sub-tasks are sequentially transmitted to the VLA model, and the VLA model converts the sub-tasks into real-time robot executable and continuous action instructions, including: The VLA model obtains a single sub-task in the task sequence; the VLA model obtains the real-time state of the robot body and visual feedback; the VLA model outputs a series of control instructions, including target joint position, speed or end effector pose and gripper opening and closing output at a specific frequency.
[0008] The method further includes a data acquisition and processing step: Hundreds of complete robot operation trajectories are collected through demonstration to form an original data set, which is composed of high-fidelity expert demonstration trajectories and error-correcting countermeasures trajectories, wherein the high-fidelity expert demonstration trajectories cover a variety of operation scenes and target objects, and are used for model learning of core skills; the error-correcting countermeasures trajectories record the process of recovering from non-ideal or error state to correct execution path of the robot, and are used to enhance the fault tolerance of the model in actual deployment; The collected continuous trajectories are segmented, and a standardized natural language instruction is assigned to each meaningful action segment, and the standardized natural language instruction-trajectory data pair constitutes the basis for model training; the standardized natural language instruction is both the output target generated by the VLM model after receiving the upper task type and the input received by the VLA model, and the output target of the VLA model corresponds to the real and collected mechanical arm joint state sequence under the standardized natural language instruction.
[0009] The second aspect of the present application provides a desktop cleaning system, comprising: a voice recognition module, which acquires a voice instruction issued by a user and converts the voice instruction into raw text through the voice recognition module; an LLM model, which performs semantic analysis on the raw text, converts the raw text into a standardized structured task instruction, and generates a reply to the customer; a VLM model, which converts the standardized structured task instruction into a task sequence comprising a plurality of subtasks; a VLA model, which sequentially passes the plurality of subtasks to the VLA model, and the VLA model converts the subtasks into robot-executable and continuous action instructions in real time; an execution unit, which issues the generated action instructions to a robot arm controller to drive the robot arm to complete a tabletop cleaning action.
[0010] The key data channel of the tabletop cleaning system adopts a TCPROS / UDPROS protocol to guarantee low-delay transmission, supports high-throughput data interaction of Majic-LLM instruction analysis, VLM / VLA model reasoning, and ROT-Control executor modules, and simultaneously realizes multi-task parallel communication through namespace isolation.
[0011] The third aspect of the present application provides an electronic device, comprising a memory and at least one processor, the memory having instructions stored therein, and the memory and the at least one processor being interconnected by a circuit; the at least one processor invokes the instructions in the memory to enable the electronic device to perform the tabletop cleaning method as described above.
[0012] The fourth aspect of the present application provides a computer-readable storage medium having instructions stored therein, which, when executed on a computer, enable the computer to perform the tabletop cleaning method as described above.
[0013] The present application has the following beneficial effects: The present system adopts a "big brain and small brain cooperation" design concept, the system "brain" relies on a visual-language model (VLM), identifies the category of desktop items and environmental information according to real-time visual data, and generates an executable task sequence and logic planning in combination with user instruction information processed by the LLM; the cerebellum is based on a visual-language-action model (VLA) and is responsible for fine action control, which decomposes the task into robot arm motion trajectory, gripper opening and closing degree, obstacle avoidance path and other robot arm bottom instructions, to ensure that the garbage is accurately put into the garbage can and the tableware is stably placed on the tray.
[0014] This invention utilizes a loosely coupled distributed architecture built upon ROS, employing a multi-node communication mechanism to achieve efficient collaboration among modules such as voice interaction, visual analysis, and motion control. It supports dynamic allocation of hardware resources and parallel task processing, enabling millisecond-level low-latency interaction between task planning and action execution, ensuring system real-time performance. Through standardized interface encapsulation, it is compatible with multiple brands of robotic arms and sensor devices, reducing system expansion and maintenance costs. Attached Figure Description
[0015] Figure 1 A photograph of the robotic arm portion provided in an embodiment of the present invention; Figure 2 This is a diagram of the instruction control interface of the present invention; Figure 3 Photo of a robotic arm performing desktop item sorting; Figure 4 A robotic arm is used to wipe photos of stains off a desktop. Detailed Implementation
[0016] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0017] This invention relates to an intelligent desktop cleaning solution based on ROS (Robot Operating System). The system receives user commands through a voice interaction module and understands user needs through a large language model. It adopts an innovative "brain-brain collaboration" hierarchical architecture and multimodal perception fusion technology to achieve efficient collaboration between the perception, decision-making, and execution modules, thereby realizing fully automated desktop cleaning. In addition, the system can dynamically adapt to complex environmental interference (such as stacked items and temporary obstacles), and has the advantages of fast response speed and high generalization ability, significantly improving cleaning efficiency and user experience in service scenarios.
[0018] This system is suitable for intelligent cleaning needs in scenarios such as catering services and leisure and entertainment, and is especially well-suited for the following environments: 1. Commercial catering establishments: such as bars, cafes, fast food restaurants, etc., quickly clear table residue (bottles, paper scraps, tableware) during peak hours to reduce customer waiting time and improve table turnover efficiency.
[0019] 2. Unattended spaces: nighttime bars or self-service areas, automated cleaning can be triggered by voice commands, reducing the cost of manual inspections.
[0020] 3. High-end home bar counter: Meets the needs of smart home scenarios, realizes the classification and storage of tableware and the immediate disposal of garbage, and avoids the growth of bacteria from food residue.
[0021] For ease of understanding, the specific process of the embodiments of the present invention is described below. The first embodiment of the desktop cleaning method of the present invention includes: The system acquires user-issued voice commands and converts them into raw text using a speech recognition module. It then activates an LLM model to perform semantic parsing on the raw text, converting it into standardized structured task commands and generating responses for the customer. Next, it activates a VLM model to transform the standardized structured task commands into a task sequence comprising multiple sub-tasks. These sub-tasks are then passed one by one to a VLA model, which converts them into continuous, executable motion commands in real time. Finally, the generated motion commands are sent to the robotic arm controller to drive the robotic arm to complete the desktop cleaning action.
[0022] In this embodiment of the invention, the core technical terms involved are defined as follows: User instructions are the intentions given by users through natural language (such as voice or text) to instruct the robotic system to perform specific tasks. Examples include "The desktop is a bit messy," "Clean the table," and "Take that bottle away." User instructions are the trigger points for the entire automation process.
[0023] Scene perception: refers to the process by which a robot system uses sensors (such as cameras) to acquire environmental information and analyze and understand that information. In this invention, it mainly refers to identifying the categories (such as trash, tableware, and recyclables), positions, postures, and environmental conditions such as stains on a table through visual information.
[0024] Task planning refers to the process by which a system, based on user instructions and scene awareness, breaks down a macroscopic, fuzzy task into a series of specific, ordered, and executable sub-task sequences. For example, "cleaning the table" can be planned as "1. grab the bottle; 2. put the bottle in the recycling bin; 3. grab the paper ball; 4. put the paper ball in the trash can".
[0025] Motion generation: This refers to the process of converting a subtask generated in the task planning step into a series of low-level control commands that a robot (such as a robotic arm) can execute. These commands include, but are not limited to, the target angles of each joint of the robotic arm, the motion trajectory of the end effector, and the opening and closing states of the gripper.
[0026] As one embodiment of the present invention, a desktop cleaning method includes the following steps: In the instruction acquisition step, the voice instructions given by the user are acquired and converted into raw text by the speech recognition module; The semantic parsing step involves starting the LLM model, performing semantic parsing on the original text, converting it into standardized structured task instructions, and generating responses for the customer. The scene perception and task planning steps involve starting the VLM model and converting the standardized structured task instructions into a task sequence comprising multiple subtasks; then, each of the subtasks is passed to the VLA model. The VLA model, which includes action generation, scene perception, and task planning steps, converts subtasks into continuous action instructions that the robot can execute in real time. The robot executes and controls the process by sending the generated motion commands to the robotic arm controller, which then drives the robotic arm to complete the desktop cleaning action.
[0027] Command Acquisition Steps: The system captures the user's voice commands through audio acquisition devices such as microphones. To eliminate environmental noise interference and improve recognition accuracy, the original audio signal is first preprocessed, such as through noise reduction and filtering. Then, a pre-trained Automatic Speech Recognition (ASR) model is used to convert the processed voice signal into a text string. This text string serves as the initial loading of the user's intent and is passed to subsequent modules for processing.
[0028] In the semantic parsing step: the system invokes a pre-configured Large Language Model (LLM) to perform intent understanding and standardization on the raw text instructions obtained in the previous step. This LLM is assigned a specific system role, such as "MBot, the bar cleaner," and is tasked with parsing the user's natural language instructions into a predefined JSON format. This JSON object clarifies the task type, records the user's original instruction, and generates a human-like, concise response. For example, for the user instruction "The table is a bit messy, please clean it up," the LLM will output a JSON object containing the fields "task," "user," and "utterance," clearly indicating that this is a "Desk Cleaning" task. This structured output provides a clear and unambiguous task entry point for subsequent modules.
[0029] In the scene perception and task planning steps, the "Brain" (VLM, Visual-Language Model) module is responsible. This module receives structured task instructions (task field) generated in the semantic parsing step and simultaneously acquires real-time multi-view RGB images from two cameras deployed on the robot's wrist and at an external fixed position. The core function of the Brain (VLM) model is to understand the association between instructions and the visual scene. It can identify all items on the table (such as beverage bottles, paper scraps, and plates) and their spatial positions, and generate a logically clear macro-level task sequence based on the task type (such as "Table Cleaning"). For example, for the "Table Cleaning" task, the model might generate the following sequence: [pick up the plate, move the arm to the left, put the plate into the tray, pick up the tissue]. This sequence defines the type of operation, the category of the target object, and the target placement position, providing clear guidance for subsequent action execution.
[0030] In the motion generation step, the cerebellum (VLA, Visual-Language-Motion Model) module is responsible. This module operates in a "cerebellum" manner, focusing on translating macroscopic task instructions from the "brain" into fine, smooth physical movements. It receives individual subtasks from the task sequence (such as picking up the plate) and real-time robot body states (such as current joint angles) and visual feedback. The cerebellum (VLA) model outputs a series of low-level control instructions, such as target joint positions, velocities, or end effector postures (pose and rotation) output at specific frequencies (such as 20Hz), as well as gripper opening and closing commands.
[0031] In the robot's execution and control steps, the entire system is coordinated based on the distributed communication architecture of the Robot Operating System (ROS). The Virtual Machine Model (VLM) and the Virtual Machine Lagrange Model (VLA) operate as independent ROS nodes. Motion commands generated by the VLA are published through a standardized ROS topic, such as ` / arm_cmd`. The robot's underlying controller node subscribes to this topic, parses the commands in real time, and drives the hardware execution. Simultaneously, sensor data from the robot (such as joint encoder readings) is published through another topic (such as ` / joint_states`), forming feedback for the VLA model to perform closed-loop control, thereby achieving real-time correction of execution errors.
[0032] As a preferred embodiment, the specific steps of the method of the present invention include: S1. Instruction Acquisition and Semantic Resolution: S11. Capture the user's voice signal A_raw through the audio interface. Then, use the speech recognition model F_asr to convert the voice signal into a text instruction T_cmd, i.e., T_cmd = F_asr(A_raw).
[0033] S12. Call the Large Language Model (LLM) M_llm to perform semantic parsing on the raw text instruction T_raw. By setting the system prompt word (Prompt), guide the model to output the structured JSON task instruction J_task. This process is represented as J_task = M_llm(Prompt, T_raw).
[0034] For example, generating: { "task": "Desk Cleaning", "user": "My desk is a bit messy, can you clean it up for me?", "utterance": "Of course, I'll clean your desktop for you."} S13. The system parses J_task, extracting the task type (e.g., "Desk Cleaning") and the utterance. The utterance can be read to the user via the text-to-speech (TTS) module.
[0035] S2. Scene perception and task planning based on the Virtual Brain Model (VLM): S21. The system synchronously obtains the task type (task_type), wrist camera image (I_wrist), and global camera image (I_global) parsed in step S1.
[0036] S22. The multimodal input (task_type, I_wrist, I_global) is fed into the Virtual Brain Model (VLM) M_vlm. The model performs inference and outputs a structured task sequence S_task, in the form S_task = M_vlm(task_type, I_wrist, I_global). This sequence contains N subtasks s_1, s_2, ..., s_n. In one specific embodiment of the present invention, the brain (VLM) model M_vlm is an end-to-end model based on a multimodal Transformer, employing an encoder-decoder architecture. Internally, it includes a ViT (Vision Transformer) encoder for extracting visual features, converting the input image into a series of feature vectors containing rich contextual information; a language encoder for processing text instructions; and a cross-modal attention fusion module for aligning and fusing visual and linguistic information. Finally, a decoder concatenates the image feature vectors output by the visual encoder with the text task instructions, serving as the input context. Subsequently, a structured task sequence S_task is generated using an autoregressive approach. During generation, the decoder utilizes an attention mechanism to associate image features with text instructions, ensuring that the generated task plan closely corresponds to the real visual scene.
[0037] S3. Generation of action commands based on the cerebellar (VLA) model: S31. The system retrieves the currently pending subtask s_i from the task sequence S_task in sequence.
[0038] S32. Input the subtask s_i, the current real-time camera image (I_wrist, I_global), and the robot's current state Q_robot (including joint angles, end-effector pose, etc.) into the cerebellum (VLA) model M_vla.
[0039] S33. The cerebellar (VLA) model M_vla outputs a series of low-level action instructions A_t over time steps, i.e., A_t = M_vla(s_i, I_wrist, I_global, Q_robot). Here, A_t can be the target joint angle and gripper state at the next time step.
[0040] In one specific embodiment of the present invention, the cerebellar (VLA) model M_vla is an end-to-end conditional generative policy network whose structure is designed to decouple high-level semantic understanding from low-level motor control. The model consists of two core modules: a multimodal context encoder and a dedicated action generation module.
[0041] The multimodal context encoder is a large-scale vision-language model (VLM). Its function is to receive subtask instructions s_i (e.g., grab a bottle) and a real-time visual image stream (I_wrist, I_global) from a camera. Leveraging its strong prior knowledge, the encoder processes this multimodal input information into a compact and context-rich feature vector. This feature vector encapsulates a comprehensive understanding of the task intent and the dynamic environment.
[0042] The motion generation module receives feature vectors from the context encoder and fuses them with the robot's real-time ontological perception state Q_robot (such as the angles of each joint). The core function of this expert module is to learn a conditional probability flow that deterministically maps random noise sampled from a simple prior distribution (such as a Gaussian distribution) to an expert-level, temporally coherent sequence of future actions A_t. During inference, by integrating this learned vector field, the model can generate smooth and accurate robot motion trajectories.
[0043] S4. Robot Execution and Closed-Loop Control: S41. Encapsulate the action command A_t output by the cerebellum (VLA) model into a ROS message and publish it through a specified topic (such as / arm_cmd).
[0044] S42. The ROS driver for the robotic arm subscribes to this topic, parses messages, and controls the robotic arm hardware to perform actions.
[0045] S43. The system continuously receives the latest state of the robot body and the environment through topics (such as / joint_states, / camera / image_raw), and feeds it back to the cerebellum (VLA) model as input for the next time step, forming a high-frequency real-time closed-loop control loop to ensure the robustness and adaptability of the action.
[0046] The desktop cleaning method in the embodiments of the present invention has been described above. The desktop cleaning system in the embodiments of the present invention is described below: The speech recognition module acquires the user's voice commands and converts them into raw text. The LLM model performs semantic parsing on the original text, converts it into standardized structured task instructions, and generates responses for customers. The VLM model transforms the standardized structured task instructions into a task sequence that includes multiple subtasks; The VLA model passes multiple subtasks to the VLA model one by one, and the VLA model converts the subtasks into continuous action instructions that the robot can execute in real time. The execution unit sends the generated motion commands to the robotic arm controller, driving the robotic arm to complete the desktop cleaning action.
[0047] This system primarily employs a master-slave robotic arm. During data acquisition, the master arm teaches and controls the slave arm to collect data. During system deployment and operation, only the slave arm is used. It also requires two cameras (capable of capturing RGB and depth images): one mounted on the wrist of the slave arm to provide local spatial information, and the other mounted on a desktop to provide a third-person, global perspective. The robotic arm equipped with the cameras is shown below. Figure 1 As shown.
[0048] To ensure the high generalization and robustness of the Virtual Brain Model (VLM) and Virtual Cerebellum Model (VLA), this invention employs a systematic data acquisition and processing method. First, hundreds of complete robot operation trajectories were collected through teaching, forming the original dataset. This dataset was carefully designed, containing two categories: high-fidelity expert demonstration trajectories and error-correcting adversarial trajectories. The expert demonstration trajectories cover diverse operation scenarios and targets, used for the model to learn core skills; the error-correcting adversarial trajectories record the robot's recovery from a non-ideal or erroneous state to the correct execution path, significantly enhancing the model's fault tolerance in actual deployment.
[0049] During the data annotation phase, the collected continuous trajectory is segmented, and a standardized natural language instruction is assigned to each meaningful action segment (such as putting a tissue into the recycling bin), for example, "put the tissue into the recycling can." These "standardized natural language instruction-trajectory" data pairs form the basis of model training. In the training process, this standardized natural language instruction is both the output target that the Virtual Brain Model (VLM) needs to generate after receiving the upper-level task type, and the input received by the Virtual Laparoscopic Model (VLA). The output target of the VLA corresponds to the actual, collected sequence of robotic arm joint states under this standardized natural language instruction.
[0050] Finally, to ensure data quality, integrity, and compatibility with the training framework, all labeled data was converted and serialized into a standardized dataset format (such as the Lerobot specification), providing a guarantee for efficient and reliable model training. Deployment process: S1: Start the ROS core.
[0051] S2: Start the hardware driver (robotic arm, camera).
[0052] S3: Start each AI model service (ASR, LLM, TTS, VLM, VLA) separately. S4: Start the main control GUI interface like Figure 2GUI Interface: Provides a powerful debugging and control interface, allowing real-time viewing of camera feeds and commands output by each model, as well as manual input for step-by-step debugging. Details are as follows: Back button: Clicking it once disables the VLA model's action output and returns to the initial position. Clicking it a second time enables VLA output. The current command area displays the command status of each model in the inference state of the robotic arm in real time. For easier debugging, a manual command setting area has been added below, allowing you to set commands for individual models. If you want to test VLA separately, you cannot enter LLM or VLM commands, because LLM commands will affect VLM commands, and VLM commands will affect VLA commands. Similarly, if you want to test VLM, do not set LLM commands.
[0053] In addition, the interface includes several buttons for setting commands and controlling the start and stop of the robotic arm. A demonstration is shown below. Figures 1-4 As shown.
[0054] After launching the master node in the VS Code terminal, starting the device, and deploying the model, the robotic arm can begin cleaning the desktop, sorting objects on the desktop: trash into the trash can, dishes into the storage box, and recyclable bottles and cans into the recycling box. When there are stains on the desktop, the robotic arm will pick up a cloth and begin wiping them off.
[0055] This system employs modular, layered collaboration to form a closed-loop intelligent cleanup task coordination process encompassing "instruction-planning-execution-feedback". Relying on the ROS multi-node communication mechanism, the system ensures efficient allocation of computing resources and loosely coupled collaboration between modules, balancing real-time performance and scalability.
[0056] This system builds its communication network based on the ROS distributed architecture, employing a loosely coupled publish / subscribe (Pub / Sub) model among multiple nodes to achieve module interaction. Core communication is accomplished through standardized topics and services, such as the vision module publishing image stream topics ( / camera_data) and control nodes subscribing to robotic arm trajectory commands ( / arm_cmd), combined with ActionLib to achieve long-term task status feedback. Critical data channels use the TCORS / UDPROS protocol to ensure low-latency transmission, supporting high-throughput data interaction between modules such as Majic-LLM instruction parsing, VLM / VLA model inference, and ROT-Control actuators. Simultaneously, namespace isolation enables parallel communication for multiple tasks, ensuring the system's dynamic scalability.
[0057] This invention also provides an electronic device, which can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and memory, and one or more storage media (e.g., one or more mass storage devices) for storing applications or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations stored in the storage media on the electronic device.
[0058] The electronic device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that the electronic device structure in this embodiment does not constitute a limitation on the electronic device itself, and may include more or fewer components, or combinations of certain components, or different component arrangements.
[0059] This invention provides an electronic device structure that can vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and memory, and one or more storage media (e.g., one or more mass storage devices) for storing applications or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations stored in the storage media on the electronic device.
[0060] The electronic device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that the structure of the electronic device does not constitute a limitation on the electronic device itself, and may include more or fewer components than described above, or combine certain components, or have different component arrangements.
[0061] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the aforementioned method.
[0062] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0063] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0064] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A desktop cleaning method, characterized in that, Includes the following steps: The instruction acquisition step involves acquiring the user's voice instructions and converting them into raw text using a speech recognition module. The semantic parsing step involves starting the LLM model, performing semantic parsing on the original text, converting it into standardized structured task instructions, and generating responses for the customer. The scene perception and task planning steps involve starting the VLM model and converting the standardized structured task instructions into a task sequence comprising multiple subtasks; then, each of the subtasks is passed to the VLA model. The VLA model, which includes action generation, scene perception, and task planning steps, converts subtasks into continuous action instructions that the robot can execute in real time. The robot executes and controls the process by sending the generated motion commands to the robotic arm controller, which then drives the robotic arm to complete the desktop cleaning action.
2. The desktop cleaning method according to claim 1, characterized in that, The instruction acquisition step includes: Capture user's voice commands using audio capture devices; Preprocess the raw audio signal; Using a pre-trained automatic speech recognition model, the processed speech signal is converted into a text string, which serves as the initial loading of the user's intent.
3. The desktop cleaning method according to claim 1, characterized in that, The semantic parsing step includes: Call a pre-configured LLM model to perform intent understanding on the acquired text string; The LLM model outputs standardized structured task instructions, which are predefined JSON objects. The JSON objects include the task type and record the user's original instructions. The LLM model also generates a human-like, concise response.
4. The desktop cleaning method according to claim 1, characterized in that, The scene perception and task planning steps include: Receive standardized structured task instructions generated by the semantic parsing step; Acquire real-time multi-view RGB images from multiple cameras deployed in different locations; The VLM model understands the association between instructions and the visual scene, identifies all items on the desktop and their spatial locations, and generates a task sequence that includes multiple sub-tasks based on the task type.
5. A desktop cleaning method according to claim 1, characterized in that, The action generation steps include: The VLA model retrieves individual subtasks from a task sequence; VLA model acquires the robot's real-time status and visual feedback; The VLA model outputs a series of control commands, including target joint position, velocity, or end effector attitude, as well as gripper opening and closing, output at a specific frequency.
6. A desktop cleaning method according to claim 1, characterized in that, The method also includes data acquisition and processing steps: Hundreds of complete robot operation trajectories were collected through teaching to form the original dataset. The original dataset consists of high-fidelity expert demonstration trajectories and error-correcting adversarial trajectories. The high-fidelity expert demonstration trajectories cover a variety of operation scenarios and targets, which are used for the model to learn core skills. The error-correcting adversarial trajectories record the process of the robot recovering from a non-ideal or erroneous state to the correct execution path, which is used to enhance the model's fault tolerance in actual deployment. The collected continuous trajectory is segmented, and a standardized natural language instruction is assigned to each meaningful action segment. The standardized natural language instruction-trajectory data pair constitutes the basis for model training. The standardized natural language instruction is both the output target that the VLM model needs to generate after receiving the upper-level task type and the input received by the VLA model. The output target of the VLA model corresponds to the real, collected sequence of robotic arm joint states under the standardized natural language instruction.
7. A desktop cleaning system, characterized in that, The system includes: The speech recognition module acquires the user's voice commands and converts them into raw text. The LLM model performs semantic parsing on the original text, converts it into standardized structured task instructions, and generates responses for customers. The VLM model transforms the standardized structured task instructions into a task sequence that includes multiple subtasks; The VLA model passes multiple subtasks to the VLA model one by one, and the VLA model converts the subtasks into continuous action instructions that the robot can execute in real time. The execution unit sends the generated motion commands to the robotic arm controller, driving the robotic arm to complete the desktop cleaning action.
8. An electronic device comprising a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the electronic device to perform the steps of the desktop cleanup method as claimed in any one of claims 1-6.
9. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the steps of the desktop cleanup method as claimed in any one of claims 1-6.
Citation Information
Cited By
A working system based on VLA and VLM dual-head cooperation
CN122416150A