A robotic arm control method, system and device for implementing multi-modal general operation tasks
By combining a robotic arm control method with multiple input modal information such as voice, vision and text, the problem of insufficient single modal input in the prior art is solved, and efficient task execution and multimodal interaction capabilities are achieved in complex environments.
Patent Information
- Application Number
- CN202510281655.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing robotic arm control method is based on a single modal input, which is difficult to adapt to complex and changeable manipulation needs, and lacks dynamic environment adaptability and multimodal interaction capabilities, resulting in limited success rate and operational accuracy of task execution.
By combining a variety of input modal information such as voice, vision and text, we provide robotic arm control methods for multimodal general operating tasks, including acquisition of multimodal data, modal alignment and encoding, instruction understanding and motion control, and adopt a strategy big model for end-to-end control model training.
It has realized that in a non-standardized and dynamically changing operating environment, the robot can effectively understand environmental information and human needs, improve the probability of success and operation accuracy of task completion, and has the ability to dynamically adjust task goals and independently correct errors.
Smart Images

Figure CN119772905B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotic arm control, and in particular to a robotic arm control method, system and device for realizing multi-modal general operation tasks. Background Art
[0002] Traditional robotic arm control systems are mostly applied in industrial production lines, such as industries like automobile assembly, food processing, electronic component assembly, etc. The tasks performed have characteristics such as a single purpose, clear processes, and a stable environment. With the progress of sensor technology and the advancement of artificial intelligence algorithms, robotic arm control has gradually gained the ability to perceive and understand the environment and human dynamic instructions, thus showing a certain processing ability in complex tasks and unstructured environments, and its application scope has also expanded to more diverse fields such as medical assistants, catering services, and home assistants.
[0003] Most existing robotic arm control methods are based on single-modal input, such as speech recognition, image recognition, or text recognition, etc. Among them, speech recognition is usually parsed by converting speech instructions into text, and finally jointly constitutes the control instructions of the robot with additional text instructions. However, in practical applications, single-modal input has many limitations and cannot adapt to complex and changeable manipulation requirements. For example, text recognition systems are difficult to comprehensively describe task information, image recognition systems are sensitive to changes in lighting and viewing angles, and speech recognition systems are easily interfered in noisy environments. In addition, in complex environments, traditional robotic arm control methods often have difficulty achieving efficient interactive operations. For example, as a home assistant robot with highly customized characteristics, a robotic arm based on single-modal or only on text and visual modalities is difficult to naturally accept user feedback guidance during the execution of user tasks, nor can it obtain personalized information about the user based on the user's speech characteristics to achieve better task planning, such as providing more user-friendly services according to the user's emotional and habitual characteristics. In existing multi-modal robotic arm control schemes, a multi-level multi-modal signal processing scheme is usually adopted, that is, modal information such as speech will be converted into unified text modal information for subsequent processing. However, affected by environmental interference and model capabilities, this hierarchical conversion scheme is very likely to introduce cumulative errors and is difficult to capture more abundant emotional intentions, voiceprint characteristics, etc. in the original speech signal, thus limiting the upper limit of model capabilities.
[0004] In summary, the following limitations still exist in existing robotic arm control methods:
[0005] (1) Most current robotic arm control strategies are based on single-modal input (such as speech recognition, image recognition, or text recognition, etc.). In non-standardized and dynamically changing operating environments, robots usually have difficulty fully understanding environmental information and complex human needs, and are easily misled by specific modal information, affecting the success rate and operation accuracy of task execution;
[0006] (2) Traditional robotic arm interaction methods are limited to simple text, picture instruction control, or pre-set program selection, lacking the ability to understand complex instructions with rich modalities, semantics, and context associations, and also lacking the ability to interact and adjust in real time by receiving rich modal information in the environment during task execution;
[0007] (3) Traditional multi-modal robotic arm control schemes lack the end-to-end processing ability for modal signals such as speech, relying on the layering of multi-modal signal conversion, processing, understanding, and planning, which may lead to the cumulative error layer by layer and the loss of important information;
[0008] (4) Traditional robotic arm control methods are still not flexible, efficient, and accurate enough in application scenarios such as home assistants and catering services that require dynamic adaptation ability, self-correction ability, and multi-modal interaction ability. Summary of the Invention
[0009] The purpose of the present invention is to provide a robotic arm control method, system, and device for realizing multi-modal general operation tasks in view of the deficiencies of the prior art. By combining various input modal information such as speech, vision, and text, more flexible and efficient robotic arm control means are provided, thereby solving the problems that single-modal input in the prior art is insufficient to meet complex operation requirements and lacks dynamic environment adaptation ability.
[0010] The purpose of the present invention is achieved through the following technical solutions. In the first aspect, the present invention provides a robotic arm control method for realizing multi-modal general operation tasks, and the method includes the following steps:
[0011] Step 1: Obtain multi-modal input data required for the task, including voice instruction data, text instruction data, visual perception data in the current task environment, and robotic arm pose perception data;
[0012] Step 2: Align the input data of multiple modalities, encode them into input representation vectors with a unified form, and splice them into a multi-modal input instruction in paragraph form;
[0013] Step 3: Based on the voiceprint features carried by the voice instruction data, retrieve and extract the task-related voice data of the corresponding user from the collected user information voice database as supplementary information for the instruction, and splice them in the same format;
[0014] Step 4: Input the multi-modal input instruction into the backbone of the policy large model for instruction understanding, and generate a set of end-effector poses of the robotic arm at several future time steps, where the end-effector pose includes the action instruction of the robotic arm at each time step;
[0015] Step Five: Based on the pose of the end effector of the robotic arm, perform motion control of the robotic arm, and then obtain new multi-modal input data to update the multi-modal input instructions for perceiving the latest task status;
[0016] Step Six: The policy large model performs closed-loop control through the comparison and understanding of the initial target instruction information and the current environmental perception information, and continuously performs new motion planning until the robotic arm completes the initially set task goal.
[0017] Further, the implementation method of the said Step One includes:
[0018] Step 101: The operator inputs text instruction data, which is wrapped with system prompt words, model dialogue templates, and supplementary prompt words to form a complete text string instruction;
[0019] Step 102: The voice instruction of the operator is converted into Mel spectrograms of several frequency bands using the Short-Time Fourier Transform (STFT), and then filled into the total number of frames of a fixed length. If the original number of frames is less than this length, zero-padding is performed at the end of the time axis; otherwise, it is truncated to this length. The number of frequency bands and the total number of frames are determined by the parameters used during the pre-training of the voice feature encoder.
[0020] Step 103: Obtain visual perception data and robotic arm pose perception data through the sensor component. The visual perception data includes side-view RGB images captured by a side-view monocular camera fixed to the base of the robotic arm and RGB images captured by a monocular camera fixed to the gripper of the robotic arm. The images are subjected to size transformation, multi-image stitching, border pixel filling, and data format conversion to form an image format that the model can process. The robotic arm pose perception data is obtained through the sensors of the robotic arm body, including the coordinates of the robotic arm, the angles of each joint, and the opening and closing information of the gripper. After reading, it is concatenated with the text instruction data in string form.
[0021] Further, the implementation method of the said Step Two includes:
[0022] Step 201: Reserve special marker symbols indicating images or voices in the text instruction data to indicate the insertion positions of various modal information in the multi-modal instruction;
[0023] Step 202: For different modal information, use independent feature encoders to convert the information into a multi-modal vector space with consistent dimensions, and perform dimension alignment using the input dimensions required by the backbone of the policy large model. After completion of the alignment, different modal information is concatenated and inserted according to the reserved special marker symbols to form a multi-modal instruction paragraph.
[0024] Further, the implementation method of the said Step Three includes:
[0025] Step 301: Use a pre-trained or open-source speech recognition model to extract the user's voiceprint information carried by the collected voice commands, and use this as the reference information for retrieving the user information voice database. The collection process of the user information voice database includes the following steps: pre-collecting user personal information, pre-coding voice information, storing voice data, and updating voice information.
[0026] Step 302: Use the method of voiceprint matching to locate the personal voice information of the user who issued the current command, extract this voice information, and then splice it to the special marker marked as the user's personal information in the multi-modal command paragraph to indicate the supplementary user information.
[0027] Furthermore, the policy large model in step four adopts a vision-language-action large model obtained by fully supervised fine-tuning of the vision-language large model. The implementation method includes: a dataset construction part, a model training part, and a deployment and inference part;
[0028] The dataset construction part includes two parts: constructing a voice dialogue dataset and constructing a robotic arm operation dataset. The voice dialogue dataset is obtained through voice synthesis of the vision-text dialogue dataset open-sourced by the community and voice recognition of the vision-speech dialogue dataset, which is used to increase the model's support for end-to-end understanding of voice modality information. The robotic arm operation dataset is realized by collecting the operation process data of professional robotic arm operators using the robotic arm to perform different tasks. The original data includes task instructions, environmental perception, and robotic arm pose perception data at each moment of each task. The robotic arm pose perception data needs to be further constructed into an action vocabulary before model training, that is, the floating-point form of the pose perception data is normalized, discretized, and mapped one by one to the vocabulary numbers in the tokenizer vocabulary according to the frequency of the words in the tokenizer vocabulary from low to high;
[0029] The model training part includes two parts: voice modality training and robotic arm operation training, which are trained using the voice dialogue dataset and the robotic arm operation dataset respectively. Among them, voice modality training is the first stage of training, and robotic arm operation training is the second stage of training. The cross-entropy loss function is used to calculate the loss value in both stages of training. The true label uses the one-hot distribution method to calculate the cross-entropy with the probability distribution predicted by the model;
[0030] The described deployment and inference part reads in real time the RGB images captured by the monocular side camera and the monocular camera at the gripper of the robotic arm, as well as the pose perception data of the robotic arm, and jointly processes them with the text instruction data, voice instruction data of the current task, and supplementary information extracted from the user information voice database into a multi-modal instruction paragraph, and inputs it into the backbone of the trained policy large model to obtain a set of poses of the end effector of the robotic arm at several future time steps. Among them, the policy large model runs on the server computer, and the client computer collects instruction data, packages and sends instruction data, and obtains server response data; the server communicates with the client through a wireless network, and the client and the robotic arm transmit information and control the robotic arm through ROS topic nodes.
[0031] Further, the implementation method of step five is to perform inverse kinematics calculation on the poses of the end effector of the robotic arm at several future time steps output by the policy large model to obtain the rotation angles of each axis of the robotic arm, and then perform motor control through a PID controller; after completing the actions at several time steps planned by the large model, read the multi-modal data in the current environment again, and update the multi-modal instructions with the newly input visual perception data, robotic arm pose perception data, text instruction data, and voice instruction data, and wait for the policy large model to perform a new round of motion planning.
[0032] Further, the implementation method of step six is to construct a closed-loop control loop with the policy large model as the controller based on steps one to five, and continuously output new action plans in a dynamically changing environment through the repeated comparison and understanding of the initial multi-modal instruction information and the current environment perception information by the model until it is judged that the current task has been completed, stop the robotic arm action, and restart when waiting for the next task command.
[0033] In a second aspect, the present invention provides a robotic arm control system for implementing multi-modal general operation tasks, which system includes a multi-modal data acquisition module, a task planning module, and a motion control and feedback module;
[0034] The multi-modal data acquisition module is used to collect various modal perception data and instruction data in the current working area, perform modal alignment and encode and splice them into a multi-modal instruction paragraph, including text instruction data, voice instruction data describing tasks or supplementary information, as well as visual perception data and robotic arm pose perception data in the current task environment. In addition, it also includes a user information voice database established by extra collection before task execution;
[0035] The task planning module consists of a server and a client. The model is deployed and inferred on the server using a GPU processor. Based on the multi-modal instruction data, a set of poses of the end effector of the robotic arm for several future time steps is generated. The pose of the end effector includes the action instructions of the robotic arm at each time step. On the client, a computer equipped with only a CPU processor and an SSD solid-state drive is used to collect, forward, and process the instruction information, and the local robotic arm is controlled in the form of a network service based on the pose of the end effector of the robotic arm.
[0036] The motion control and feedback module is used to receive the planning information processed by the client in the task planning module, that is, the poses of the end effector of the robotic arm for several time steps. The rotation angles required to control each axis of the robotic arm are obtained through inverse kinematics calculation, and the motion control of the robotic arm is completed using a PID controller. After completing the planned action, the perception data and instruction data of the working area at the current moment are read again through the multi-modal data acquisition module and fed back to the task planning module for a new action plan.
[0037] In a third aspect, the present invention provides a robotic arm control device for implementing multi-modal general operation tasks. The device includes: a vision sensor, an audio sensor, a text input interface, a client computer, a server computer, and a robotic arm kit;
[0038] The vision sensor is a set of monocular cameras, which consists of a side-view monocular camera fixed to the base of the robotic arm and a monocular camera fixed above the gripper of the robotic arm. The fixed camera at the base is used to provide a stable side-view panoramic view for detecting the spatial positions of the environment and objects. The camera at the gripper can change with the movement of the robotic arm and provide real-time close-up images of the object.
[0039] The audio sensor uses a microphone array. After the audio signal is denoised, key features are extracted, including user speech and ambient sound.
[0040] The text input interface uses a mobile phone or a computer keyboard device to collect the text instruction data of the user by sending the text instructions to be executed to the client computer through a wireless network or typing them on the keyboard.
[0041] The client computer is used to store and execute part of the program code. It is equipped with a CPU processor and an SSD solid-state drive, and is directly connected to the robotic arm kit, the vision sensor, the audio sensor, and the text input interface to send control signals and collect sensor data, and communicates with the server computer through the network.
[0042] The server computer is used to deploy the trained policy large model, complete the deployment and inference of the model using a GPU processor, generate a set of poses of the end effector of the robotic arm for several future time steps, obtain the required rotation angles for controlling each axis of the robotic arm through inverse kinematics solution, and complete the motion control of the robotic arm using a PID controller; after completing the planned action, a new action plan is carried out;
[0043] The robotic arm kit includes a robotic arm with an actuator and a controller, which is responsible for executing the motion control signal processed by the client to achieve specific physical operations.
[0044] The beneficial effects of the present invention are as follows:
[0045] (1) The present invention realizes a robotic arm control strategy that supports rich modal information and an end-to-end control model training method, enabling the robot to effectively understand environmental information and human needs through multi-modal information fusion in a non-standardized and dynamically changing operating environment, and correctly execute the current task;
[0046] (2) The present invention realizes a deeper and more comprehensive multi-modal data understanding and instruction construction. Through end-to-end speech signal understanding (instead of converting to text) and the construction of instruction paragraphs with multiple modalities interleaved, it can capture more comprehensive and rich semantic information, avoid cumulative errors and information loss caused by hierarchical processing, thereby increasing the success probability of task completion;
[0047] (3) The present invention realizes a more efficient human-computer interaction and instruction construction method. Through closed-loop control planning with a policy large model as the controller and information retrieval and extraction from the user information voice database, the task information is supplemented and updated, enabling the robotic arm to have the ability to dynamically adjust the task target and correct itself in a timely manner from the wrong execution process.
[0048] (4) The present invention realizes a resource-efficient deployment and inference method. The inference model is deployed as a cloud service, and the inference requirements of one or more local devices are sent through network requests and processed in parallel, realizing the full utilization of the costly server computing power resources;
[0049] (5) The system proposed by the present invention improves the dynamic adaptation ability, autonomous error correction ability and multi-modal interaction ability of the robotic arm, effectively expands the available scenarios of the robotic arm, increases the upper limit of the task completion ability, and enables the robotic arm system to be applied in more scenarios. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0051] Figure 1 It is a flowchart of the robotic arm control method of the present invention.
[0052] Figure 2 It is a flowchart of multi-modal data processing and planning generation of the present invention.
[0053] Figure 3 It is a block diagram of multi-modal signal fusion and feedback control loop of the present invention.
[0054] Figure 4 It is a block diagram of the composition of the robotic arm system of the present invention.
[0055] Figure 5 It is a schematic diagram of the robotic arm system device provided by the embodiment of the present invention. Detailed implementation manners
[0056] The following further elaborates on the detailed implementation manners of the present invention with reference to the accompanying drawings.
[0057] Embodiment 1:
[0058] As Figure 1 shown, a robotic arm control method for implementing multi-modal general operation tasks provided by the present invention includes the following steps:
[0059] Step 1: Obtain multi-modal instruction data required for the task, including the user's voice instruction data and text instruction data input via the keyboard or sent via the wireless network, and obtain visual perception data and robotic arm pose perception data in the current task environment through the sensor component. Among them, the visual perception data includes RGB images taken by cameras from different perspectives, and the robotic arm pose perception data includes information such as the coordinates of the robotic arm, the angles of each joint, and the opening and closing of the gripper;
[0060] Step 2: Align the input data of multiple modalities, encode them into input representation vectors with a unified form, and splice them into a multi-modal input instruction in paragraph form;
[0061] Step 3: Based on the voiceprint features carried by the voice instruction data, retrieve and extract the task-related voice data of the corresponding user from the collected user information voice database as supplementary information for the instruction, and splice it in the same format;
[0062] Step 4: Input the instruction information into the backbone of the policy large model for instruction understanding, and generate a set of poses of the end effector of the robotic arm for several future time steps. The pose of the end effector includes the action instructions of the robotic arm at each time step, such as angles, coordinates, gripper opening degrees, etc.;
[0063] Step 5: Process the poses of the end effector of the robotic arm for the above-mentioned several time steps and send them into the controller for motion control of the robotic arm. Subsequently, read the data of the new sensor components and the user input data to update the multi-modal input instructions to perceive the latest task status;
[0064] Step 6: The policy large model performs closed-loop control through the comparison and understanding of the initial target instruction information and the current environmental perception information, and continuously performs new motion planning until the robotic arm completes the initially set task goal.
[0065] As Figure 2 shown, the implementation method of the above Step 1 includes:
[0066] Step 101: The operator inputs text instruction data through methods such as keyboard and network sending, and forms a complete text instruction string after packaging with system messages, model dialogue templates, and reasonable supplementary prompts.
[0067] The system message, that is, the systematic guidance for the policy large model, is used to represent the basic rules that the model should abide by, the basic expectations of the user for the model, and the role that the model needs to play, and needs to be consistent during both the training process and the inference service process;
[0068] The model dialogue template, that is, a formatted template representing the interaction between the user and the model. Due to limited training resources, in actual task scenarios, the model is usually trained and its capabilities are transferred by means of fine-tuning. Therefore, it is necessary to keep the dialogue template consistent with that used in the pre-training process of the model. For example, a dialogue template in the ChatML format can be adopted to divide the system message, user instruction data, and model answer into three roles of "system", "user", and "assistant", and distinguish them with special characters, forming a dialogue text in the form of "<|im_start|>system\n(System message)<|im_end|>\n<|im_start|>user\n(User instruction)<|im_end|>\n<|im_start|>assistant\n(Model answer)<|im_end|>", where the model answer part is generated by the model;
[0069] The supplementary prompt is used to help the user input additional supplementary requirements or information to the model and splice it into the above user instruction data. It is usually used to provide supplementary information that the model needs to know additionally under a current difficult task. For example, it can inform whether there are additional action restrictions in the current environment, the time limit of the task, or expand and explain unclear user instructions.
[0070] Step 102: The voice instruction of the operator is converted into a Mel spectrogram of several frequency bands using the Short-Time Fourier Transform (STFT), and then filled into the total number of frames with a fixed length. If the original number of frames is less than this length, zeros are padded at the end of the time axis; otherwise, it is truncated to this length. The number of frequency bands and the total number of frames are determined by the parameters used during the pre-training of the voice feature encoder. The calculation formulas for the Short-Time Fourier Transform and the Mel spectrogram generation process are as follows:
[0071]
[0072]
[0073]
[0074]
[0075] Among them, is the complex frequency spectrum calculated by STFT, is the original time-domain voice signal, is the window function, is the frame index, is the frame shift, is the frequency point index, is the number of points of FFT (Fast Fourier Transform). is the power spectrum calculated from the complex frequency spectrum. is the linear frequency used in the STFT process, is the Mel scale converted using the Mel frequency scale, represents the th weight of the Mel filter, covering several frequency bands, is the Mel spectrogram obtained by weighted summation of the power spectrum using the Mel filter bank.
[0076] Step 103: For the visual perception data and the robotic arm pose perception data obtained through the sensor component, the former includes the side-view RGB image captured by the side-view monocular camera fixed to the robotic arm base and the RGB image captured by the monocular camera fixed to the robotic arm gripper. The images are subjected to size transformation, border pixel filling, multi-image stitching, and data format conversion to form an image format that can be processed by the subsequent model. The latter is obtained through the robotic arm body sensor, including the coordinates of the robotic arm, the angles of each joint, and the opening and closing information of the gripper. After being read, it is concatenated with the text instruction data in the form of a string.
[0077] As Figure 2 shown, the implementation method of the second step includes:
[0078] Step 201: Reserve special marker symbols such as those representing images or voices in the formed text instructions to facilitate the insertion of various modal information during the construction of multi-modal instructions.
[0079] The special marker symbols are used to identify the positions where other modal information needs to be replaced and inserted during the subsequent instruction concatenation stage. Commonly used special marker symbols include " <audio>” (representing voice data), " ” (representing image data), etc. For the case of multiple consecutive pictures or multiple voices, spaces or line breaks can be used for separation.
[0080] Step 202: For different preprocessed modality information, use independent feature encoders to convert the information into their respective multi-modal vector spaces, and then use linear layers or MLP (Multi-Layer Perceptron) for dimension conversion and alignment, defaulting to the input dimensions required by the backbone of the policy large model. After alignment, different modality information is spliced and inserted according to the reserved special marker symbols to form a multi-modal instruction paragraph.
[0081] The feature encoders respectively include: a voice feature encoder, a text feature encoder, and an image feature encoder. All three encoders adopt the Transformer architecture. To reduce R & D costs and transfer the capabilities of existing powerful models as much as possible, open-source models that have been pre-trained on a large scale in academia or industry or low-cost model APIs can be used.
[0082] Such as Figure 2 As shown, the implementation method of the third step includes:
[0083] Step 301: Use a pre-trained or open-source speech recognition model in advance to extract the user voiceprint information carried by the collected voice instructions, and use this as the reference information for retrieving the user information voice database.
[0084] Step 302: Use the voiceprint matching method to locate the personal information voice data of the user who issued the current instruction, extract the voice data and then splice it to the special marker marked as the user's personal information in the multi-modal instruction paragraph, which is used to indicate the supplementary user information. If it is empty, the special marker is directly deleted.
[0085] The acquisition process of the user information voice database includes the following steps:
[0086] (1) Pre-collect user personal information: Through the method of recording, let all users in the current environment store important personal information related to the execution of tasks, such as personal operation preferences, scene item ownership, potential usage motivations, safety inclinations, etc., which are used to supplement text and voice instructions that may have unstructured and insufficient information problems;
[0087] (2) Voice information pre-encoding: Use the voice feature encoder to perform pre-encoding on all voice information, and process it into voice feature vectors that can be processed by the subsequent backbone of the policy large model, which is convenient for the splicing of multi-modal instructions. Additionally, use a speech recognition model to extract the voiceprint information in different users' voice information to improve the implementation efficiency of the voiceprint matching process.
[0088] (3) Voice data storage: The user information voice database is stored in the form of a JSON (a lightweight data interchange format) file, using the user ID as the key and the voiceprint information and user voice information as the values when indexed by the user number.
[0089] (4) Voice information update: According to the failure cases that occur when the robotic arm performs tasks and the dynamic changes in the information in the current environment, the user information voice database needs to be dynamically adjusted. For newly entered user information, the existing key-value pairs can be matched according to the user's voiceprint characteristics for user information supplementation. If the returned matching result is non-existent, new key-value pairs are created. For information that needs to be deleted or adjusted, specific user information is matched according to the voiceprint characteristics, and the user can independently select the part to be deleted or adjusted in the existing voice information storage file.
[0090] The policy large model in step four adopts a voice-vision-language-action large model obtained by full supervision fine-tuning based on VLM (Vision-Language Model), and the implementation method includes: a dataset construction part, a model training part, and a deployment and inference part.
[0091] The dataset construction part includes two parts: the construction of the voice dialogue dataset and the construction of the robotic arm operation dataset.
[0092] For the construction of the voice dialogue dataset, since the basic model for training is an open-source vision-language model that has been trained and fine-tuned on a large scale, to add support for end-to-end understanding of the voice modality without voice recognition, it is necessary to train and adjust the data in the form of voice questions and text responses. The visual voice dialogue dataset is obtained through voice synthesis using the community open-source visual text dialogue dataset and voice recognition using the visual speech dialogue dataset. It is divided into a training set and a validation set .
[0093] For the construction of the robotic arm operation dataset, it is achieved by collecting the operation process data of professional robotic arm operators performing different tasks with the robotic arm. When the operator performs a task, for each moment of each task, a side-view monocular camera fixed to the base of the robotic arm and a monocular camera fixed above the robotic arm gripper are used for shooting to obtain the perception dataset of the policy large model. . Record the task instructions in text form executed by the operator, use the pose information read by the manipulator body sensor at the current moment to form the manipulator pose perception data in text form, use the pose information read by the manipulator body sensor after being operated by the operator at several subsequent moments to form the training ground truth data in text form, and splice the above texts using the dialogue template to obtain the text instruction dataset of the policy large model. . For the situation where voice is used to supplement task information or describe task instructions, collect and store the original voice clips to obtain the voice instruction dataset of the policy large model. . In addition, use the user information voice database collection process described in Step 3 to perform additional user information data collection and construct a vector database. , which is used to assist in the construction of multimodal instruction information. Combine the obtained datasets according to the executed tasks and execution moments to form the manipulator operation dataset. , and each piece of data uses the task instruction in text form at the current moment, the task instruction in voice form at the current moment, the camera perception data at the current moment, and the manipulator pose perception data at the current moment as input data, and uses the manipulator pose perception data at several subsequent moments as the training ground truth. Among them, the task instruction in voice form and the camera perception data are both represented in the form of the storage address of the file. Divide the constructed manipulator operation dataset into a training set. and a validation set. .
[0094] The manipulator pose perception data needs to be further processed to construct an action vocabulary for model training. This process includes: converting the manipulator pose perception data back to floating-point numbers, normalizing the data, discretizing it into specific integers within a certain length interval, and mapping the floating-point form of the pose perception data to the vocabulary numbers in the tokenizer vocabulary one by one according to the frequency of the vocabulary in the tokenizer vocabulary from low to high. Through the construction of the action vocabulary, the difficult-to-traverse floating-point pose data can be unified into a finite number of numbers within a certain length interval.
[0095] The model training part includes two parts: voice modality training and manipulator operation training. Among them, voice modality training is the first stage of training, and manipulator operation training is the second stage of training.
[0096] The voice modality training takes the training set as input. To the model, load the weights of the pre-trained vision-language large model and the speech feature encoder weights, use the multi-modal data processing process in Step 1, the modal alignment and multi-modal instruction splicing process in Step 2 to obtain the multi-modal instruction paragraph, input it into the backbone of the policy large model to get the text response. During the training process, only use the text response part to calculate the loss function value to update the model weights, so that it has the normal text response ability to the input visual speech questions. After the training is completed, output the training weight file, and through the validation set Verify the effect of the obtained weight file, and select the weight file with the best performance as the optimal weight file .
[0097] For the manipulator operation training, input the training set To the model, load the weights trained through the speech modality , also use the multi-modal data processing process in Step 1, the modal alignment and multi-modal instruction splicing process in Step 2 to obtain the multi-modal instruction paragraph, input it into the backbone of the policy large model to get the model response. For the non-model response part, use masking to avoid participating in the loss function calculation, and only calculate the loss function value between the model response part and the training ground truth to update the model weights. After the training is completed, output the training weight file, and through the validation set Verify the effect of the obtained weight file, and select the weight file with the best performance as the optimal weight file .
[0098] Both of the above two training processes use the cross-entropy loss function to calculate the loss value. The true label uses the one-hot distribution method to calculate the cross-entropy with the probability distribution predicted by the model. The calculation formula is expressed as:
[0099]
[0100] where is the total loss value that needs to be optimized to minimize it, represents the length of the input sequence, represents the predicted token at the th position in the sequence, represents the context before the th position, represents the probability distribution predicted by the model, indicating that under the given condition, the probability that the model predicts the next token as is
[0101] The deployment and inference part reads in real time the RGB images captured by the monocular side camera and the monocular camera at the gripper of the robotic arm, as well as the robotic arm pose perception data, and processes them together with the text instruction data, voice instruction data of the current task, and supplementary information extracted from the user information voice database into a multi-modal instruction paragraph, and inputs it into the backbone of the trained policy large model to obtain a set of poses of the end effector of the robotic arm at several future time steps. The policy large model runs on the server computer, and the client computer collects instruction data, packs and sends instruction data, and obtains server response data. The server communicates with the client via WIFI (wireless network), and information is transmitted between the client and the robotic arm and the robotic arm is controlled via ROS topic nodes.
[0102] The autoregressive generation process of the policy large model under a given multi-modal instruction paragraph can be modeled as:
[0103]
[0104] Where, represents the prefix part in the prompt, corresponding to the multi-modal instruction part provided by the user, represents the subsequent text content predicted based on the given context, corresponding to the poses of the end effector of the robotic arm at several future time steps predicted by the policy large model, represents the current context content, including the prompt and all the tokens that have been predicted, represents the next token to be predicted based on the current context, represents the total probability of the model predicting the subsequent text content, represents the probability of predicting each subsequent token based on the current context.
[0105] The implementation method of step five is to perform inverse kinematics calculation on the poses of the end effector of the robotic arm at several future time steps output by the policy large model to obtain the rotation angles of each axis of the robotic arm, and then perform motor control through a PID controller. After completing the actions at several time steps planned by the large model, read in various modal data in the current environment again, update the multi-modal instructions with the newly input visual perception data, robotic arm pose perception data, text instruction data, and voice instruction data, and wait for the policy large model to perform a new round of motion planning.
[0106] Such as Figure 3 As shown in the figure, the implementation method of Step 6 is to construct a closed-loop control loop with the policy large model as the controller based on Steps 1 to 5. Through repeated comparison and understanding of the initial multi-modal instruction information and the current environmental perception information by the model, new action plans are continuously output in the dynamically changing environment until it is determined that the current task has been completed, the robotic arm action is stopped, and it waits to restart when the next task command arrives.
[0107] Embodiment 2:
[0108] As Figure 4 shown in the figure, a robotic arm control system for implementing multi-modal general operation tasks provided by the present invention includes: a multi-modal data acquisition module, a task planning module, and a motion control and feedback module.
[0109] The multi-modal data acquisition module uses sensors to collect perception data and instruction data of various modalities in the current working area, including multi-view RGB images captured by a side-view monocular camera fixed to the base of the robotic arm and a monocular camera fixed above the gripper of the robotic arm, text instruction data obtained through a keyboard and network transmission, and voice instruction data for describing tasks or supplementary information obtained using a microphone array. In addition, it also includes user information voice data collected and stored additionally before task execution.
[0110] The task planning module consists of a server and a client. By using a GPU processor with a high cost in the server to complete the deployment and inference of the large model with a large amount of computation, and using a low-performance computer equipped only with a CPU processor and an SSD solid-state drive in the client to collect, forward, and process instruction information, the contradiction between the large amount of computing resources required for multi-modal large model inference and the limited processor performance locally is decoupled. Based on the pose of the end effector of the robotic arm in the form of a network service, local control of the robotic arm can be achieved, and the control system cost can still be kept low under a relatively high number of requests.
[0111] Furthermore, the client in the task planning module is responsible for information transmission and parsing of local control instructions. During task execution, it first sends modality data such as text, images, and voice collected through the network to the server for processing. For the pose of the end effector of the robotic arm fed back by the server, after effective information extraction, error checking, and format conversion, it is transmitted to the motion control and feedback module for execution.
[0112] Furthermore, the server in the task planning module is responsible for model deployment, inference, and multi-modal information processing. During task execution, the multi-modal information received is processed by the model deployed on the server to form a multi-modal instruction paragraph, and the pose of the end effector of the robotic arm at multiple future time steps is obtained by inputting it into the backbone of the policy large model. Subsequently, this planning information is sent to the client through the network for processing.
[0113] The motion control and feedback module receives the planning information processed by the client in the task planning module, that is, the pose of the end effector of the robotic arm at several time steps. The rotation angles required to control each axis of the robotic arm are obtained through inverse kinematics calculation, and the motion control of the robotic arm is completed using a PID controller. After completing the planned action, the working area perception data and instruction data at the current moment are read again through the multi-modal data acquisition module and fed back to the task planning module for new action planning.
[0114] Embodiment 3:
[0115] As Figure 5 shown, a robotic arm control device for realizing multi-modal general operation tasks provided by the present invention includes: a visual sensor, an audio sensor, a text input interface, a client computer 2, a server computer 4, and a robotic arm kit.
[0116] The visual sensor is a set of monocular cameras 7, which consists of a side-view monocular camera fixed to the base of the robotic arm and a monocular camera fixed above the gripper of the robotic arm. The fixed camera at the base is used to provide a stable side-view panoramic view for detecting the spatial positions of the environment and objects, and the camera at the gripper can move with the robotic arm 5 and provide a close-up image of the object 8 to be grasped in real time.
[0117] The audio sensor uses a microphone array 6. After the audio signal is noise-reduced, key features such as user speech and ambient sound are extracted.
[0118] The text input interface uses devices such as a mobile phone 3 and a computer keyboard 1 to collect the user's text instruction data by sending text instructions to be executed to the client computer 2 via the network or by typing on the keyboard 1.
[0119] The client computer 2 is used to store and execute part of the program code, equipped with a CPU processor and an SSD solid-state drive, directly connected to the robotic arm 5, the visual sensor, the audio sensor, and the text input interface to send control signals and collect sensor data, and communicate with the server computer 4 via the network. Among them, information transmission and control of the robotic arm 5 between the client computer 2 and the robotic arm kit are carried out through ROS topic nodes.
[0120] The server computer 4 is used to deploy the trained policy large model, complete the deployment and inference of the model using a high-performance GPU processor, generate a set of poses of the end effector of the robotic arm for a number of future time steps, obtain the required rotation angles for controlling each axis of the robotic arm through inverse kinematics calculation, and use a PID controller to complete the motion control of the robotic arm; after completing the planned action, a new action plan is carried out; provide efficient and fast result response and high-concurrency robotic arm control services.
[0121] The robotic arm kit includes a robotic arm 5 (including an actuator), a controller, etc., and is responsible for executing the motion control signal processed by the client computer 2 to achieve specific physical operations.
[0122] The above embodiments are used to explain the present invention rather than limit the present invention. Any modification and change made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.< / audio>
Claims
1. A method for controlling a robotic arm to realize a multi-modal general operation task, characterized in that: The method comprises the following steps: Step 1: Obtain the multimodal input data required for the task, including voice command data, text command data, visual perception data in the current task environment, and robotic arm posture perception data; Step 2: Perform modal alignment on the input data of multiple modalities, encode them into input representation vectors with a unified form, and concatenate them into multimodal input instructions in the form of paragraphs; Step 3: Based on the voiceprint features of the voice command data, retrieve and extract the task-related voice data of the corresponding user from the collected user information voice database as supplementary information of the command, and splice them in the same format; Step 4: Input the multimodal input command into the strategy large model backbone for command understanding, and generate a set of robot end effector poses for several future time steps, wherein the end effector poses include the action commands of the robot at each time step; Step 5: Control the motion of the robot arm based on the position of the end effector of the robot arm, and then obtain new multimodal input data to update the multimodal input instructions to perceive the latest task status; Step 6: The strategy model performs closed-loop control by comparing and understanding the initial target command information and the current environmental perception information, and continuously makes new action plans until the robotic arm completes the initially set mission objectives.
2. A method for controlling a robotic arm to realize a multi-modal general operation task according to claim 1, characterized in that: The implementation method of step 1 includes: Step 101: The operator inputs text instruction data, and uses system prompt words, model dialogue templates, and supplementary prompt words to package them into a complete text string instruction; Step 102: The operator's voice command is converted into Mel spectrograms of several frequency bands using short-time Fourier transform (STFT), and then padded to a total frame number of a fixed length. If the original frame number is less than the length, zero padding is performed at the end of the time axis, otherwise it is truncated to the length; wherein the number of frequency bands and the total number of frames are determined by the parameters used in pre-training of the speech feature encoder used; Step 103: Obtain visual perception data and robot arm posture perception data through the sensor component. The visual perception data includes a side-view RGB image taken by a side-view monocular camera fixed on the base of the robot arm and an RGB image taken by a monocular camera fixed on the gripper of the robot arm. The images are resized, multi-image stitched, border pixel filled, and data format converted to form an image format that can be processed by the model. The robot arm posture perception data is obtained through the robot arm body sensor, including the coordinates of the robot arm, the angles of each joint, and the opening and closing information of the gripper. After reading, it is spliced with the text instruction data in the form of a character string.
3. A method for controlling a robotic arm to realize a multi-modal general operation task according to claim 1, characterized in that: The implementation method of step 2 includes: Step 201: reserve a special mark symbol representing an image or voice in the text instruction data to indicate the insertion position of multiple modal information in the multimodal instruction; Step 202: For different modal information, use independent feature encoders to convert the information into a multimodal vector space with consistent dimensions, and use the input dimensions required by the strategy large model backbone to perform dimensional alignment; after the alignment is completed, splice and insert different modal information according to the reserved special marking symbols to form a multimodal instruction segment.
4. A method for controlling a robot arm to realize a multi-modal general operation task according to claim 3, characterized in that: The implementation method of step three includes: Step 301: Use a pre-trained or open-source speech recognition model to extract the user voiceprint information carried by the collected voice command, and use it as reference information for retrieving the user information voice database; the collection process of the user information voice database includes the following steps: pre-collecting user personal information, pre-coding voice information, storing voice data, and updating voice information; Step 302: Use voiceprint matching to locate the personal voice information of the user who currently issues the instruction, extract the voice information and then splice it into the special mark marked as user personal information in the multimodal instruction paragraph to indicate the supplementary user information.
5. The method for controlling a robot arm to realize a multi-modal general operation task according to claim 1, characterized in that: The strategy big model in step 4 adopts the vision-language-action big model obtained by fine-tuning the vision-language big model with full supervision, and the implementation method includes: a data set construction part, a model training part and a deployment reasoning part; The data set construction part includes two parts: speech dialogue data set construction and robot arm operation data set construction. The speech dialogue data set is obtained by performing speech synthesis on the community open source visual text dialogue data set and performing speech recognition on the visual speech dialogue data set, which is used to increase the model's support for end-to-end understanding of speech modal information. The robot arm operation data set is realized by collecting the operation process data of professional robot arm operators using the robot arm to perform different tasks. The original data includes task instructions, environmental perception and robot arm posture perception data for each task and each moment. The robot arm posture perception data needs to be further constructed with an action word list before model training can be performed, that is, the floating point posture perception data is normalized and discretized according to the vocabulary usage frequency in the tokenizer word list from low to high, and mapped one by one to the vocabulary sequence number in the tokenizer word list. The model training part includes two parts: speech modality training and robotic arm operation training. The speech dialogue data set and the robotic arm operation data set are used for training respectively, wherein the speech modality training is used as the first stage of training, and the robotic arm operation training is used as the second stage of training. The two-stage training process uses the cross entropy loss function to calculate the loss value, and the real label uses the one-hot distribution method to calculate the cross entropy with the probability distribution predicted by the model; The deployment reasoning part reads the RGB images taken by the side-view monocular camera and the monocular camera at the gripper of the robot arm, and the robot arm posture perception data in real time, and processes them together with the text instruction data of the current task, voice instruction data, and supplementary information extracted from the user information voice database into a multimodal instruction paragraph, and inputs the trained strategy large model backbone to obtain a set of robot arm end effector postures for several future time steps, wherein the strategy large model runs on the server computer, and the client computer collects instruction data, packages and sends instruction data, and obtains server response data; the server and the client communicate through a wireless network, and the client and the robot arm transmit information and control the robot arm through ROS topic nodes.
6. A method for controlling a robot arm to realize a multi-modal general operation task according to claim 1, characterized in that: The implementation method of step five is to perform inverse kinematics solution on the robot arm end effector posture for several future time steps output by the strategy large model to obtain the rotation angle of each axis of the robot arm, and then control the motor through a PID controller; after completing several time step actions planned by the large model, re-read the multiple modal data in the current environment, use the newly input visual perception data, robot arm posture perception data, text command data and voice command data to update the multimodal instructions, and wait for the strategy large model to perform a new round of motion planning.
7. A method for controlling a robotic arm to realize a multi-modal general operation task according to claim 1, characterized in that: The implementation method of step six is to construct a closed-loop control circuit with the strategy large model as the controller based on steps one to five, and through the model, repeatedly compare and understand the initial multimodal command information and the current environmental perception information, continuously output new action plans in a dynamically changing environment, until it is determined that the current task has been completed, stop the robot arm action, and wait for the next task command to restart.
8. A robotic arm control system for realizing multi-modal general operation tasks, characterized in that: The system includes a multimodal data acquisition module, a task planning module, and a motion control and feedback module; The multimodal data acquisition module is used to collect perception data and instruction data of various modes in the current working area, perform modal alignment and encode and splice them into multimodal instruction paragraphs, including text instruction data, voice instruction data describing tasks or supplementary information, and visual perception data and robot arm posture perception data in the current task environment. In addition, it also includes a user information voice database that is additionally collected and established before the task is executed; The task planning module consists of a server and a client. The model deployment and reasoning are completed on the server using a GPU processor, and a set of robot arm end effector postures for several future time steps are generated based on multimodal instruction data. The end effector postures include the action instructions of the robot arm at each time step. The client uses a computer equipped with only a CPU processor and an SSD solid-state hard disk to collect, forward and process instruction information, and controls the local robot arm based on the robot arm end effector posture in the form of a network service. The motion control and feedback module is used to receive the planning information processed by the client in the task planning module, that is, the position of the end effector of the robotic arm at several time steps, obtain the rotation angle required to control each axis of the robotic arm through inverse kinematics solution, and use the PID controller to complete the motion control of the robotic arm; after completing the planned action, the current working area perception data and instruction data are re-read through the multimodal data acquisition module, and fed back to the task planning module for new action planning.
9. A robot arm control device for realizing multi-modal general operation tasks, characterized in that: The device includes: visual sensors, audio sensors, text input interface, client computer, server computer and robotic arm kit; The visual sensor is a set of monocular cameras, which are composed of a side-view monocular camera fixed to the base of the robotic arm and a monocular camera fixed above the gripper of the robotic arm. The fixed camera at the base is used to provide a stable side-view panoramic view for detecting the spatial position of the environment and the object, and the camera at the gripper can change with the movement of the robotic arm to provide a close-up image of the object in real time. The audio sensor uses a microphone array, and after the audio signal is processed by noise reduction, key features are extracted, including user voice and ambient sound; The text input interface uses a mobile phone or computer keyboard device to collect the user's text instruction data by sending it to the client computer through a wireless network or by typing the text instruction to be executed on the keyboard; The client computer is used to store and execute part of the program code, is equipped with a CPU processor and an SSD solid state drive, is directly connected to the robotic arm kit, the visual sensor, the audio sensor, and the text input interface to send control signals and collect sensor data, and communicates with the server computer via the network; The server computer is used to deploy the trained strategy model, use the GPU processor to complete the deployment reasoning of the model, generate a set of robot arm end effector positions for several future time steps, obtain the rotation angle required to control each axis of the robot arm through inverse kinematics solution, and use the PID controller to complete the motion control of the robot arm; after completing the planned action, perform new action planning; The robotic arm kit includes a robotic arm with an actuator and a controller, which is responsible for executing the motion control signal processed by the client to realize specific physical operations.
Citation Information
Patent Citations
Home service robot cloud multimode dialogue method, device and system
CN109658928A
Systems and methods for operations a robotic system and executing robotic interactions
CN112088070A