Mechanical arm natural language instruction control system and method based on large language model
By combining a large language model with a visual perception module, a parameterized atomic skill library and motion control module are constructed, which solves the shortcomings of the robotic arm control system in flexible interaction and adaptation to complex scenarios. It enables the efficient conversion of natural language instructions from non-professionals into robotic arm action sequences, improving the efficiency of human-machine collaboration and the system's cross-scenario generalization capabilities.
Patent Information
- Application Number
- CN202511047774.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-17
AI Technical Summary
Existing robotic arm control systems have shortcomings in flexible interaction, adaptation to complex scenarios, and human-machine collaboration efficiency. It is difficult to efficiently convert natural language instructions from non-professionals into action sequences that can be executed by the robotic arm. In addition, the functional modules of traditional robotic arm control systems are independent and cannot fully combine the advantages of flexible human decision-making and precise robot execution.
By combining a large language model with visual perception and motion control modules, the system achieves precise mapping from natural language to structured action commands through the construction of a parameterized atomic skill library, efficient LoRA parameter fine-tuning technology, and dynamic prompt engineering to optimize command parsing. The visual perception module uses a YOLOv8 lightweight detection network optimized through transfer learning, combined with a binocular structured light depth camera for object detection. The motion control module builds a robotic arm model based on forward kinematics. The system's backend implements module collaboration through ROS middleware, while the frontend builds a human-machine interface based on Gradio and Gazebo.
It realizes intelligent human-machine collaboration in assembly scenarios, improves the accuracy of command parsing and response speed, has logical explainability, supports voice input and three-dimensional visualization of natural language commands, meets the rapid deployment needs of flexible production scenarios, and lowers the operating threshold for workers.
Smart Images

Figure BDA0005522236040000042 
Figure BDA0005522236040000044 
Figure BDA0005522236040000051
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of industrial automation, and particularly relates to a mechanical arm natural language instruction intelligent interaction method fusing a large language model, visual perception and mechanical arm control. BACKGROUND
[0002] In the field of high-end equipment manufacturing, the aerospace and precision instrument industries show a trend of multi-variety and small-batch production, which puts forward higher requirements on the flexibility and rapid response capability of production systems. At the same time, the operation requirements of the production site are becoming more and more complex, and an interactive mode that is convenient for non-professionals to quickly issue complex instructions is needed to efficiently complete assembly, debugging and other tasks. In the traditional human-machine cooperation mode, human-machine collaborative work is often isolated by physical fences and time staggered work methods to ensure safety, which not only affects the efficiency of human-machine interaction, but also is difficult to fully combine the advantages of human flexible decision-making and robot precise execution, resulting in difficulty in improving production efficiency. In terms of mechanical arm control, traditional mechanical arms mostly rely on preset programming to execute tasks, and when facing new scenes and tasks, more time and manpower need to be invested for program updating and debugging, and the adaptability is difficult to meet the production needs of rapid iteration of high-end equipment manufacturing.
[0003] With the major breakthrough of large language models in the field of natural language processing and the continuous development of machine vision and intelligent control technology, how to break through the traditional technical bottlenecks and realize the intelligent and flexible control of mechanical arms has become the focus of the industry. Most existing mechanical arm control systems have relatively independent functional modules, and there are obvious deficiencies in the coordination of natural language instruction analysis, environment perception and motion control, which cannot efficiently convert complex natural language instructions issued by non-professionals into executable action sequences of mechanical arms. Therefore, it is of great significance to develop a mechanical arm control system and method that can adapt to the flexible needs of multi-variety and small-batch production, support natural language interaction of non-professionals, and ensure safe and efficient human-machine collaboration, to promote the upgrading of the high-end equipment manufacturing industry. SUMMARY
[0004] In view of the deficiencies of the existing mechanical arm control system in flexible interaction, complex scene adaptation and human-machine collaborative efficiency, the purpose of the present application is to provide a mechanical arm natural language instruction control system and method based on a large language model. The method realizes intelligent human-machine collaboration in the assembly scene by constructing intelligent mapping and execution of natural language instructions to mechanical arm action sequences. First, a parameterized atomic skill library for the assembly scene is established, and common human-machine collaboration instructions are abstracted into four basic action units of detection, grasping, moving and releasing. On this basis, a DeepSeek-R1-Distil-Llama-8B large language model is used as the semantic understanding core, and the mechanical arm control field knowledge is injected through the LoRA parameter efficient fine-tuning technology to build a three-level analysis architecture including a task intent extraction layer, an action sequence decomposition layer and a parameter configuration generation layer. The dynamic prompt engineering is combined to optimize the instruction analysis accuracy, and the accurate mapping from natural language to structured action instructions is realized. The visual perception module uses a lightweight detection network YOLOv8 optimized by transfer learning, and a three-dimensional scene is built with a binocular structured light depth camera. The motion control module builds a mechanical arm model based on the robot forward kinematics, and uses MovIet! to realize the motion planning of atomic actions. The system backend realizes the communication and cooperation of various modules through the ROS middleware, and the front end is based on Gradio to integrate the Dolphin speech text large model and the Gazebo physical simulation engine to build a virtual-real combined human-machine interaction interface, supporting voice input of natural language instructions, three-dimensional visualization, forming a control system.
[0005] In order to achieve the above purpose, the present application adopts the following technical solutions:
[0006] The mechanical arm natural language instruction control system and method based on a large language model disclosed by the present application comprise the following steps:
[0007] Step 1: Assembly scene mechanical arm atomic skill library building and natural language instruction analysis, decomposing the common actions of the mechanical arm in the assembly scene for the worker to deliver tools and parts as atomic skills, and pre-training the large model to understand the natural language instruction and convert it into an atomic skill sequence.
[0008] Step 2: Build a mechanical arm intelligent perception system to realize target detection and three-dimensional coordinate acquisition of the target using a binocular camera.
[0009] Step 3: Obtain the coordinate transformation matrix of the mechanical arm by kinematic modeling, and realize the motion control and planning of the mechanical arm by solving the motion trajectory of the mechanical arm through inverse kinematics.
[0010] Step 4: Build a human-machine interaction platform, add a voice-to-text module and a virtual-to-real module, and build a front-end web interaction interface.
[0011] The implementation method of step 1 is
[0012] Step 1.1: Parameterized atomic skill library construction
[0013] The present application aims to construct a parameterized atomic skill library for common instructions in human-machine collaboration process under assembly scene, and provides basic semantic units for the mapping of natural language instructions to robot arm actions. First, based on the kinematic characteristics of robot arm and the requirements of assembly process, the operation instructions are abstracted into four basic atomic action types: (1) Detect: based on visual system to identify the three-dimensional coordinates of target object; (2) Grasp: the end effector establishes physical contact with the target object and forms a stable constraint; (3) Move: the robot arm executes trajectory planning and position change in space; (4) Release: the constraint relationship between the end effector and the target object is released. The mathematical model of atomic skill is as follows:
[0014] A = {E, P, F}
[0015] E: represents the environmental information input, including the environmental state parameters related to atomic action in the assembly scene, specifically including the pose of the target object, the initial pose of the robot arm end effector and the joint torque limit of the robot arm. P: represents the action planning and execution, the atomic action execution strategy generated based on the environmental information E, including the robot arm motion path planning and control parameter configuration. F: represents the action execution feedback, which is the state feedback parameter for judging whether the atomic action is successfully completed, taking the discrete set {success, failure, in progress} as the value.
[0016] Step 1.2: Action instruction semantic analysis dataset construction;
[0017] First, based on the semantic analysis requirements of robot arm operation, a set of instruction classification related to industrial scenes such as assembly and transportation is constructed. In data collection, typical standard operation instructions and natural language descriptions of historical operations are integrated to cover a variety of application scenarios as much as possible. In the data structure processing stage, a unified triple annotation framework is designed, and each sample contains instruction text (instruction), operation description (input) and control function sequence (output). Among them, the control function uses the parameterized atomic skill expression described in step 1.1, such as detect(obj), move(pos), gripper_closing(), gripper_openning(), which realizes the atomic decomposition of complex tasks. In the annotation process, the instruction text should be expressed in natural language as much as possible to avoid using professional terms and make the instruction more close to daily expression; the operation description part should clearly define the target object and accurately describe its spatial relationship to facilitate understanding of the task requirements; the control function sequence should correspond to the actual executable atomic action of the robot arm, and the parameter value should also meet the kinematics constraint condition to ensure that the annotated content has actual operability.
[0018] To further enhance the generalization ability of the model, data enhancement strategies are implemented from the language and operation levels. At the language level, the expression form of the instruction is enriched through synonym replacement, sentence restructuring, entity generalization, etc.; at the operation level, parameter perturbation is used, such as setting a position offset of ±5mm, changing the action sequence, and increasing the complexity of the task scene, etc. to generate semantically equivalent but different variant samples. In addition, for instructions with easily confused semantics, contrast learning samples are designed, such as "put A on B" and "put B on A" contrast training samples, to help the model better understand and distinguish different spatial relationships.
[0019] Step 1.3: Fine-tune the distilled large model using low-rank adaptive technology;
[0020] In the model architecture modification link, a small parameter model is obtained by knowledge distillation of a large parameter large language model. LoRA technology is used to modify the attention mechanism of the Transformer layer, and the weight update calculation method is as follows:
[0021] W 1 =W 0 +B·A
[0022] In the formula, W 0is the initial weight parameter; B is the dimension reduction matrix with the dimension of dxr(d is the input dimension and r is the low rank value); A is the dimension increasing matrix with the dimension of rxk(k is the output dimension). The specific operation of LoRA fine-tuning is: freezing all parameters of the original model, training only the adapter weight; adding LoRA module in the linear transformation layer of Query matrix and Value matrix; adopting bilinear interpolation to realize the dynamic fusion of original weight and adapter weight. The parameters required to configure the LoRA module include rank and scaling factor lora_alpha; the target fine-tuning module is locked in the "q_proj" and "v_proj" of the attention mechanism; in addition, the lora_dropout rate also needs to be adjusted according to the actual training effect. The task type is set to causal language model. In the model training, the precision training and gradient accumulation step number need to be set according to the hardware memory condition; the learning rate, warm-up ratio and weight decay parameters are adjusted according to the training log. In the training and verification strategy, multiple indexes such as command analysis accuracy, action sequence generation F1 value and parameter prediction error are monitored, the early stopping mechanism of stopping when the performance of continuous 5 verification cycles does not improve is set, and the ensemble learning strategy is used to integrate 3 checkpoint models with the best performance.
[0023] 3. The large language model-based robot arm natural language instruction control system and method of claim 1, wherein the step 2 is implemented by,
[0024] Step 2.1: target detection data set construction;
[0025] First, a camera is used to capture videos of the assembly scene, ensuring complete recording of tool states and environmental information during robot operation. Video capture covers different lighting conditions (such as natural light, artificial lighting), shooting angles (front view, side view, overhead view) and tool placement postures to enhance the scene diversity of the data set. In the image frame extraction stage, an interval sampling strategy is adopted to obtain basic samples at fixed intervals (such as extracting 1 frame every 5 frames); then the internal parameter matrix obtained based on camera calibration is used to perform distortion correction on the extracted image. In the labeling stage, the LabelImg tool is used for manual labeling, following the COCO data set labeling specification. The tool target needs to be framed in the image, the corresponding class label (such as screwdriver, wrench, clamp, etc.) is input, and the pose information (position coordinates and attitude quaternion) of the tool in the robot coordinate system is recorded. In the data set enhancement stage, methods such as geometric enhancement, photometric enhancement, contrast transformation and Gaussian noise addition are used to expand the original labeled image in multiple dimensions. Finally, the data set is divided into training set, validation set and test set in proportion, and each sample contains labeled image file, corresponding XML label file and pose parameter text file.
[0026] Step 2.2: YOLOv8 model training combined with transfer learning;
[0027] The universal features extracted by the backbone network in the pre-trained model are reserved, the weights in the YOLOv8 pre-trained model based on the COCO dataset are taken as the initial weights of the training, the backbone network is frozen before the training, only the head network of the model is trained, and the extracted features are used to complete the classification task of common tools and parts, that is, the transfer learning process. The loss function comprehensively considers the coordinate loss and recognition classification loss of the common tool and part positioning. The classification loss calculation process of the tool and part is as follows:
[0028] In the formula, n is the number of tools and parts; σ(x i ) is a Sigmoid function; y i is a symbol function, which takes a value of 1 if the recognized key part i belongs to the real category, otherwise 0. The coordinate loss calculation of the common tool and part positioning adopts CIoU loss, and the calculation process is as follows:
[0029]
[0030] In the formula, A and B are the coverage areas of the predicted frame and the true value respectively; p is the diagonal distance of the minimum closed area of the predicted frame and the true frame; b is the center point of the tool and part recognition predicted frame; b gt is the center point of the true frame; a is a weight coefficient; v is used to measure the consistency of the aspect ratio, and its calculation process is as follows:
[0031]
[0032] In the formula, w and h are the width and height of the predicted frame respectively; w gt and h gt are the width and height of the true frame respectively.
[0033] Step 2.3: Building a binocular vision module;
[0034] First, the binocular camera is calibrated, and the parameters of the left and right cameras are obtained using the Zhang calibration method. The specific steps of the Zhang calibration method are as follows: (1) Data collection: Prepare a checkerboard calibration plate with clear and identifiable corner points. Place the calibration plate in different directions and positions, and use the camera to take multiple photos from multiple angles. At least three photos are required, but in order to improve the accuracy of calibration, it is usually recommended to use more images. (2) Corner detection: Automatically detect the corner points of the checkerboard in each image. These corner points are pixels in the image coordinate system, which can be automatically found and accurately located by image processing software (such as OpenCV). (3) Establish coordinate relationship: Assuming that the size of each square of the checkerboard is known, the physical position of each corner point can be mapped to the world coordinate system. Because the calibration plate is placed on a plane, the Z coordinates of these corner points are usually set to 0. (4) Calculate camera parameters: Use a mathematical model to describe the imaging process of the camera, which involves internal and external parameters. The internal parameters include focal length, principal point coordinates and distortion coefficients. The external parameters include the rotation and translation matrix of the camera relative to the calibration plate in each image. The camera intrinsic parameter matrix is as follows:
[0035]
[0036] The distortion coefficient k matrix is shown below:
[0037] k=[k1 k2 p1 p2 k3]
[0038] On this basis, the homogeneous transformation matrix H is calculated, which includes the rotation matrix R and the translation vector T. The specific form of the homogeneous transformation matrix H is as follows:
[0039]
[0040] After calibration, the images taken by the left and right cameras are corrected. Based on the intrinsic and extrinsic matrix obtained by calibration, the correction transformation matrices R1, R2 and projection matrices P1, P2 are calculated, and then the original image is mapped to the corrected image plane so that the polar lines of the left and right images are collinear and parallel to the horizontal axis of the image, eliminating image distortion and unifying the coordinate system. In the disparity calculation link, the Semi-Global Block Matching (SGBM) algorithm is used for the corrected left and right images. It is necessary to set the matching cost calculation window size and the disparity search range, and adjust the uniqueness ratio threshold and smoothing parameter according to the generated disparity map. Finally, the disparity is converted into depth information according to the principle of triangulation. The depth calculation formula is
[0041]
[0042] Where f is the camera focal length, b is the baseline distance of the binocular camera, and d is the parallax value. By substituting the parameters into the depth value, the corresponding depth image can be generated.
[0043] Step 2.4: Target three-dimensional coordinate acquisition
[0044] Take the center point of the target obtained by the target detection model, and use the depth vision module built in step 2.3 to get the coordinates of the target point in the camera reference system. Through hand-eye calibration conversion, the coordinates of the target point in the robot base coordinate system are obtained. Using the eye-on-hand scheme, the solution formula of the hand-eye calibration parameter is as follows:
[0045]
[0046] In the formula is the conversion matrix from the fixture coordinate system to the robot base coordinate system; is the conversion matrix from the camera coordinate system to the fixture coordinate system; is the conversion matrix from the calibration board coordinate system to the camera coordinate system. can be solved by robot kinematics, can be obtained by camera calibration, and Tsai method can be used to solve the hand-eye calibration parameter.
[0047] 4. The large language model-based robot natural language instruction control system and method of claim 1, wherein the step 3 implementation method is
[0048] Step 3.1: Robot kinematics model building
[0049] A conversion matrix describing the joint position and direction is constructed using forward kinematics. Using the Denavit-Hartenberg (D-H) method, the conversion from each joint to the next joint can be represented by four parameters: link length a, link twist a, link offset d, and joint angle q. The URDF file describes the coordinate system of each joint and its constraint conditions. Nonlinear equations are solved using inverse kinematics. Set the target pose as P=[x d ,y d ,z d ,R x ,R y ,R z ], and the joint variable q=[q1,q2,…,q n ] needs to be solved. By separating the position and attitude components, a nonlinear equation system F(q)=0 is constructed. From the multiple solutions, the solution with the smoothest joint space trajectory is selected, and a collision-free path is generated by OMPL.
[0050] Step 3.2: Robot atomic action design
[0051] Four types of actions, detection, movement, grabbing, and release, are designed for the atomic skills of the robot arm described in step 1.1. Detection action: input the target name of detection, call the trained target detection model and depth vision module, output the three-dimensional coordinates of the target relative to the base of the robot arm, and the hardware is a binocular camera; Movement action: input the target pose of the end effector movement of the robot arm, call the MoveIt! robot arm motion planner, and control the end effector of the robot arm to move to the specified position; Grabbing action: perform the closing jaw operation, and the hardware is the end effector of the robot arm; Release action: perform the opening jaw operation, and the hardware is the end effector of the robot arm.
[0052] 5. The large language model-based robot arm natural language instruction control system and method of claim 1, wherein the implementation method of step 4 is,
[0053] Step 4.1: Voice instruction input layer design;
[0054] The voice signal is transmitted to the main control unit through the USB interface, and the acoustic feature extraction is performed through the locally deployed Dolphin voice-to-text large model. The model uses Mel-frequency cepstral coefficients (MFCC) combined with a log-mel filter bank to convert time-domain signals into feature vector sequences, generating an acoustic feature matrix. In the semantic conversion stage, the acoustic feature matrix is input into the encoder module of the Dolphin model, and the feature extraction and context modeling are performed through the Transformer structure. The decoder uses an attention mechanism combined with pre-trained language model parameters to generate the most probable text sequence. Finally, it is automatically converted into full-width characters, unified punctuation format, and adds metadata such as time stamp and recognition confidence, forming a structured input that can be parsed by the downstream natural language processing module.
[0055] Step 4.2: Visualize the virtual motion process;
[0056] Firstly, the virtual model of the manipulator is built in Gazebo, and the URDF (Unified Robot Description Format) language is used for three-dimensional modeling of the manipulator. According to the actual geometric parameters of the manipulator, the size, mass distribution and inertia matrix of each link are accurately defined; the XACRO macro is used to define the kinematic parameters of the joints, including joint type (revolute joint, prismatic joint), motion range and limit constraint. Through the SDF (Simulation Description Format) file, the physical properties are embedded, the friction coefficient between links and the collision detection threshold are set, and the simulation time step is configured to ensure that the dynamic characteristics of the virtual model are highly consistent with the actual system. Then, based on the ROS (Robot Operating System) architecture, the data communication link is built, the data acquisition node is deployed in the actual manipulator control cabinet, the joint angle and end effector pose data are collected through the real-time drive interface, the ROS topic is established, and the collected data is published in binary format in real time. The ROS subscription node is created on the virtual model side, and the callback function mechanism is used to trigger the virtual manipulator state update when the actual manipulator data is received, and the joint angle and end effector pose of the actual manipulator are mapped to the virtual manipulator.
[0057] Step 4.3: Front-end web interface construction;
[0058] The interface function integration is realized by calling Gradio library through Python language, and the interface design adopts three-column layout, mainly including text direct input, speech to text and large model reply three modules. The text direct input module uses <gr.Textbox> component to create input box, which is bound with large model reply module through <gr.Button> click event. When the button is clicked, the input text is transmitted to the deployed large language model inference interface; the speech to text module creates a microphone input interface through <gr.Audio> component, sets the real-time recording function and opens the specified audio data file path format. The window button triggers the dolphin speech recognition model to convert the audio file to text and output to <gr.Textbox> component to display the recognition result. The buttons of the text box window are also bound with the large model reply module through <gr.Button> click event.
[0059] Beneficial effects:
[0060] 1. The large language model-based mechanical arm natural language instruction control system and method disclosed in the present application, by constructing intelligent mapping and execution of natural language instructions to mechanical arm action sequences, realizes intelligent human-machine collaboration in assembly scenarios. The method proposed in the present application breaks through the limitations of traditional control systems in flexible interaction and scene generalization ability, and has the characteristics of high instruction analysis accuracy, fast response speed and strong logical interpretability.
[0061] 2. The present application fuses lightweight large language model and domain knowledge enhancement technology, uses LoRA parameter efficient fine-tuning method to adapt the basic model obtained by large parameter large language model through knowledge distillation technology to mechanical arm control field, and combines dynamic prompting engineering to construct context-aware instruction analysis framework, effectively solves the analysis problem of fuzzy expression and implicit constraint in multi-modal instruction, and significantly improves the cross-scene generalization ability and logical interpretability of the system.
[0062] 3. The parameterized atomic skill library architecture proposed in the present application abstracts the basic operation in the assembly scene into four types of combinable action units of "detection-grab-move-release", which realizes common human-machine collaboration operation. The atomic skill library can be expanded to quickly improve the collaboration ability of the mechanical arm system and realize specific tasks.
[0063] 4. The present application realizes multi-module collaboration of semantic analysis, visual perception and motion control through ROS middleware, supports voice input of natural language instructions, three-dimensional visual pre-performance in Gazebo environment and real-time interaction feedback, meets the rapid deployment needs of flexible production scenes and significantly reduces the operation threshold of workers. BRIEF DESCRIPTION OF DRAWINGS
[0064] The present application will be further described below in conjunction with the drawings and examples, wherein:
[0065] Figure 1 is the large language model-based mechanical arm natural language instruction control human-machine collaboration system framework design diagram of the present application.
[0066] Figure 2 is the loss change diagram of the LoRA fine-tuned large language model in the example of the present application.
[0067] Figure 3 is the tool and part delivered in the example of the present application.
[0068] Figure 4 is the classification loss change curve of the target detection model trained in the example of the present application.
[0069] Figure 5 is the precision, recall rate and average accuracy change curve of the target detection model trained in the example of the present application.
[0070] Figure 6 is the stereoCameraCalibration calibration toolbox interface of the dual target in the present embodiment example.
[0071] Figure 7 is the depth disparity map of the binocular camera in the present embodiment example.
[0072] Figure 8 is the Panda robot arm link mechanism diagram used in the present embodiment example.
[0073] Figure 9 is the visualization of the Panda robot arm in Gazebo in the present embodiment example.
[0074] Figure 10 is the flowchart of the virtual robot arm implementation in the present embodiment example.
[0075] Figure 11 is the interactive front-end web page interface built in the present embodiment example. Implementation method
[0076] To make the purpose, technical solutions and advantages of the present application clearer, the content of the invention is further described below in combination with the drawings and examples.
[0077] As shown in Figure 1 , the robot natural language instruction control man-machine collaboration system and method based on a large language model disclosed in the present embodiment example, the specific implementation steps are as follows:
[0078] Step 1: Assemble the scene robot atomic skill library and parse the natural language instruction. For example, the worker transfers the nut to the robot through the voice command, decomposes the action of the robot to the worker to transfer the nut as an atomic skill, and pre-trains the large model to understand the natural language instruction and converts it into an atomic skill sequence. The pre-trained large model needs to build a corresponding data set.
[0079] Step 1.1: Parameterized atomic skill library construction
[0080] The action of the robot in the assembly scene to transfer the nut to the worker can be decomposed into four steps: moving the end effector to the nut position, grasping the nut, moving the nut to the worker's workbench, and releasing the nut. The prerequisite for moving the end effector to the nut position is to detect the target position. Based on the above robot movement process, the operation instruction is abstracted into four basic atomic action types: (1) Detect: Based on the vision system to identify the three-dimensional coordinates of the target object; (2) Grasp: The end effector establishes physical contact with the target object and forms a stable constraint; (3) Move: The robot performs trajectory planning and position change in space; (4) Release: Release the constraint relationship between the end effector and the target object.
[0081] Step 1.2: Action instruction semantic parsing dataset construction
[0082] First, based on the action of the robot arm delivering nuts to the worker in the assembly scene, a unified dataset triple annotation framework is constructed, and each sample contains instruction text (instruction), operation description (input), and control function sequence (output). Among them, the instruction text is in the form of "deliver the nuts on the box to me"; the operation description is to convert the natural language instruction into a robot arm control function sequence; the control function uses the parameterized atomic skill expression described in step 1.1, such as detect(obj), move(pos), gripper_closing(), and gripper_openning(), to realize the atomic decomposition of complex tasks. When constructing the dataset, the instruction text is described in natural language to avoid professional terms. In addition, in order to increase the generalization of the dataset, at the language level, synonym replacement, sentence restructuring, and entity generalization are used, such as adding the sample "please give me the nuts" to "deliver the nuts to me"; at the operation level, parameter perturbation, action sequence transformation, and task scene complexity are introduced, such as adding the sample "deliver the nuts on the black box to me" to "deliver the nuts to me"; at the same time, for instructions that are easy to confuse, contrast learning samples are designed, such as the contrast training of "put the nuts on the toolbox" and "put the toolbox on the nuts". Finally, a dataset containing 219 samples is obtained and saved in json format.
[0083] Step 1.3: Low-rank adaptation technology for pre-training large language model fine-tuning
[0084] In the model training phase, an efficient fine-tuning technique based on low-rank adaptation (LoRA) is used to adapt the pre-trained language model DeepSeek-R1-Distill-Llama-8B to the task. The training process is based on mixed precision (BF16) and gradient accumulation strategy to optimize memory usage and improve training efficiency. The maximum length of the model input sequence is set to 2048, and the data preprocessing stage uses parallel processing (16 threads) to standardize the cutting of the training set manipulation.json, ensuring that the input format matches the deepseek3 template. The training is performed for a total of 15 complete cycles (epochs), with an effective batch size of 16 (single device batch of 2, 8 times gradient update), a total of 210 training steps, and a parameter update interval of every 8 micro-batches for gradient backpropagation. The learning rate scheduling uses a cosine decay strategy with an initial learning rate of 5e-4 and a 21-step linear warm-up (10% of the total steps), and the gradient norm clipping threshold is set to 1.0 to prevent gradient explosion. The rank in the LoRA low-rank adapter is defined as 4, the scaling factor a is 8, and the proportion of the two is 2:1, the adaptation module covers all trainable layers of the model (lora_target=all), and the Dropout probability is set to 0.2 to alleviate the risk of overfitting in small sample scenarios. The training process records indicators every 5 steps, the loss curve is tracked in real time through the visualization module, and the model checkpoint is saved every 100 steps to achieve fault recovery and later verification. The loss curve is shown in Figure 2 As shown. Finally, the hardware 3090Ti training running time is 421 seconds, and the loss function converges to 0.2537, which meets the expected nonlinear optimization goal.
[0085] Step 2.1: Target detection data set construction
[0086] Videos of each tool and part are taken, and OPENCV is used to extract frames from the video. The video is expanded to have 1000 images of each tool and part in the data set, and there are 5 categories in the data set, including wrench, screwdriver, nut, needle-nose pliers, and nail. Since the YOLO algorithm is a supervised deep learning, LABELIMG is used to label the data set, and each tool and part is labeled. The position and category of the label are normalized and arranged in txt document format to form a label file, i.e. each tool and part data set.
[0087] Step 2.2: YOLOv8 model training combined with transfer learning
[0088] The pre-trained YOLO model is trained using transfer learning: the backbone network weights of the pre-trained yolov8n.pt model are frozen, the batch size is set to 10, the image size is set to 640x640, the learning rate is set to 0.01, and the model is trained for 100 epochs on hardware NVIDIARTX 3090. The loss curve during training is shown in Figure 4 As shown in the figure, the bounding box regression loss decreases rapidly and stabilizes at 0.7, indicating that the accuracy of target position prediction has improved; the classification loss decreases from the initial 2.5 to 0.5, reflecting the enhanced ability to distinguish the five target classes (nuts, screws, screwdrivers, wrenches, pliers); and the loss centered on the distribution smoothly decreases to 0.95, indicating that the bounding box distribution modeling has improved. The precision of the model and the evaluation index curve during training are shown in Figure 5 During training, the precision quickly reaches 1.0 and remains stable, indicating strong sample recognition ability. The recall rate quickly rises from a low value to 1.0 and remains stable, reflecting effective target detection ability. The average precision quickly rises and stabilizes at 0.8, indicating balanced performance at different IoU thresholds.
[0089] Step 2.3: Building a binocular vision module
[0090] The calibration chessboard used in this example has a size of 9x12 and a side length of 20mm. The chessboard is photographed from multiple angles using a binocular camera, and a total of 20 photos are taken. The left and right views are saved separately after being divided. The MATLAB software stereoCameraCalibration calibration toolbox is used for calibration, and images with a reprojection error greater than 0.5 are deleted, as shown in Figure 6 The left and right camera parameters after calibration are as follows:
[0091] The intrinsic matrix of the left camera is:
[0092]
[0093] The distortion coefficients of the left camera are: [-0.0791 0.2001 1.7313e-04 4.1040e-04 0
[0094] The intrinsic matrix of the right camera is:
[0095]
[0096] The distortion coefficients of the right camera are: [-0.0762 0.1860 -9.3685e-04 -6.1099e-04 0
[0097] The rotation matrix is:
[0098]
[0099] Translation vector:
[0100] -117.6817 -1.0345 1.0424
[0101] Then the left and right photos are corrected for distortion and stereoscopic correction, the pixels of left and right views are matched using SGBM algorithm, and then the depth map is calculated using OPENCV, as shown in Figure 7
[0102] Step 2.4: Target three-dimensional coordinate acquisition
[0103] In this example, the eye-in-hand scheme is adopted, the calibration board and the base of the robot arm are fixed, the robot arm end effector is moved multiple times, the corresponding calibration board picture and robot arm pose matrix are used to solve the function cv2.calibrateHandEye of OPENCV, and finally the conversion matrix is obtained as follows:
[0104]
[0105] Step 3.1: Kinematic model of robot arm
[0106] In this example, the Panda robot arm designed by Franka Emika Company is used, and its linkage mechanism diagram is shown in Figure 8 . Set the rotation axis of each joint as z0 to z7, use D-H method, and get its D-H parameter table as follows:
[0107]
[0108]
[0109] Motion planning is realized through MoveIt! motion planner.
[0110] Step 3.2: Design of robot arm atomic action
[0111] Four types of actions of detection, movement, grabbing and releasing are designed for the robot arm atomic skill in step 1.1. Python is used to realize the design, set the detection function detect(obj), where obj is the detection target, call the robot vision perception module, and output the three-dimensional coordinates of the detection target; set the move function move(position), where position is the target position of movement, call the MoveIt! motion planner to generate a motion path; set the gripper_closing() function to control the gripper closing; set the gripper_opening() function to control the gripper to open to the maximum width.
[0112] Step 4.1: Design of voice instruction input layer
[0113] The Dolphin speech-to-text large model is deployed locally based on the PyTorch framework, and a sliding window mechanism is used to frame the real-time audio stream. When the end-of-speech flag is detected, the complete text sequence is output, and a timestamp and confidence score are automatically added. The final generated text sequence is encapsulated in JSON format, including original speech signal features, recognized text, and confidence metadata, which will then be transmitted to the large language model core module.
[0114] Step 4.2: Visualization of virtual-to-real motion process
[0115] A virtual model of the robotic arm is built in Gazebo, and the Franka Panda robotic arm is modeled in three dimensions using URDF language. According to the actual geometric parameters of the Franka panda robotic arm, the dimensions, mass distribution, and inertia matrix of each link are defined; the joint kinematics parameters are defined using XACRO macros, including joint type, motion range, and limit constraints. The physical properties are embedded in the SDF file. The visualization of the panda robotic arm in Gazebo is shown in Figure 9 The cross-platform data channel is established through the distributed communication framework of ROS, and the virtual-to-real motion process visualization process is shown in Figure 10The custom data collection node is deployed in the embedded system of the actual robot control cabinet, which calls the real-time driver API Franka Control Interface provided by the manufacturer of Franka robot to read the absolute angle values (unit: rad) of each joint encoder feedback at a sampling frequency of 1 kHz, and uses the Eigen library to solve the forward kinematics to obtain the 6-DOF pose data of the end effector in the base coordinate system. Then create a ROS publisher node, encapsulate the joint state data collected as sensor_msgs / JointState message type, fill the joint angle sequence to the position array in physical joint order, and pass the end pose through geometry_msgs / Pose message nesting, and broadcast in real time through the channel named / panda_joint_states topic in binary serialization, the message publishing frequency is strictly synchronized with the robot control cycle 500Hz. The virtual simulation environment is Ubuntu 20.04 system, and the URDF model isomorphic with the real robot is built by using the MoveIt! framework. When creating an independent subscription node, initialize the Python node instance by using rospy.init_node, and bind to the / panda_joint_states topic by using rospy.Subscriber function. In the callback processing, first perform CRC check and time alignment on the received JointState message, convert the end pose to TF transform tree in the base coordinate system by using tf2_ros.TransformBroadcaster, and then call robot_state_publisher to dynamically update the joint angle data to the JointStateController of the virtual model. To realize zero delay rendering, enable the Effort Display plug-in in the RViz visualization tool and configure the joint torque color mapping, and at the same time, set the use_sim_time parameter to false to ensure that the virtual model time reference is strictly synchronized with the actual physical system.
[0116] Step 4.3: Front-end web interface construction
[0117] Python language is used to call Gradio library for system development, and ollama and Dolphin modules are introduced to realize large model interaction and speech recognition function respectively. The AIbot function is defined to realize the large model interaction function, and the robot_AI_based_on_deepseek_8b model trained in step 1.3 is called. The dialog interaction mode is used to process the user input text. The function encapsulates the user input into a message structure of {“role”: user, “content”: text} format, sends it to the model inference service through the chat interface of the ollama library, obtains and returns the generated robot control instructions. The speech recognition function is realized through the ALbot function, and the pre-trained “small” speech model is loaded by using the Dolphin library. The function converts the audio file collected by the microphone into waveform data, and outputs the text recognition result after model processing, supporting accurate analysis of Chinese speech instructions. The main container is created by gr.Blocks(), and the speech interaction, text interaction and result display modules are integrated. The gr.Audio component is used in the speech interaction area to configure the microphone input, and the gr.Textbox is used to display the recognized text; the multi-line gr.Textbox is used in the text interaction area to realize manual input of instructions; the result display area presents the control instructions generated by the large model through the gr.Textbox with extended height. Each function area is configured with operation buttons, and the corresponding function functions are bound through the click event to realize speech recognition, instruction inference and interface cleaning operations. The web front-end interface is as shown in Figure 11
[0118] The above specific description further details the purpose, technical solutions and benefits of the application. It should be understood that the above description is only a specific embodiment of the application and is not used to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the application should be included in the protection scope of the application.
Claims
1. A natural language command control system and method for a robotic arm based on a large language model, characterized in that: The steps include: Step 1: Build an atomic skill library for robotic arms in assembly scenarios and parse natural language instructions. Decompose the common actions of robotic arms in assembly scenarios that transfer tools and parts to workers into atomic skills. Pre-train a large model to understand natural language instructions and convert them into atomic skill sequences. Step 2: Build the intelligent perception system of the robotic arm and use binocular cameras to realize target detection and obtain the three-dimensional coordinates of the target. Step 3: Use forward kinematics modeling to obtain the coordinate transformation matrix of the robot arm, and solve the robot arm's motion trajectory through inverse kinematics to achieve motion control and planning of the robot arm. Step 4: Build a human-computer interaction platform, add voice-to-text and virtual-to-reality modules, and build a front-end web interaction interface.
2. The large language model-based natural language command control system and method for a robotic arm according to claim 1, characterized in that: Step 1 is implemented as follows: Step 1.1: Construction of parameterized atomic skill library; The present invention constructs a parameterized atomic skill library for common instructions in the human-machine collaboration process in assembly scenarios, providing basic semantic units for mapping natural language instructions to robotic arm actions. First, based on the kinematic characteristics of the robotic arm and the assembly process requirements, the operation instructions are abstracted into four basic atomic action types: (1) Detect: Identify the three-dimensional coordinates of the target object based on the visual system; (2) Grasp: The end effector establishes physical contact with the target object and forms a stable constraint; (3) Move: The robotic arm performs trajectory planning and position changes in space; (4) Release: Release the constraint relationship between the end effector and the target object. The mathematical model of atomic skills is as follows: A={E,P,F} E: represents the environmental information input, including the environmental state parameters related to the atomic action in the assembly scenario, specifically the target object pose, the initial pose of the robot end effector, and the robot joint torque limits. P: represents action planning execution, which generates the atomic action execution strategy based on the environmental information E, including the robot motion path planning and control parameter configuration. F: represents action execution feedback, a state feedback parameter used to determine the successful completion of the atomic action. Its value is a discrete set {success, failure, in progress}. Step 1.2: Constructing action instruction semantic parsing dataset; First, based on the semantic parsing requirements of robotic arm operations, a set of instruction classifications covering industrial scenarios such as assembly and handling was constructed. For data collection, typical standard operating instructions and natural language descriptions of historical operations were integrated to maximize data coverage across a wide range of application scenarios. During the data structuring phase, a unified triplet annotation framework was designed. Each sample consists of an instruction (instruction), an operation description (input), and a control function sequence (output). Control functions are expressed using the parameterized atomic skills described in step 1.1, such as detect(obj), move(pos), gripper_closing(), and gripper_opening(), enabling the atomic decomposition of complex tasks. During the annotation process, the instruction text must be expressed in natural language, avoiding the use of specialized terminology as much as possible to make the instructions more accessible to everyday users. The operation description must clearly define the target object and accurately describe its spatial relationship to facilitate understanding of the task requirements. The control function sequence must correspond to the actual atomic actions that the robotic arm can perform, and the parameter values must satisfy kinematic constraints to ensure practical operability. To further enhance the model's generalization capabilities, data augmentation strategies were implemented at both the language and operational levels. At the language level, the expressive forms of commands were enriched through synonym replacement, sentence reorganization, and entity generalization. At the operational level, parameter perturbations, such as setting position offsets of ±5mm, changing the order of actions, and increasing the complexity of task scenarios, were employed to generate semantically equivalent but distinct variant samples. Furthermore, for commands with easily confusable semantics, comparative learning samples were designed, such as "Put A on B" and "Put B on A," to help the model better understand and distinguish different spatial relationships. Step 1.3: Fine-tune the distilled large model using low-rank adaptation techniques; In the model architecture transformation phase, knowledge distillation was used to convert a large-parameter, large-language model into a small-parameter model. LoRA technology was used to modify the attention mechanism of the Transformer layer. The weight update calculation method is as follows: W 1 =W 0 +B·A Where W 0 are the initial weight parameters; B is a dimensionality reduction matrix with dimensions d×r (d is the input dimension, r is the low-rank value); A is an dimensionality increase matrix with dimensions r×k (k is the output dimension). The specific operations of LoRA fine-tuning are as follows: all parameters of the original model are frozen, and only the adapter weights are trained; a LoRA module is added to the linear transformation layer between the query matrix (Query) and the value matrix (Value); and bilinear interpolation is used to dynamically fuse the original weights with the adapter weights. The parameters that need to be configured for the LoRA module include the rank and the scaling factor lora_alpha; the target fine-tuning module is locked to the "q_proj" and "v_proj" of the attention mechanism; and the lora_dropout rate needs to be adjusted based on the actual training results. The task type is set to causal language model. During model training, the accuracy training and gradient accumulation steps need to be set according to the hardware memory availability; the learning rate, warmup ratio, and weight decay parameters need to be adjusted based on the training log. In terms of training and verification strategies, we simultaneously monitor multiple indicators such as instruction parsing accuracy, action sequence generation F1 value, and parameter prediction error. We set an early stopping mechanism that terminates when there is no performance improvement for five consecutive verification cycles, and use an integrated learning strategy to fuse the three best-performing checkpoint models.
3. The large language model-based natural language command control system and method for a robotic arm according to claim 1, characterized in that: Step 2 is implemented as follows: Step 2.1: Target detection dataset construction; First, a camera is used to capture video of the assembly scene, ensuring a complete record of tool status and environmental information during the robotic arm's operation. The video capture covers various lighting conditions (e.g., natural light, artificial lighting), shooting angles (front view, side view, and top view), and tool placement to enhance the dataset's scene diversity. During the image capture phase, an interval sampling strategy is employed, acquiring basic samples at fixed intervals (e.g., one frame every five frames). The captured images are then dedistorted based on the intrinsic parameter matrix obtained through camera calibration. The LabelImg tool is used for manual annotation, adhering to the COCO dataset annotation specifications. Tool objects are selected in the image, and corresponding category labels (e.g., screwdriver, wrench, clamp, etc.) are entered. The tool's pose information (position coordinates and attitude quaternion) in the robotic arm's coordinate system is then recorded. During the dataset enhancement phase, the original annotated images are augmented with multi-dimensional data using methods such as geometric enhancement, photometric enhancement, contrast transformation, and Gaussian noise addition. The final dataset is divided into training set, validation set and test set in proportion. Each sample contains an annotated image file, a corresponding XML annotation file and a pose parameter text file. Step 2.2: YOLOv8 model training combined with transfer learning; We retain the universal features extracted by the pre-trained model's backbone network and use the weights from the YOLOv8 pre-trained model based on the COCO dataset as the initial training weights. Before training, we freeze the backbone network and train only the model's head network. This extracted feature is then used to complete the classification task of common tools and parts, a process known as transfer learning. The loss function comprehensively considers the coordinate loss for locating common tools and parts, as well as the recognition and classification loss. The tool and part classification loss calculation process is as follows: Where n is the number of tools and parts; σ(x i ) is the Sigmoid function; y i is a symbolic function. If the identified key component i belongs to the true category, the value is 1, otherwise it is 0. The coordinate loss calculation for common tool and component positioning uses CIoU loss, and the calculation process is as follows: Where, A and B are the coverage areas of the prediction box and the true value respectively; ρ is the diagonal distance of the minimum closure area between the prediction box and the true box; b is the center point of the prediction box for tool and part identification; b gt is the center point of the real frame; α is the weight coefficient; v is used to measure the consistency of the aspect ratio, and its calculation process is shown in the following formula: Where w and h are the width and height of the prediction box respectively; w gt With h gt are the width and height of the real frame respectively. Step 2.3: Building the binocular vision module; First, the binocular camera is calibrated, and the parameters of the left and right cameras are obtained using the Zhang calibration method. The specific steps of the Zhang calibration method are as follows: (1) Data collection: Prepare a checkerboard calibration plate with clear and identifiable corner points. Place the calibration plate in different directions and positions, and use the camera to take multiple photos from multiple angles. At least three photos are required, but in order to improve the accuracy of calibration, it is usually recommended to use more images. (2) Corner detection: Automatically detect the corner points of the checkerboard in each image. These corner points are pixels in the image coordinate system, which can be automatically found and accurately located by image processing software (such as OpenCV). (3) Establish coordinate relationship: Assuming that the size of each square of the checkerboard is known, the physical position of each corner point can be mapped to the world coordinate system. Because the calibration plate is placed on a plane, the Z coordinates of these corner points are usually set to 0. (4) Calculate camera parameters: Use a mathematical model to describe the imaging process of the camera, which involves internal and external parameters. The internal parameters include focal length, principal point coordinates and distortion coefficients. The external parameters include the rotation and translation matrix of the camera relative to the calibration plate in each image. The camera intrinsic parameter matrix is as follows: The distortion coefficient k matrix is shown below: k=[k1 k2 p1 p2 k3] On this basis, the homogeneous transformation matrix H is calculated, which includes the rotation matrix R and the translation vector T. The specific form of the homogeneous transformation matrix H is as follows: After calibration, the images taken by the left and right cameras are corrected. Based on the intrinsic and extrinsic matrix obtained by calibration, the correction transformation matrices R1, R2 and projection matrices P1, P2 are calculated, and then the original image is mapped to the corrected image plane so that the polar lines of the left and right images are collinear and parallel to the horizontal axis of the image, eliminating image distortion and unifying the coordinate system. In the disparity calculation link, the Semi-Global Block Matching (SGBM) algorithm is used for the corrected left and right images. It is necessary to set the matching cost calculation window size and the disparity search range, and adjust the uniqueness ratio threshold and smoothing parameter according to the generated disparity map. Finally, the disparity is converted into depth information according to the principle of triangulation. The depth calculation formula is Where f is the camera focal length, b is the binocular camera baseline distance, and d is the disparity value. By traversing each pixel in the disparity map and substituting the parameters into the calculated depth value, the corresponding depth image can be generated. Step 2.4: Obtain the target three-dimensional coordinates; Take the center point of the target obtained by the target detection model, use the depth vision module built in step 2.3 to obtain the coordinates of the target point in the camera reference system, and convert it into the coordinate system of the robot base through hand-eye calibration. Using the eye-in-hand solution, the solution formula for the hand-eye calibration parameters is as follows: In the formula is the transformation matrix from the fixture coordinate system to the robot base coordinate system; is the transformation matrix from the camera coordinate system to the fixture coordinate system; is the conversion matrix from the calibration plate coordinate system to the camera coordinate system. It can be solved by robot kinematics. It can be obtained through camera calibration, and the hand-eye calibration parameters can be solved using the Tsai method.
4. The large language model-based natural language command control system and method for a robotic arm according to claim 1, wherein: Step 3 is implemented as follows: Step 3.1: Build the kinematic model of the robotic arm; Forward kinematics is used to construct a transformation matrix that describes the joint positions and orientations. Using the Denavit-Hartenberg (DH) method, the transformation from one joint to the next can be expressed using four parameters: link length a, link twist α, link offset d, and joint angle θ. The URDF file describes the coordinate system of each joint and its constraints. Inverse kinematics is used to solve the nonlinear equations. Let the target pose be P = [x d ,y d ,z d ,R x ,R y ,R z ], it is necessary to solve the joint variable q=[q1,q2,…,q n ], by separating the position and attitude components, a nonlinear system of equations F(q) = 0 is constructed. The smoothest solution of the joint space trajectory is selected from multiple solutions, and a collision-free path is generated using OMPL. Step 3.2: Robotic Arm Atomic Motion Design Design four types of actions for the robotic arm's atomic skills described in step 1.1: detect, move, grasp, and release. For the detect action, input the target name to be detected, invoke the trained object detection model and deep vision module, and output the target's 3D coordinates relative to the robotic arm base. The hardware is a binocular camera. For the move action, input the target pose for the robotic arm's end effector and invoke the MoveIt! robotic arm motion planner to control the end effector to move to the specified position. For the grasp action, close the gripper. The hardware is the robotic arm's end effector. For the release action, open the gripper. The hardware is the robotic arm's end effector.
5. The large language model-based natural language command control system and method for a robotic arm according to claim 1, characterized in that: Step 4 is implemented as follows: Step 4.1: Design the voice command input layer; The voice signal is transmitted to the main control unit via the USB interface, and acoustic features are extracted by the locally deployed Dolphin speech-to-text model. The model uses Mel-frequency cepstral coefficients (MFCC) combined with a logarithmic Mel filter bank to convert the time domain signal into a feature vector sequence and generate an acoustic feature matrix. In the semantic conversion stage, the acoustic feature matrix is input into the encoder module of the Dolphin model, and feature extraction and context modeling are performed through the Transformer structure. The decoder uses an attention mechanism, combined with pre-trained language model parameters, to generate the text sequence with the highest probability. Finally, it is automatically converted to full-width characters, unified punctuation format, and metadata such as timestamp and recognition confidence are added to form a structured input that can be parsed by the downstream natural language processing module. Step 4.2: Visualize the motion process by using virtual objects to reflect the real ones; First, build a virtual model of the robotic arm in Gazebo, and use the URDF (Unified Robot Description Format) language to perform three-dimensional modeling of the robotic arm. Based on the actual geometric parameters of the robotic arm, the dimensions, mass distribution, and inertia matrix of each link are precisely defined. XACRO macros are used to define the kinematic parameters of the joints, including joint type (revolute joint, prismatic joint), range of motion, and limit constraints. Physical properties are embedded in an SDF (Simulation Description Format) file, the friction coefficient between links and the collision detection threshold are set, and the simulation time step is configured to ensure that the dynamic characteristics of the virtual model are highly consistent with the actual system. A data communication link is then established based on the ROS (Robot Operating System) architecture. A data acquisition node is deployed in the actual robotic arm control cabinet. This data is collected through a real-time driver interface, and a ROS topic is established to publish the collected data in binary format in real time. A ROS subscription node is created on the virtual model side, and a callback function mechanism is implemented to trigger a virtual robotic arm status update upon receiving data from the actual robotic arm. The actual robotic arm joints and end-effector poses are mapped to the virtual robotic arm. Step 4.3: Build the front-end web interface; The interface function integration is realized by calling Gradio library through Python language. The interface design adopts a three-column layout, which mainly includes three modules: text direct input, speech to text and large model reply.<gr.Textbox> The component creates an input box, which is connected to the large model reply module through<gr.Button> When the trigger button is clicked, the input text is passed to the deployed large language model inference interface; the speech-to-text module is bound to the click event of<gr.Audio> The component creates a microphone input interface, sets the real-time recording function and opens the specified audio data file path format. When the window button is triggered, the speech recognition model dolphin will be called to convert the audio file into text and output it to<gr.Textbox> The component displays the recognition results, and the buttons of the text box window are also displayed through<gr.Button> The click event is bound to the large model reply module.
Citation Information
Cited By
Autonomous planning method based on large language model
CN121179443A
Autonomous planning method based on large language model
CN121179443B
Personal intelligent multi-source data quality evaluation and verification method, device, medium and product
CN121188440A
Double-arm robot flat cable buckling method based on multi-modal sensing
CN121777159A
Robot URDF model construction method and device, electronic equipment and storage medium
CN121934817A