Intelligent mobile grabbing robot control method and system based on modular architecture
Through a modular architecture combining visual perception, natural language interaction and intelligent decision-making of large language models, the problems of complex interaction and limited decision-making capabilities of industrial robots are solved, and efficient and flexible industrial task execution is achieved.
Patent Information
- Application Number
- CN202510476944.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
The existing industrial robots have complex interaction methods, limited decision-making capabilities, insufficient adaptability of the system architecture, and it is difficult to cope with dynamically changing task requirements.
It adopts a modular architecture, combining visual perception, natural language interaction and intelligent decision-making of large language models, and through the cloud-edge-native collaborative system, decoupling of object recognition, natural language processing and action control is achieved.
It improves the naturalness and intelligent decision-making capabilities of robot interaction, enhances the flexibility and adaptability of the system, and meets the real-time needs of industrial industries.
Smart Images

Figure CN120382482A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent robots, and more specifically, relates to a method and system for an intelligent mobile grasping robot based on a modular "vision-language-action" architecture. Background Art
[0002] Industrial production is transforming from traditional mechanization and automation to intelligence and flexibility. As one of the key technologies for realizing intelligent manufacturing, intelligent robots play a vital role in industrial production.
[0003] Therefore, industrial robots urgently need to evolve towards natural human-machine interaction, rapid perception, and intelligent decision-making. However, there are still many limitations in the development of industrial robots:
[0004] 1) Complex interaction methods: Although traditional physical buttons or teach pendant interactions can achieve precise operations, they have high requirements and require pre-arranged or recorded action sequences. They are difficult to cope with dynamically changing task requirements, have a high operating threshold, and lack flexibility.
[0005] 2) Limited decision-making capabilities: The decision-making mechanism based on the rule engine relies on a large number of predefined templates. When faced with instructions in natural language or vague expressions, it is easy to produce semantic ambiguity, leading to decision failure and poor intelligence.
[0006] 3) Insufficient adaptability of system architecture: Existing intelligent robot systems mostly adopt an end-to-end integrated design with highly coupled perception, decision-making and execution. This not only requires a large amount of resources in model construction and deployment, but also requires major modifications to the existing system during upgrades and expansions, resulting in insufficient adaptability.
[0007] In recent years, the rise of large language models has opened up new possibilities for natural language interaction and decision-making. With their powerful understanding and generation capabilities, large language models enable natural human-robot conversations, process complex natural language instructions, and generate dynamic decisions. However, how to efficiently integrate large language models with robot control systems remains an understudied problem. Summary of the Invention
[0008] In response to the above-mentioned defects or improvement needs of the prior art, the present invention provides a control method and system for an intelligent mobile grasping robot based on a modular architecture, aiming to solve problems such as complex interactions and limited decision-making capabilities of traditional robots.
[0009] To achieve the above objectives, according to one aspect of the present invention, a control method for an intelligent mobile grasping robot based on a modular architecture is provided, comprising the following steps:
[0010] S1: Real-time collect the image information of the target object, process the image data in real time, identify the target object and output the position and category of the target object, obtain the depth information of the detected target object, and provide the basic detection result for subsequent decision-making;
[0011] S2: Send natural language text through the user terminal, the cloud server receives the text information, uses the large language model for natural language interaction to generate language reply information and decompose actions to generate decomposed action instructions; return the language reply information to the user terminal through the cloud server; send the decomposed action instructions to the edge computing device, combine the decomposed action instructions with the depth information of the target object obtained in step S1 to obtain the final action instruction sequence, and realize intelligent decision-making;
[0012] S3: Send the action instruction sequence generated in step S2 to the IPC, and the IPC controls the mobile grasping robot to complete related actions of moving and grasping to ensure the accuracy and stability of the actions.
[0013] Preferably, step S1 specifically includes the following steps:
[0014] S11: Collect images through web crawlers and camera shooting methods, manually annotate each image, and generate annotation files for object detection;
[0015] S12: Perform preprocessing and data augmentation operations on the original images generated in step S11, expand the dataset, and divide the dataset into a training set, a validation set, and a test set;
[0016] S13: Use a depth camera to collect the RGB image and depth map of the target object in real time;
[0017] S14: Select the improved YOLOv10 model and train the training set obtained in step S12 on the cloud server. After training is completed, obtain the model file; quantize the model file to obtain the quantized model, deploy it to the edge computing device to achieve local operation; at the same time, based on the quantized model, detect the RGB image collected in step S13 in real time to obtain the bounding box and classification label of the target object, and combine the depth map collected in step S13 to obtain the depth value of the object, and output the detection result.
[0018] Preferably, the specific steps of the preprocessing and data augmentation operations in step S12 are: automatically remove the background from some images, extract the target element map, and generate diverse synthetic images and corresponding annotation files by using random transformation and background synthesis to enhance the diversity and robustness of the dataset.
[0019] Preferably, the improved YOLOv10 model in step S14 includes a ResSPPF module, an OGCA module, and a C2f_ParNetAttention module.
[0020] Among them, the ResSPPF module optimizes the multi-scale feature extraction ability by introducing a cascaded residual structure, and optimizes the detection performance of small targets and multi-scale targets.
[0021] The OGCA module improves the model's adaptability to complex scenarios by introducing a global feature fusion mechanism and an occlusion-aware branch.
[0022] The C2f_ParNetAttention module replaces the labeled bottleneck block with an improved bottleneck block based on ParNetAttention, enhancing the spatial perception of features and the modeling ability of inter-channel relationships, thereby enhancing the detection performance.
[0023] Preferably, in step S14, the quantization method is to convert the model weights from floating-point numbers to integers, perform format conversion and quantization processing on the trained model, and generate a quantization model suitable for NPU operation.
[0024] Preferably, step S2 specifically includes the following steps:
[0025] S21: The user terminal inputs natural language text information through the public platform, and the message interface of the public platform sends the instruction to the cloud server, which receives and preprocesses it.
[0026] S22: Call the large language model to perform natural language processing on the text information obtained in step S21, utilize the semantic understanding ability of the large language model, and generate language reply information based on the CoT and few-shot prompting methods, as well as decompose actions to generate an executable action instruction sequence.
[0027] S23: Call the passive reply interface of the public platform to return the language reply information generated in step S22 to the user terminal to achieve a two-way interaction experience.
[0028] S24: Send the action instruction sequence generated in step S22 to the edge computing device through the cloud server, combine the decomposed action instructions with the target category, bounding box, and depth information of the target object obtained in step S1, and construct a complete action instruction sequence.
[0029] S24: For the problem that no target object is detected in step S1 and the instruction sequence parameters generated in step S24 are incorrect, send them to the cloud server through the edge computing device, and the cloud server calls the template message interface of the public platform to actively send them to the user terminal.
[0030] Preferably, step S3 specifically includes the following steps:
[0031] S31: Write a control program to implement path planning and motion control of the mobile grasping robot.
[0032] S32: Transmit the action instruction sequence generated in step S2 from the edge computing device to the IPC, and the IPC controls the mobile grasping robot to complete the relevant actions of movement and grasping to ensure the accuracy and stability of the actions.
[0033] To achieve the above object, according to another aspect of the present invention, there is provided an intelligent mobile grasping robot control system based on a modular architecture, including:
[0034] A vision module that collects image information of the target object in real time, processes the image data in real time, identifies the target object and outputs the position and category of the target object, obtains the depth information of the detected target object, and provides a basic detection result for subsequent decision-making;
[0035] A language module. The user terminal sends natural language text, and the cloud server receives the text information, uses a large language model for natural language interaction to generate information for language reply and decomposes actions to generate decomposed action instructions; returns the information for language reply to the user terminal through the cloud server; sends the decomposed action instructions to the edge computing device, combines the decomposed action instructions with the depth information of the target object obtained by the vision module to obtain a final action instruction sequence, and realizes intelligent decision-making;
[0036] An action module that sends the action instruction sequence generated by the language module to the IPC, and the IPC controls the mobile grasping robot to complete the relevant actions of movement and grasping to ensure the accuracy and stability of the actions.
[0037] Preferably, the vision module includes:
[0038] An image acquisition unit that collects high-quality images through web crawlers and camera shooting means, and manually annotates each image to generate an annotation file for object detection;
[0039] An image preprocessing unit that preprocesses and performs data augmentation operations on the original images generated by the image acquisition unit, expands the data set, and divides the data set into a training set, a validation set, and a test set;
[0040] A depth information acquisition unit that uses a depth camera to collect RGB images and depth maps of the target object in real time;
[0041] A model training unit that selects an improved YOLOv10 model and trains the model on a cloud server, sets the model training parameters, and obtains a model file after training; quantifies the model file to obtain a quantized model, deploys it to the edge computing device to achieve local operation; at the same time, based on the quantized model, it detects RGB images in real time, obtains the bounding box and classification label of the target object, and combines the depth map to obtain the depth value of the object, and outputs the detection result.
[0042] Preferably, the language module includes:
[0043] An input unit. The user inputs natural language text information through a public platform. The message interface of the public platform sends instructions to the cloud server, and the cloud server receives and preprocesses them.
[0044] A processing unit that calls a large language model to perform natural language processing on the text information. Utilizing the semantic understanding ability of the large language model, it performs natural language interaction to generate language reply information based on the CoT and few-shot prompting methods, and decomposes actions to generate an executable action instruction sequence.
[0045] An interaction unit that calls the passive reply interface of the public platform to return the language reply information obtained by the processing unit to the user, realizing a two-way interaction experience.
[0046] A decomposed action unit that sends the action instruction sequence obtained by the processing unit from the cloud server to the edge computing device, combines the decomposed action instructions with the target category, bounding box, and depth information of the target object obtained by the vision module to construct a complete action instruction sequence, and transmits it to the action module.
[0047] An error handling unit. For the problems that the vision module fails to detect the target object and the instruction sequence parameters generated by the decomposed action unit are incorrect, it will be sent to the cloud server through the edge computing device, and the cloud server will actively send it to the user through the template message interface of the public platform.
[0048] Preferably, the CoT and few-shot prompting techniques are to construct a complete prompt word by designing system prompts, task instructions, few-shot examples, and their thought chains.
[0049] Preferably, the action module includes:
[0050] A control unit that writes a control program to implement path planning and motion control of the mobile grasping robot.
[0051] A communication unit. The action instruction sequence generated by the language module is transmitted from the edge computing device to the IPC, and the IPC controls the mobile grasping robot to complete mobile and grasping related actions to ensure the accuracy and stability of the actions.
[0052] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0053] Through modular design, the method and system of the present invention decouple and optimize the integration of visual perception, natural language interaction, large language model intelligent decision-making, and action control execution functions. It can not only identify the category, bounding box, and depth information of objects, but also meet the industrial real-time requirements, while solving the problems of complex traditional robot interaction, limited decision-making ability, complex construction of embodied intelligent robots, and poor flexibility. Description of the Drawings
[0054] Figure 1 It is a schematic diagram of the overall system architecture provided by an embodiment of the present invention;
[0055] Figure 2 It is a schematic diagram of the visual module structure provided by an embodiment of the present invention;
[0056] Figure 3 It is a schematic diagram of the improved YOLOv10 network structure provided by an embodiment of the present invention;
[0057] Figure 4 It is a schematic diagram of the language module structure provided by an embodiment of the present invention;
[0058] Figure 5 It is a schematic diagram of the action module structure provided by an embodiment of the present invention. Detailed Embodiments
[0059] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0060] As Figure 1 shown, in view of the problems of complex traditional robot interaction, limited intelligent decision-making ability, complex construction of embodied intelligent robots, and poor flexibility, the present invention provides an intelligent mobile grasping robot control method based on a modular vision-language-action architecture. By decoupling and synergistically integrating visual perception, natural language interaction, large language model intelligent decision-making, and action control execution, the intelligent mobile grasping task in an industrial scenario is realized. This architecture fully combines cloud computing capabilities, edge real-time performance, and local stability to form an efficient cloud-edge-device collaborative system.
[0061] The method includes the following steps:
[0062] S1: Real-time collect the image information of the target object, real-time process the image data, identify the target object, and output the position and category of the target object, and obtain the depth information of the detected target object to provide a basic detection result for subsequent decision-making;
[0063] Step S1 specifically includes the following steps:
[0064] S11: Collect images through web crawlers and camera shooting, manually annotate each image, and generate annotation files for object detection;
[0065] S12: Perform preprocessing and data augmentation operations on the original images generated in step S11, expand the dataset, and divide the dataset into a training set, a validation set, and a test set;
[0066] The specific steps of the preprocessing and data augmentation operations are as follows: Automatically remove the background from some images, extract the target element map, and generate diverse synthetic images and corresponding annotation files by using random transformation and background synthesis to enhance the diversity and robustness of the dataset;
[0067] S13: Use a depth camera to collect RGB images and depth maps of the target object in real time;
[0068] S14: Select the improved YOLOv10 model and train the training set obtained in step S12 on the cloud server. After training is completed, obtain the model file; Quantize the model file to obtain the quantized model, deploy it to the edge computing device to achieve local operation; At the same time, based on the quantized model, detect the RGB images collected in step S13 in real time to obtain the bounding boxes and classification labels of the target objects, and combine the depth maps collected in step S13 to obtain the depth values of the objects, and output the detection results.
[0069] The improved YOLOv10 model in step S14 includes a ResSPPF module, an OGCA module, and a C2f_ParNetAttention module.
[0070] Among them, the ResSPPF module optimizes the multi-scale feature extraction ability by introducing a cascaded residual structure, and optimizes the detection performance of small targets and multi-scale targets;
[0071] The OGCA module improves the model's adaptability to complex scenes by introducing a global feature fusion mechanism and an occlusion perception branch;
[0072] The C2f_ParNetAttention module enhances the detection performance by replacing the standard bottleneck block with a bottleneck block improved based on ParNetAttention to enhance the spatial perception of features and the modeling ability of the relationship between channels.
[0073] The quantization method in step S14 is to convert the model weights from floating-point numbers to integers, perform format conversion and quantization processing on the trained model, and generate a quantized model suitable for NPU operation.
[0074] S2: The client sends natural language text, the cloud server receives the text information, and uses a large language model for natural language interaction to generate language reply information and perform action decomposition to generate decomposed action instructions; the language reply information is returned to the client through the cloud server, and the decomposed action instructions are sent to the edge computing device, so that the decomposed action instructions are combined with the target object depth information obtained in step S1 to obtain the final action instruction sequence and achieve intelligent decision-making;
[0075] Step S2 specifically includes the following steps:
[0076] S21: The client inputs natural language text information through the public platform, and the message interface of the public platform sends the instruction to the cloud server, and the cloud server receives and preprocesses it;
[0077] S22: Invoke the large language model to perform natural language processing on the text information obtained in step S21, and use the semantic understanding ability of the large language model to perform natural language interaction based on the CoT and few-shot prompting methods to generate language reply information and perform action decomposition to generate an executable action instruction sequence;
[0078] S23: Invoke the passive reply interface of the public platform to return the language reply information generated in step S22 to the client to achieve a two-way interaction experience;
[0079] S24: Send the action instruction sequence generated in step S22 to the edge computing device through the cloud server, so that the decomposed action instructions are combined with the target category, bounding box, and depth information of the target object obtained in step S1 to construct a complete action instruction sequence and transfer it to the action module;
[0080] S24: For the problems that no target object is detected in step S1 and the instruction sequence parameters generated in step S24 are incorrect, send them to the cloud server through the edge computing device, and the cloud server actively sends them to the client by invoking the template message interface of the public platform.
[0081] S3: Send the action instruction sequence generated in step S2 to the IPC, and the IPC controls the mobile grasping robot to complete the relevant actions of moving and grasping to ensure the accuracy and stability of the actions.
[0082] Step S3 specifically includes the following steps:
[0083] S31: Write a control program to implement path planning and motion control of the mobile grasping robot;
[0084] S32: Transmit the action instruction sequence generated in step S2 from the edge computing device to the IPC, and the IPC controls the mobile grasping robot to complete the relevant actions of moving and grasping to ensure the accuracy and stability of the actions.
[0085] Meanwhile, the present invention provides an intelligent mobile grasping robot control system based on a modular vision-language-action architecture, including:
[0086] A vision module that real-time collects image information of the target object, processes the image data in real-time, identifies the target object, and outputs the position and category of the target object, obtains the depth information of the detected target object, and provides a basic detection result for subsequent decision-making;
[0087] A language module where the user terminal sends natural language text, the cloud server receives the text information, uses a large language model for natural language interaction to generate information for language responses and decomposes actions to generate decomposed action instructions; returns the information for language responses to the user terminal through the cloud server; sends the decomposed action instructions to the edge computing device, combines the decomposed action instructions with the depth information of the target object obtained by the vision module to obtain a final action instruction sequence, and realizes intelligent decision-making;
[0088] An action module that sends the action instruction sequence generated by the language module to the IPC, and the IPC controls the mobile grasping robot to complete mobile and grasping-related actions to ensure the accuracy and stability of the actions.
[0089] The vision module includes:
[0090] An image acquisition unit that collects high-quality images through web crawlers and camera shooting means, and each image is manually annotated using the LabelImg tool to generate an annotation file for object detection in the format of a YOLO-compatible txt file;
[0091] An image preprocessing unit that preprocesses and performs data augmentation operations on the original images generated by the image acquisition unit, expands the dataset, and divides the dataset into a training set, a validation set, and a test set in a ratio of 8:1:1;
[0092] A depth information acquisition unit that uses a depth camera to real-time collect the RGB image and depth map of the target object;
[0093] A model training unit that selects an improved YOLOv10 model and trains the model on the cloud server using the PyTorch framework, sets the model training parameters, and obtains a model file after training; quantizes the model file to obtain a quantized model, deploys it to the edge computing device to achieve local operation; simultaneously, based on the quantized model, real-time detects the RGB image, obtains the bounding box and classification label of the target object, and combines the depth map to obtain the depth value of the object, and outputs the detection result in JSON format.
[0094] Post-Training Quantization (PTQ) quantizes the model after training is completed. Based on PTQ technology, a small amount of calibration data is used to guide the quantization process. There is no need to retrain the original model. It can be directly converted into a fixed-point computing network with only a small number of hyperparameters. The quantization method is to convert the model weights from 32-bit floating-point numbers to 8-bit integers based on PTQ technology, convert the trained pt format model to an ONNX format model, and perform quantization processing on the ONNX model based on the RKNNToolkit2 development kit to generate a quantization model suitable for NPU operation.
[0095] The specific steps of the preprocessing and data augmentation operations are as follows: Use the Pixian.AI tool to automatically remove the background from some images and extract the target element map. Generate diverse synthetic images and corresponding annotation files by using random transformations (such as rotation, scaling, brightness adjustment, etc.) and background synthesis (such as industrial scenes, office scenes, natural environments, etc.) to enhance the diversity and robustness of the dataset.
[0096] The depth information acquisition unit uses a depth camera as the image acquisition device, which has the ability to acquire RGB images and depth information, and provides a dual-mode data stream of RGB images and depth maps.
[0097] In the model training unit, the basic model selects the improved YOLOv10; the improved YOLOv10 includes the ResSPPF module, the OGCA module, and the C2f_ParNetAttention module.
[0098] The ResSPPF (Residual Spatial Pyramid Pooling Fast) module optimizes the multi-scale feature extraction ability by introducing a cascaded residual structure, as Figure 3 shown. The input features of the ResSPPF module first generate initial features through a 1×1 convolution where N, C, H, and W are the batch size, number of channels, height, and width respectively, to retain the original information for subsequent operations, and store the initial features as the residual branch residual. Different from the SPPF module, the ResSPPF module does not halve the number of channels and retains all input channel information. Subsequently, the module uses a max pooling layer for multi-scale pooling operations and introduces a cascaded residual connection, specifically: after each pooling operation, add the pooling result to the current feature element by element, as shown in the formula:
[0099] x = MaxPool(x) + x
[0100] Among them, MaxPool is the maximum pooling layer, which generates multi-scale features through three iterations. Four groups of features [x0, x1, x2, x3] including the original features are concatenated along the channel dimension to generate a feature vector with the shape of where x0 is the original feature, and x1, x2, and x3 are the results of three pooling operations. The number of channels is adjusted back to the target value through 1×1 convolution to generate fused features. Finally, a global residual connection is introduced to add the fused features to the initial features to further retain the original information and generate the final output, as shown in the formula:
[0101] x = x + residual
[0102] Based on retaining the efficient multi-scale feature extraction ability of the SPPF module, the ResSPPF module further optimizes the detection performance of small targets and multi-scale targets. The cascaded residual connection alleviates the problem of small target spatial resolution compression caused by consecutive pooling by dynamically fusing the current pooling features with the features of the previous stage in each pooling operation, enhances the expression ability of the edges and texture details of small targets, and simultaneously optimizes the feature hierarchy consistency of different scales through iteration. The global residual connection further compensates for the potential weakening of small target details during the pooling process by adding the fused features to the initial features, and can improve the detection accuracy and localization accuracy of small targets. ResSPPF does not follow the design of halving the number of channels in the SPPF module and retains all input channel information. At the same time, the cascaded residual connection and the global residual connection are only implemented through element-wise addition operations without introducing additional parameters.
[0103] To overcome the deficiencies of the CA module in two-dimensional spatial relationship modeling and occlusion scenarios, an occlusion-aware global coordinate attention module (OGCA) based on the improvement of the CA module is proposed. By introducing a global feature fusion mechanism and an occlusion-aware branch, the adaptability of the model to complex scenarios is further improved.
[0104] The input feature map of the OGCA module is where N, C, H, and W are the batch size, number of channels, height, and width respectively. First, adaptive average pooling operations are performed along the height and width dimensions respectively to generate the height-direction feature and the width-direction feature Then the height-direction feature x h and the width-direction feature x w are concatenated into through the Concat operation and the channels are adjusted through 1×1 convolution. To enhance the perception ability of global context information, the OGCA module introduces a global pooling operation to generate the global feature and expands it to be the same as x cThe dimension of feature alignment is the same as x c They are added together to fuse the global context information, enabling the generation of attention weights to comprehensively consider the global and local spatial dependencies. The fused features are processed by batch normalization and activation functions to provide non-linear expression capabilities. The processed features are split along the height and width dimensions into and and respectively generate attention weights a h in the height and width directions through 1×1 convolutions w . To address the problems of instance overlap and occlusion, the OGCA module designs an occlusion-aware branch that reduces the input features to a single channel through 1×1 convolutions to generate a preliminary occlusion map and performs local refinement through 3×3 convolutions to further enhance the perception ability of overlapping regions, and normalizes it to the range of [0,1] through the Sigmoid function to obtain the occlusion probability. Subsequently, the occlusion map is adaptively averaged along the height and width dimensions to generate the occlusion distribution in the height direction and the occlusion distribution in the width direction . After the attention weights are normalized by the Sigmoid function, they are combined with the occlusion distribution to adjust the attention weights, reducing the attention weights in the overlapping regions and enhancing the attention to non-overlapping target regions, as shown in the formula:
[0105] a h = Sigmoid(a h )·(1-occ h )
[0106] a w = Sigmoid(a w )·(1-occ w )
[0107] Finally, the OGCA module applies the adjusted attention weights a h and a w to the input feature map x, and generates the output features through element-wise multiplication operations, as shown in the formula:
[0108] x out = x·a h ·a w
[0109] The OGCA module significantly improves the accuracy of feature representation and the robustness of the model through the synergy of the global feature fusion mechanism and the occlusion-aware branch. The global pooling operation enhances the model's ability to model the two-dimensional space by integrating height and width information, making the generation of attention weights more globally semantically based; while the occlusion-aware branch accurately suppresses the influence of overlapping regions by dynamically generating occlusion distributions, ensuring that the model focuses on the unoccluded target regions. Compared with the CA module, it not only retains the lightweight design advantage but also significantly improves the accuracy of feature representation through the introduction of global context information and the ability to perceive occlusion regions, taking into account the complex scene adaptability and computational efficiency while enhancing the spatial position perception ability.
[0110] To overcome the deficiencies of the C2f module in multi-scale feature extraction, the C2f_PNA module is proposed. By replacing the standard bottleneck block with a bottleneck block improved based on ParNetAttention (ParNet Block Attention), it further enhances the spatial perception and inter-channel relationship modeling ability of features, thereby improving the detection performance. The input features of the C2f_PNA module first generate initial features through 1×1 convolution where N, C, H, and W are the batch size, number of channels, height, and width respectively. The initial features are divided into two parts along the channel dimension, and one part of the features is directly retained, and the other part of the features is processed through n Bottleneck_PNA modules. In the Bottleneck_PNA module, the input features first adjust the number of channels through 1×1 convolution, then extract local spatial features through 3×3 convolution, and are then enhanced through the ParNetAttention module. ParNetAttention adopts a multi-branch parallel structure. The first branch generates channel-enhanced features through 1×1 convolution The second branch extracts local spatial features through 3×3 convolution The third branch uses the SSE mechanism to generate global features through adaptive average pooling, and through 1×1 convolution and the Sigmoid function to generate spatial attention weights, which are multiplied by the original features to obtain [x 21 ,x 22 ,x 23 After the three-branch features are fused, if the residual connection is enabled, they are added to the input feature x2 to generate the output of the bottleneck block. Finally, the number of channels is adjusted to the target number of channels through 1×1 convolution to output the final features
[0111] The C2f_PNA module significantly enhances the expressive ability of the spatial context and channel relationships of features by introducing ParNetAttention, enabling it to capture more comprehensive feature information than C2f in complex scenarios. The multi-branch design of ParNetAttention combines 1×1 and 3×3 convolutions to extract multi-level features, and uses the SSE mechanism to enhance the global perception ability, making C2f_PNA show stronger robustness in dense or small object detection. At the same time, its lightweight structure optimizes the computational efficiency, ensuring its applicability in real-time detection tasks. Therefore, the C2f_PNA module not only integrates the cross-stage fusion and multi-branch characteristics of C2f, maintaining the ability to capture multi-scale objects, but also achieves a balance between feature expressive ability and computational efficiency through the introduction of ParNetAttention, providing an efficient and robust solution for object detection in complex scenarios.
[0112] That is, the improved YOLOv10 significantly enhances the detection accuracy, robustness and adaptability to complex scenarios by integrating the ResSPPF module, the OGCA module and the C2f_ParNetAttention module, while maintaining high efficiency in real-time. The ResSPPF module enhances the expressive ability of the edges and texture details of small objects, improves the hierarchical consistency of multi-scale features and detection accuracy, especially showing higher accuracy in small object localization and recognition. The OGCA module improves the accuracy of feature expression and the model's perception ability of two-dimensional spatial relationships, effectively suppressing the interference of occluded areas and enhancing the focusing ability on unoccluded objects, thus showing stronger robustness in dense or overlapping object scenarios. The C2f_ParNetAttention module significantly improves the feature capture ability in complex scenarios through more comprehensive spatial context and channel relationship modeling, especially showing excellent performance in dense and small object detection, while optimizing the computational efficiency to meet real-time requirements. The synergistic effect of the three modules enables YOLOv10 to achieve a comprehensive breakthrough in small object detection accuracy, multi-scale object capture ability and adaptability to complex scenarios, providing an efficient, robust and high-precision solution for industrial-level real-time detection tasks.
[0113] The language module includes:
[0114] An input unit. The user side (WeChat) inputs natural language text information through a public platform (such as the WeChat official account platform), and the message interface of the public platform sends the instructions to the cloud server in the form of an HTTPS request. The cloud server deploys the FLASK framework for reception and preprocessing;
[0115] The processing unit calls the API of the large language model to perform natural language processing on the text information. Utilizing the semantic understanding ability of the large language model, it conducts natural language interaction based on the CoT and few-shot prompting methods to generate language reply information, and decomposes actions to generate an executable action instruction sequence.
[0116] The interaction unit calls the passive reply interface of the public platform (such as WeChat official account) in JSON format with the language reply information obtained by the processing unit and returns it to the user, realizing a two-way interaction experience.
[0117] The decomposed action unit sends the action instruction sequence obtained by the processing unit to the edge computing device from the cloud server based on the Scoket object of the TCP protocol, combines the decomposed action instructions with the target category, bounding box, and depth information of the target object obtained by the visual module to construct a complete action instruction sequence, and transfers it to the action module.
[0118] CoT and few-shot prompting techniques construct a complete prompt by designing system prompts, task instructions, few-shot examples, and their thought chains. Among them, the designed system prompt, such as clearly identifying the identity as a robot motion planning assistant, focuses on the mobile grasping task in the industrial scenario; provides environmental information and assumes that there may be dynamic changes in the scenario. Task instructions include robot movement, robotic arm movement, target detection, gripper opening and closing, etc. Few-shot examples include providing diverse examples covering scenarios such as conversations without action intent, ambiguous instructions, single tasks, and complex tasks, guiding the model to learn the action decomposition logic and output specifications. The thought chain designs a step-by-step reasoning process: 1. Identify the user's intent and determine whether there are action instructions; 2. Extract core information such as actions, target objects, and position parameters, and clarify if not clear; 3. Generate an action sequence by combining the principles of safety and dependence, generate a natural language reply and an action sequence, explain the plan and ask for confirmation.
[0119] The language module also includes an error handling unit. For problems such as the visual module not detecting the target object and incorrect instruction sequence parameters generated by the decomposed action unit, the problem situation is sent to the cloud server through the edge computing device based on the Scoket object of the TCP protocol, and the cloud server actively sends it to the user side by calling the template message interface of the public platform.
[0120] The action module includes:
[0121] The control unit writes a control program based on the HPAC programming platform to achieve path planning and motion control of the mobile grasping robot.
[0122] The communication unit transmits the action instruction sequence generated by the language module from the edge computing device to the IPC via the Modbus protocol. The Modbus protocol adopts the RTU mode to ensure the stability and real-time performance of communication.
[0123] The Modbus protocol is a simple bus protocol used in industry for communication. It has the advantages of excellent compatibility, high reliability, high-speed data transmission, high scalability, and simplicity of use.
[0124] Specifically, the present invention has the following beneficial effects:
[0125] (1) The vision module of this preferred embodiment uses a depth camera to collect RGB images and depth information, and combines a pre-trained object detection model and post-training quantization technology to achieve high-precision recognition and positioning of target objects. The system can not only identify the category, bounding box, and depth information of objects, but also the local deployment of edge computing devices can meet the real-time requirements of industry.
[0126] (2) Through the combination of the WeChat official account platform and the large language model in this preferred embodiment, a natural language interaction method without training is achieved. Users do not need to learn complex operation interfaces or predefined instructions, and can communicate with the robot bidirectionally only through daily language.
[0127] (3) Utilizing the powerful semantic understanding ability of the large language model and combining the chain of thought and few-shot prompting techniques in this preferred embodiment significantly improves the naturalness of human-computer interaction, the understanding of natural language instructions, and the task decomposition ability. With CoT and few-shot prompting, the action decomposition ability of the large language model and the adaptability to diverse instructions are enhanced, not only improving the intelligence level of the system, but also enhancing the robustness and flexibility of unstructured language input.
[0128] (4) In this preferred embodiment, functions such as visual perception, natural language interaction, intelligent decision-making of the large language model, and action execution are decoupled through a modular architecture. Each module operates independently and collaborates through a standardized interface in JSON format. This design greatly improves the adaptability and scalability of the system, providing convenience for the rapid deployment and iteration of different industrial scenarios. The modular architecture adopts a cloud-edge-device collaboration mechanism, and each module runs on a cloud server, an edge computing device, and an IPC respectively, making full use of the computing power of the cloud, the real-time performance of the edge computing device, and the stability of local hardware.
[0129] The following further describes the present invention with specific embodiments.
[0130] Example 1: Take the screwdriver on the table to the table on the right
[0131] The vision module, as Figure 2 shown, specifically is:
[0132] Collect high-quality images of objects such as screwdrivers (not limited to screwdrivers) through web crawlers and camera shooting. Each image is manually annotated using the LabelImg tool to generate annotation files for object detection in the format of YOLO-compatible txt files.
[0133] Perform preprocessing and data augmentation operations on the original images to expand the dataset. The dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0134] Select the improved YOLOv10 as the basic object detection model and train the model using the PyTorch framework on a cloud server. Set training parameters such as batch size, initial learning rate, optimizer, and number of training epochs. After training is completed, obtain the model file.
[0135] Based on the PTQ technology, convert the model weights from 32-bit floating-point numbers to 8-bit integers. Convert the trained pt format model to the ONNX format model, and perform quantization processing on the ONNX model based on the RKNN Toolkit2 development kit to generate an RKNN model suitable for running on the NPU, reducing the storage space of the model and improving the running efficiency of the model on the target hardware.
[0136] Deploy the quantized model on the LubanCat-RK3588 edge computing device, and use a depth camera as the image acquisition device to collect RGB images and depth maps in real time. Write code to detect RGB images in real time based on the YOLOv8 quantized model, obtain the bounding boxes and classification labels of objects, and at the same time combine the depth map to obtain the depth values of objects. The detection results are output in JSON format, for example, {"object":"Screwdriver","bbox":[20,30,100,150],"confidence":0.95,"depth":50,}, where "object" is the object label, "bbox" is the bounding box position parameter, "confidence" is the confidence of the object, and "depth" is the depth value of the object.
[0137] The language module, such as Figure 4 shown, specifically:
[0138] The user inputs the natural language text message "Take the screwdriver on the table to the right table" through the WeChat official account platform and sends it. The message interface of the WeChat developer platform sends the instruction from the WeChat server to the cloud server in the form of an HTTPS request.
[0139] The FLASK framework is deployed on the cloud server side to receive the HTTPS request from the WeChat server. After receiving the request, extract the text field from the JSON data to obtain "Take the screwdriver on the table to the right table" and perform preprocessing.
[0140] The preprocessed instructions are deployed on a large language model such as GLM-4 in the cloud through API calls. To improve the decomposition accuracy, CoT and few-shot prompting techniques are adopted, that is, a complete prompt is constructed by designing system prompts, task instructions, few-shot examples and their thought chains to generate language responses for natural language interaction and generate an executable action instruction sequence for action decomposition.
[0141] The reply message is encapsulated in JSON format and returned to the user through the passive reply interface of the WeChat official account to achieve a two-way interaction experience.
[0142] The decomposed action instruction sequence is sent from the cloud server to LubanCat-RK3588 through a Scoket object of the TCP protocol, and the data transmission format is JSON. After receiving the instruction sequence, LubanCat-RK3588 combines the position and depth information of the screwdriver obtained by the vision module to supplement and improve the parameters of the action instruction sequence. For example {
[0143] "action":
[0144] {"object_detection":{"object":"Screwdriver","bbox":[20,30,100,150],"confidence":0.95,"depth":50,}},
[0145] {"gripper":"open"},
[0146] {"arm_move":[20,30,100,150,50]},
[0147] {"gripper":"close"},
[0148] {"arm_move_safe":[]},
[0149] {"move":[200,10],}
[0150] {"object_detection":{"object":"table","bbox":[5,5,620,610],"confidence":0.95,"depth":60,}},
[0151] {"arm_move":[5,5,620,610,60]},
[0152] {"gripper":"open"},
[0153] {"arm_move_safe":[]},
[0154] {"gripper":"close"},
[0156] }。
[0157] If problems such as the vision module failing to detect the target object or incorrect instruction sequence parameters generated by the decomposition action unit occur, the LubanCat-RK3588's Scoket object based on the TCP protocol will send the problem situation to the cloud server, and the cloud server will actively notify the user by calling the template message interface of the WeChat official account.
[0158] The action module, as Figure 5 shown, specifically includes:
[0159] Write control programs for robot movement, robotic arm movement, etc. based on the HPAC programming platform to achieve path planning and motion control of the mobile grasping robot.
[0160] The action instructions and parameters obtained by the language module are transmitted from the LubanCat-RK3588 to the IPC based on the Modbus protocol. Modbus uses the RTU mode to ensure the stability and real-time performance of communication.
[0161] For the statement "Get the screwdriver on the table to the table on the right", the mobile grasping robot sequentially performs the following actions: Detect the position and depth of the screwdriver, open the end gripper, move the end of the robotic arm to the position of the screwdriver, close the end gripper, move the end of the robotic arm to a safe position, move the robot near the table on the right, detect the position and depth of the table, move the end of the robotic arm to the position of the tabletop, open the end gripper, move the end of the robotic arm to a safe position, and close the end gripper.
[0162] Through modular design, decouple and optimize the integration of visual perception, natural language interaction, large language model intelligent decision-making, and action control execution functions. It can not only identify the category, bounding box, and depth information of objects, but also meet the industrial real-time requirements, and at the same time solve problems such as complex interaction of traditional robots, limited decision-making ability, and complex construction and poor flexibility of embodied intelligent robots.
[0163] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A control method for an intelligent mobile grasping robot based on a modular architecture, characterized in that It includes the following steps: S1: Collect the image information of the target object in real time, process the image data in real time, identify the target object and output the position and category of the target object, and obtain the depth information of the detected target object; S2: Send natural language text through the user terminal. The cloud server receives the text information, uses the large language model for natural language interaction to generate language reply information, and decomposes actions to generate decomposed action instructions; The language reply information is returned to the user terminal through the cloud server, and the decomposed action instructions are sent to the edge computing device, so that the decomposed action instructions are combined with the depth information of the target object obtained in step S1 to obtain the final action instruction sequence; S3: Send the action instruction sequence generated in step S2 to the IPC, and the IPC controls the mobile grabbing robot to complete the relevant actions of moving and grabbing.
2. The control method of the intelligent mobile grasping robot based on the modular architecture according to claim 1, wherein, Step S1 specifically includes the following steps: S11: Collect images through web crawlers and camera shooting methods, and manually annotate each image to generate annotation files for object detection; S12: Perform preprocessing and data augmentation operations on the original images generated in step S11 to expand the dataset, and divide the dataset into a training set, a validation set, and a test set; S13: Use a depth camera to collect the RGB image and depth map of the target object in real time; S14: Select the improved YOLOv10 model and train the training set obtained in step S12 on the cloud server. After training is completed, obtain the model file; Quantify the model file to obtain the quantized model, and deploy it to the edge computing device to achieve local operation; At the same time, based on the quantized model, the RGB image collected in step S13 is detected in real time to obtain the bounding box and classification label of the target object, and the depth value of the object is obtained by combining the depth map collected in step S13, and the detection results are output.
3. The control method of the intelligent mobile grasping robot based on a modular architecture according to claim 2, characterized in that The specific steps of the preprocessing and data augmentation operations in step S12 are: automatically remove the background of some images, extract the target element map, and generate diverse synthetic images and corresponding annotation files by using random transformation and background synthesis.
4. The control method of the intelligent mobile grasping robot based on the modular architecture according to claim 2, wherein, The improved YOLOv10 model in step S14 includes a ResSPPF module, an OGCA module, and a C2f_ParNetAttention module. Among them, the ResSPPF module optimizes the multi-scale feature extraction ability by introducing a cascaded residual structure, and optimizes the detection performance of small targets and multi-scale targets; The OGCA module improves the adaptability of the model to complex scenes by introducing a global feature fusion mechanism and an occlusion perception branch; The C2f_ParNetAttention module enhances the detection performance by replacing the standard bottleneck block with a bottleneck block improved based on ParNetAttention to enhance the spatial perception of features and the modeling ability of the relationship between channels; The quantization method in step S14 is to convert the model weights from floating-point numbers to integers, perform format conversion and quantization processing on the trained model, and generate a quantized model suitable for NPU operation.
5. The control method of the intelligent mobile grasping robot based on the modular architecture according to claim 1, characterized in that, Step S2 specifically includes the following steps: S21: The user terminal inputs natural language text information through the public platform. The message interface of the public platform sends the instruction to the cloud server, and the cloud server receives and preprocesses it; S22: Invoke the large language model to perform natural language processing on the text information obtained in step S21. Utilize the semantic understanding ability of the large language model to perform natural language interaction based on the CoT and few-shot prompting methods to generate language reply information and perform action decomposition to generate an executable action instruction sequence; S23: Invoke the passive reply interface of the public platform to return the language reply information generated in step S22 to the user terminal; S24: Send the action instruction sequence generated in step S22 to the edge computing device through the cloud server, so that the decomposed action instructions are combined with the target category, bounding box, and depth information of the target object obtained in step S1 to construct a complete action instruction sequence; S24: For the problems that no target object is detected in step S1 and the instruction sequence parameters generated in step S24 are incorrect, send them to the cloud server through the edge computing device, and the cloud server invokes the template message interface of the public platform to actively send them to the user terminal.
6. The control method of the intelligent mobile grasping robot based on the modular architecture according to claim 1, wherein, Step S3 specifically includes the following steps: S31: Write a control program to implement path planning and motion control of the mobile grasping robot; S32: Transmit the action instruction sequence generated in step S2 from the edge computing device to the IPC, and the IPC controls the mobile grasping robot to complete the relevant actions of moving and grasping.
7. An intelligent mobile grasping robot control system based on a modular architecture, characterized in that, Adopt the method described in any one of claims 1-6. The system includes: A vision module that real-time collects image information of the target object, real-time processes the image data, identifies the target object and outputs the position and category of the target object, and obtains the depth information of the detected target object; A language module. The user terminal sends natural language text, and the cloud server receives the text information. Use the large language model to perform natural language interaction to generate language reply information and perform action decomposition to generate decomposed action instructions; return the language reply information to the user terminal through the cloud server; send the decomposed action instructions to the edge computing device, so that the decomposed action instructions are combined with the depth information of the target object obtained by the vision module to obtain the final action instruction sequence; An action module that sends the action instruction sequence generated by the language module to the IPC, and the IPC controls the mobile grasping robot to complete the relevant actions of moving and grasping.
8. The intelligent mobile grasping robot control system based on a modular architecture according to claim 7, characterized in that, The vision module includes: An image acquisition unit that collects images through web crawlers and camera shooting means, and manually annotates each image to generate an annotation file for object detection; An image preprocessing unit that preprocesses and performs data augmentation operations on the original images generated by the image acquisition unit, expands the dataset, and divides the dataset into a training set, a validation set, and a test set; A depth information acquisition unit that uses a depth camera to real-time collect the RGB image and depth map of the target object; The model training unit selects the improved YOLOv10 model and trains the model on the cloud server. After the training is completed, a model file is obtained. The model file is quantized to obtain a quantized model, which is deployed to the edge computing device to achieve local operation. At the same time, based on the quantized model, RGB images are detected in real time to obtain the bounding boxes and classification labels of the target objects. Meanwhile, the depth values of the objects are obtained by combining with the depth map, and the detection results are output.
9. The intelligent mobile grasping robot control system based on a modular architecture according to claim 7, characterized in that, The language module includes: The input unit. The user terminal inputs natural language text information through the public platform. The message interface of the public platform sends the instructions to the cloud server, and the cloud server receives and preprocesses them. The processing unit calls the large language model to perform natural language processing on the text information. Utilizing the semantic understanding ability of the large language model, natural language interaction is carried out based on the CoT and few-shot prompting methods to generate the information of the language reply, and the action instructions sequence that can be executed is generated by decomposing the actions. The interaction unit calls the passive reply interface of the public platform to return the information of the language reply obtained by the processing unit to the user. The decomposed action unit sends the action instructions sequence obtained by the processing unit from the cloud server to the edge computing device, so that the decomposed action instructions are combined with the target category, bounding box and depth information of the target object obtained by the vision module to construct a complete action instructions sequence. The error handling unit will send the problems that the vision module fails to detect the target object and the parameter errors of the instruction sequence generated by the decomposed action unit to the cloud server through the edge computing device, and the cloud server will actively send them to the user terminal by calling the template message interface of the public platform.
10. The intelligent mobile grasping robot control system based on a modular architecture according to any one of claims 7-9, characterized in that The action module includes: The control unit writes a control program to achieve the path planning and motion control of the mobile grasping robot. The communication unit transmits the action instructions sequence generated by the language module from the edge computing device to the IPC, and the IPC controls the mobile grasping robot to complete the relevant actions of moving and grasping.
Citation Information
Cited By
Human-computer interaction method and system based on vision-language-action model
CN121245915A
Robot control method and system and training method and system based on dual systems
CN122401441A