Multi-mode sensing multi-mechanical-arm cooperative control method and system and robot

Through multimodal perception technology combined with 3D vision, 2D vision and instruction set, the motion path and grasping posture of the robot arm are dynamically adjusted, solving the problem of the traditional single-arm robot's perception accuracy decrease in complex dynamic environments, and achieving high-precision and flexible multi-manipulator coordinated control.

CN120244968APending Publication Date: 2025-07-04CHANGSHU INSTITUTE OF TECHNOLOGY +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510494796.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional single-arm robots perceive the problem of reduced accuracy and task failure in complex dynamic production environments. They lack efficient data fusion and modal switching mechanisms, making it difficult to meet the needs of complex dynamic production environments.

Method used

Multimodal perception technology is adopted, combining 3D vision, 2D vision and instruction set, and the module is selected through task type judgment, data is collected and fused in real time, the motion path and grabbing posture of the robot arm are dynamically adjusted, the instruction set module is set to support manual intervention, and the weighted average algorithm and modal optimization algorithm are used to optimize the task strategy.

Benefits of technology

It improves the operating accuracy and flexibility of the system in complex dynamic environments, enhances the stability and adaptability of collaborative control of multiple robotic arms, meets diverse operational needs, and realizes high-precision task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244968A_ABST
    Figure CN120244968A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode sensing multi-mechanical-arm cooperative control method and system and a robot. Task types are judged according to power consumption and path requirements of a target task; performing module selection according to the task type; data such as material positions, material shapes, obstacle positions and obstacle shapes are collected in real time; the real-time data collected by the sensing module is fused, and the motion path, the motion speed, the motion acceleration, the grabbing posture and the grabbing force of the mechanical arm are determined; storing the execution data and task feedback information, analyzing an operation effect, and optimizing modal selection and parameter setting of a next task; through 3D vision, 2D vision and instruction set multi-modal perception and data fusion, the motion path of the robot is continuously optimized, multi-source data are integrated in real time, and the accuracy and robustness of material recognition and state description are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi - robotic - arm collaborative control, and particularly to a multi - robotic - arm collaborative control method, system and robot with multi - modal perception. Background Art

[0002] With the rapid development of intelligent manufacturing, robots have become indispensable core equipment in assembly - line production, and are widely used in manufacturing industries such as automobiles, electronics, and new energy. Customized production is completed on the assembly line of the production assembly line through the intelligent control of robots.

[0003] When dealing with complex tasks, the traditional single - robotic - arm loading and unloading system is often limited by the limitations of single - modal perception and is difficult to meet the requirements of complex dynamic production environments. There are already some multi - modal perception solutions in the prior art, such as combining visual sensors and force sensors to improve perception accuracy. However, these solutions are usually applicable to single - robotic - arm robots and have deficiencies in dynamic adjustment and task collaboration. In addition, the lack of efficient data fusion and modality switching mechanisms makes the system prone to problems such as accuracy degradation or task failure in high - speed dynamic environments. Therefore, it is necessary to introduce advanced multi - modal perception technologies, dynamic adjustment mechanisms, and efficient data fusion and modality switching strategies, so as to effectively improve the task execution ability of single - robotic - arm robots in complex dynamic production environments and meet the requirements of higher accuracy and more efficient collaboration. Summary of the Invention

[0004] The present application provides a multi - robotic - arm collaborative control method, system and robot with multi - modal perception. In the mode of complex dynamic working environments and complex work instructions, the multi - robotic - arms of the robot can achieve multi - modal dynamic real - time and precise adjustment, meet the collaborative control of multi - robotic - arms, multi - modalities, and multi - instructions. Through multi - modal perception such as 3D vision, 2D vision, and instruction sets, combined with efficient data fusion and processing, the flexibility and stability of the system can be significantly improved.

[0005] An embodiment of the present application provides a multi - robotic - arm collaborative control method with multi - modal perception, including:

[0006] S1. Task type judgment: Judge the task type according to the power consumption and path requirements of the target task. The task types include: low - power consumption strategy, complex path optimization strategy, simple path optimization strategy, fixed task strategy, complex task strategy, simple task strategy;

[0007] S2. Module selection: Select modules according to the task type. The modules include a perception module and an instruction set module; the perception module includes: 2D vision module, 3D vision module;

[0008] S3. Real-time data acquisition, where the sensing module acquires real-time data such as the position of the material, the shape of the material, the position of the obstacle, and the shape of the obstacle;

[0009] S4. Data fusion and path planning:

[0010] When the task type is complex path optimization strategy, simple path optimization strategy, complex task strategy, or simple task strategy, fuse the real-time data collected by the sensing module to determine the movement path, movement speed, movement acceleration, grasping posture, and grasping force of the robotic arm;

[0011] When the task type is low-power strategy or fixed task strategy, call the instructions predefined in the instruction set module to determine the movement path, movement speed, movement acceleration, grasping posture, and grasping force of the robotic arm;

[0012] S5. Modal optimization and learning, store the data and task feedback information of this execution, analyze the operation effect, and optimize the modal selection and parameter settings for the next task.

[0013] Preferably, in step S2, module selection is performed according to the task type, specifically:

[0014] For the low-power strategy, select to enable the instruction set module;

[0015] For the complex path optimization strategy, select to enable the 3D vision module;

[0016] For the simple path optimization strategy, select to enable the 2D vision module;

[0017] For the fixed task strategy, select to enable the instruction set module;

[0018] For the complex task strategy, select to enable the 3D vision module;

[0019] For the simple task strategy, select to enable the 2D vision module.

[0020] Preferably, the 2D vision module is used to output the planar coordinates of the material.

[0021] Preferably, the 3D vision module is used to output the spatial coordinates of the material.

[0022] Preferably, in step S4, fusing the real-time data collected by the sensing module specifically means:

[0023] The fusion process uses the weighted average algorithm: D fusion =λ 2D ·D 2D +λ 3D ·D 3D

[0024] where λ 2Dis the weight coefficient of 2D visual data; λ 3D is the weight coefficient of 3D visual data; D 2D is two-dimensional visual data; D 3D is three-dimensional visual data; D fusion is the fused perception data.

[0025] Preferably, the weighted average algorithm:

[0026] When the surface texture of the material is poor, increase the weight coefficient of 3D visual data;

[0027] When the material is partially occluded or the texture information is lost, increase the weight coefficient of 2D visual data.

[0028] Preferably, the method for determining the grasping posture in step S4 is:

[0029] θ adjust = θ desire + K · (θ target - θ actual )

[0030] Where: θ adjust is the adjusted joint angle of the robotic arm; θ desire is the desired angle; θ target , θ actual are the target posture and the current posture respectively; K is the gain coefficient, which determines the speed and amplitude of the adjustment.

[0031] Preferably, the optimization of the modality selection for the next task in step S5 is specifically:

[0032]

[0033] Where: M is the optimized modality; f(T, C, R) is the optimized modality selection function; T is the task type; C is the resource condition; R is the feedback result;

[0034] Ω is the set of all available modalities; M k is the currently evaluated modality; E accuracy (M k ), E time (M k ), E resource (M k ) respectively represent the accuracy error, time consumption, and resource occupancy of the modality; α, β, γ are adjustable weight coefficients, and α + β + γ = 1.

[0035] This application also proposes a multi-robot collaborative control system with multi-modal perception for executing the above multi-robot collaborative control method with multi-modal perception, including:

[0036] A task type judgment module, which is used to judge the task type according to the power consumption and path requirements of the target task. The task types include: low-power strategy, complex path optimization strategy, simple path optimization strategy, fixed task strategy, complex task strategy, and simple task strategy;

[0037] A modality selection module, which is used to select modules according to the task type. The modules include a perception module and an instruction set module; the perception module includes: a 2D vision module and a 3D vision module;

[0038] A data acquisition module, which is used to collect data such as the position of materials, the shape of materials, the position of obstacles, and the shape of obstacles in real time;

[0039] A data fusion and path planning module, which is used to fuse the collected real-time data to determine the movement path, movement speed, movement acceleration, grasping posture, and grasping force of the robotic arm;

[0040] A modality optimization module, which is used to store the data and task feedback information of the current execution, analyze the operation effect, and optimize the modality selection and parameter settings of the next task.

[0041] This application also proposes a multi-robotic-arm collaborative control robot with multi-modal perception, including a processor, a memory, and a computer program stored on the memory and executable on the processor, which is used to execute the above multi-modal perception multi-robotic-arm collaborative control method.

[0042] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0043] 1. This application divides the task types according to the task objectives into 5 types, namely low-power strategy, complex path optimization strategy, simple path optimization strategy, fixed task strategy, complex task strategy, and simple task strategy, and selects corresponding path planning methods for each type of task, introducing a multi-modal perception solution of 3D vision, 2D vision, and instruction set module. By integrating 3D vision, 2D vision, and instruction set multi-modal perception, this application realizes accurate perception and operation of the dynamic production environment, and has characteristics such as high-precision task execution, multi-modal perception ability, and flexible adaptability.

[0044] 2. This application combines 3D vision and 2D vision to obtain all and partial information of the materials. Compared with the traditional single perception mode, it can simultaneously obtain the three-dimensional geometric features, color texture features, and task requirement information of the materials.

[0045] 3. The present application also sets up an instruction set module, which can receive and parse the host computer instructions in real time and support the task requirements of manual intervention. This design greatly enhances the flexibility and adaptability of the system for different tasks, is applicable to a variety of industrial scenarios, and meets diverse operation requirements.

[0046] 4. Through the weighted fusion algorithm, the multi-modal data fusion and dynamic switching mechanism of the present application ensure the operation accuracy of the dual-arm robot in a complex environment, integrate multi-source data in real time, and improve the accuracy and robustness of material recognition and status description.

[0047] 5. The present application adopts a modal optimization algorithm to store, calculate, analyze, and iterate each task instruction and feedback. Through the dynamic optimization algorithm, the system can automatically adjust the modal weights and task strategies under different working conditions, so as to flexibly respond to dynamic production requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of the multi-robot collaborative control method with multi-modal perception in the first embodiment of the present application;

[0049] Figure 2 is an architecture diagram of the multi-robot collaborative system with multi-modal perception in the second embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0051] As Figure 1 shown, the multi-robot collaborative control method with multi-modal perception includes:

[0052] Step 1: After the system is started, receive task instructions through the instruction set module, parse the task requirements, such as material type, target location, path requirements, etc., and determine the perception mode and control strategy:

[0053] Fixed tasks, such as repetitive handling: For tasks with fixed environment and path, the system directly calls the predefined instruction set module, without complex perception and dynamic adjustment, reducing the calculation amount and power consumption.

[0054] Complex tasks, such as complex shapes and unknown characteristics: Enable the 3D vision module to obtain the three-dimensional coordinates and shape information of the material for high-precision positioning and grasping.

[0055] Simple tasks, such as color classification or grasping with simple characteristics: Enable the 2D vision module to quickly identify the two-dimensional features of the material, such as color, texture, edges, and perform classification or grasping.

[0056] Modal selection function z: The system determines according to the task typeT Dynamically evaluate and select appropriate modalities

[0057] M = f(T), T ∈ {Fixed, Complex, Simple}

[0058] where the modality M can take values from the instruction set module, 2D vision module, or 3D vision module.

[0059] Step 2: Visual perception data acquisition and processing

[0060] After completing the modality selection, the system starts the corresponding visual perception module according to the task requirements for data acquisition and processing. If the 2D vision module is enabled, the industrial camera will rapidly capture two-dimensional images of the area where the material is located, and perform preprocessing such as denoising, enhancing contrast, and binarization on the images. Subsequently, key features such as the color, shape, and edges of the material are extracted through object detection algorithms (such as YOLO or SSD). For tasks that require higher precision, the 3D vision module will use a depth camera or lidar to generate point cloud data of the material, reducing data redundancy and improving quality through spatial filtering and keypoint extraction. The processed point cloud data is transformed into the world coordinate system to form a three-dimensional model of the material, locate the grasping points, and support subsequent motion planning. In a multi-modal scenario, the texture information of the two-dimensional image and the geometric information of the three-dimensional point cloud complement each other, enhancing the data integrity and perception accuracy. Pixel coordinates p = (x, y) are converted to three-dimensional coordinates X = (X, Y, Z)

[0061]

[0062] Step 3: When the task requirements involve multi-modal perception, the system will perform fusion processing on the 2D and 3D data. The fusion process adopts a weighted average algorithm to combine the two-dimensional visual data D 2D and the three-dimensional point cloud data D 3D into a unified perception input D fusion . The weight coefficient λ 2D λ 3D is dynamically adjusted according to the task requirements to ensure the accuracy of the fusion result. The fusion process adopts a weighted average algorithm:

[0063] D fusion = λ 2D · D 2D + λ 3D · D 3D

[0064] When there is occlusion or noise, two-dimensional data is preferentially used. To further optimize the perception results, the system evaluates the deviation between the fused data and the actual target position through an error minimization objective function, and adjusts the weights through an iterative optimization algorithm to reduce the perception error E. This dynamic adjustment mechanism makes the fused data more accurate and provides reliable input for task execution.

[0065]

[0066] Among them, (X fusion , Y fusion , Z fusion ) is the spatial position of the fused data, and (X target , Y target , Z target ) is the target position. Through an iterative optimization algorithm (such as gradient descent), the system dynamically adjusts the weights to minimize the error.

[0067] In the weighted average fusion algorithm, the setting of the weight quantization standard needs to be combined with the specific application scenario and sensor characteristics. Here, weight quantization indicators are established for different strategies.

[0068] 1. Texture complexity: Calculated through image gradient / entropy value

[0069] 2. Object occlusion ratio: Judged through 2D image segmentation or 3D point cloud missing rate

[0070] 3. Task complexity: Graded according to obstacle density or operation accuracy requirements

[0071] The specific adjustment methods and bases are shown in the following table:

[0072]

[0073] Regarding the dynamic adjustment algorithm:

[0074] In the dynamic adjustment algorithm, the Sigmoid function can be used. Due to its properties such as monotonic increase and monotonic increase of the inverse function, the Sigmoid function is often used as the activation function of neural networks to map variables between 0 and 1.

[0075]

[0076] τ: Texture score of the current scene (can be calculated through image entropy and gradient, the larger the value, the richer the texture).

[0077] τ0: Threshold (preset texture score benchmark).

[0078] k: Slope coefficient (controlling the sensitivity of weight change).

[0079] Step 4: After obtaining accurate perception data, the system controls the coordinated actions of the dual-arm manipulator according to the grasping point position and task path planning. First, the system analyzes the shape of the material and the grasping point position, and adjusts the end pose of the robotic arm to ensure a stable and precise grasping action. Subsequently, the dual-arm robot coordinates its motion trajectory according to the target path generated from multi-modal data, avoiding collisions and ensuring smooth operation. During path planning, the system continuously monitors for dynamic obstacles that may appear in the environment and uses obstacle avoidance algorithms to adjust the motion route. At the same time, the motion of the two arms adjusts the pose through a dynamic gain control formula to adapt to the actual requirements of material grasping, ensuring that the system always performs stably and reliably in high-complexity tasks. The grasping pose adjustment follows the following formula:

[0080] θ adjust = θ desire + K · (θ target - θ actual )

[0081] Where:

[0082] θ adjust is the adjusted joint angle of the robotic arm.

[0083] θ desire is the desired angle.

[0084] θ target , θ actual are the target pose and the current pose respectively.

[0085] K is the gain coefficient, which determines the speed and amplitude of the adjustment.

[0086] The dual-arm manipulator plans its path through a cooperative optimization algorithm to avoid motion conflicts and material damage. In path planning, the objective function is the minimization of the path length L:

[0087]

[0088] This optimization process ensures that the robotic arm can efficiently execute complex grasping tasks.

[0089] Step 5: After grasping, the system plans the optimal transportation path according to the target position and task requirements and controls the dual-arm robot to transport the material. During transportation, the system adjusts the motion path in real time to avoid dynamic interference from the external environment and ensure the safety of the material. For fragile or specially shaped materials, the path planning will particularly consider reducing the motion speed of the robotic arm and optimizing the acceleration curve to reduce the impact of vibration and shock on the material. During the placement phase, the robotic arm precisely controls the pose of the material and safely places the material at the target position by gradually adjusting the angle and position of the end effector. The entire process focuses on stability and precision to ensure the successful completion of the task and the material remains undamaged.

[0090] Step 6: After the task is completed, the system stores the key data during the execution process, such as completion time, path accuracy, and grasping success rate, in the task feedback module, and evaluates and optimizes the system performance based on these data. The task feedback information not only includes the execution status of the current task, but also covers the dynamic records of environmental changes, material characteristics, and grasping accuracy. These feedback data are used to optimize the modality selection function to improve the efficiency and accuracy of future tasks. The optimized modality selection function f(T, C, R) (task type, resource conditions, feedback result) aims to more intelligently match task requirements with resource conditions, and its update process is based on the following formula:

[0091]

[0092] where: Ω is the set of all available modalities (e.g., 2D vision, 3D vision, fixed task modality). M k is the currently evaluated modality. E accuracy (M k ), E time (M k ), E resource (M k ) represent the precision error, time consumption, and resource occupancy of the modality respectively. α, β, and γ are adjustable weight coefficients, and α + β + γ = 1. These data are used as references for subsequent task planning. Through the modality selection optimization function, the system can dynamically adjust the modality selection strategy and parameter settings in future tasks. For example, if a certain task type frequently experiences grasping failures, the system will update the modality selection logic or perception algorithm for this type to reduce the error rate in similar tasks. This adaptive optimization mechanism enables the system to continuously improve task efficiency and execution accuracy, while enhancing its self-learning and adaptation capabilities in complex dynamic environments. Specific embodiment:

[0094] In an intelligent manufacturing workshop, a dual-arm robot is used to complete a repetitive and complex material handling and assembly task. The goal of the task is to grasp different types of components (such as electronic components, batteries, circuit boards) from the bins and assemble them into the product frame on the production line. The task involves material grasping, path planning, material placement, and assembly.

[0095] Step 1: Task startup and task instruction parsing

[0096] After the system starts up, the dual-arm robot receives task instructions through the instruction set module. These task instructions include:

[0097] Material type (such as battery, circuit board, etc.)

[0098] Target positions (such as assembly positions, temporary storage positions)

[0099] Path requirements (such as whether to bypass certain obstacles, shortest path, etc.)

[0100] The task type is identified as a complex task because of the variety of material types and complex shapes. The system evaluates and selects appropriate sensing modalities and control strategies according to the task requirements.

[0101] Task type analysis: Due to the complex shape of the materials, the system determines it as a complex task and chooses to enable the 3D vision module to obtain the three-dimensional coordinates and shape information of the materials to ensure grasping accuracy.

[0102] Step 2: Visual perception data acquisition and processing

[0103] After the system starts, the depth camera of the 3D vision module begins to work, scanning the material area to obtain depth images and point cloud data. These data are used for:

[0104] Accurately obtaining the three-dimensional positions and shapes of each material (such as batteries, circuit boards).

[0105] The processed point cloud data is transformed into the world coordinate system to form an accurate three-dimensional model of the material.

[0106] The deviation between the three-dimensional model of the battery and the target position is identified, and redundant data is reduced through spatial filtering and point cloud optimization algorithms to ensure position accuracy.

[0107] Step 3: Data fusion and dynamic adjustment

[0108] In this task, two-dimensional visual information is used to identify simple features such as material color and shape, while three-dimensional visual information is used to obtain accurate spatial positions. To improve the perception accuracy, the system uses a weighted average algorithm to fuse the two-dimensional and three-dimensional data.

[0109] Data fusion process: When the surface texture of the material is poor, the system increases the weight of the three-dimensional data. When the material is partially occluded or texture information is lost, the system preferentially uses the two-dimensional image data to supplement the information. Through the error minimization algorithm, the system adjusts the weight coefficients in the data fusion in real time to minimize the deviation between the material position and the target position. If the point cloud data is severely occluded, the system will adjust the weight and rely more on the clear two-dimensional image data.

[0110] Step 4: Dual-arm collaborative control and motion planning

[0111] Based on the perception data, the system plans the motion paths for the dual-arm manipulator and controls the two arms to work in coordination.

[0112] Task path planning: When grasping the battery and circuit board, the two arms need to bypass other obstacles on the production line. The system uses an obstacle avoidance algorithm to plan the shortest path that avoids obstacles.

[0113] Two-arm synchronization and attitude adjustment: The two-arm manipulator dynamically adjusts its attitude through a gain control algorithm (such as PID control) to ensure stability and synchronization during the grasping process. For example, when the right arm grasps the battery, the left arm adjusts its dynamic gain to ensure that it does not collide and maintains synchronous movement.

[0114] Step 5: Material handling and placement

[0115] After the grasping is completed, the two-arm robot needs to transport the material from the bin to the designated position on the assembly line. Optimization of the handling path: The path planning is adjusted in real time to avoid other robots and workers on the production line. For fragile materials (such as batteries), the system pays special attention to slowing down the movement speed of the manipulator and optimizing the acceleration curve to reduce the impact of vibration on the material.

[0116] Material placement: When the two arms reach the target position, the system precisely controls the angle and position of the end effector to safely place the material at the designated position.

[0117] Step 6: Task feedback and optimization

[0118] After the task is completed, the system stores the key data during the execution process (such as completion time, path accuracy, grasping success rate) in the task feedback module.

[0119] Task feedback record: including environmental changes, material types, grasping accuracy, etc.

[0120] Modal selection optimization: The system evaluates the execution effect of the task based on the feedback data. If the grasping task of the battery often fails, the system will analyze the perception algorithm of this task and adjust the modal selection logic according to the historical feedback.

[0121] Embodiment 2

[0122] As Figure 2 shown, this embodiment proposes a multi-modal perception multi-manipulator collaborative control system, including:

[0123] A task type judgment module, used to judge the task type according to the power consumption and path requirements of the target task, and the task types include: low-power strategy, complex path optimization strategy, simple path optimization strategy, fixed task strategy, complex task strategy, simple task strategy;

[0124] A modal selection module, used to select modules according to the task type, and the modules include a perception module and an instruction set module; the perception module includes: a 2D vision module, a 3D vision module;

[0125] A data acquisition module, configured to acquire in real time data such as the position of the material, the shape of the material, the position of the obstacle, and the shape of the obstacle;

[0126] A data fusion and path planning module, configured to fuse the acquired real-time data to determine the movement path, movement speed, movement acceleration, grasping posture, and grasping force of the robotic arm;

[0127] A modal optimization module, configured to store the data and task feedback information of the current execution, analyze the operation effect, and optimize the modal selection and parameter setting of the next task.

[0128] Embodiment III

[0129] An embodiment of the present application provides a multi-modal perception multi-robotic arm collaborative control robot, including a processor, a memory, and a computer program stored on the memory and executable on the processor, for implementing the multi-modal perception multi-robotic arm collaborative control method described in Embodiment I.

[0130] The above computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0131] The embodiments of this specific implementation manner are all preferred embodiments of the present invention, and do not limit the protection scope of the present invention accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention. Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A multi-arm collaborative control method for multi-modal perception, characterized in that, Including: S1. Task type judgment: Judge the task type according to the power consumption and path requirements of the target task. The task types include: low-power strategy, complex path optimization strategy, simple path optimization strategy, fixed task strategy, complex task strategy, simple task strategy; S2. Module selection: Select modules according to the task type. The modules include a sensing module and an instruction set module; The sensing module includes: 2D vision module, 3D vision module; S3. Real-time data acquisition: The sensing module acquires data such as the position of materials, the shape of materials, the position of obstacles, and the shape of obstacles in real time; S4. Data fusion and path planning: When the task type is complex path optimization strategy, simple path optimization strategy, complex task strategy, or simple task strategy, fuse the real-time data collected by the sensing module to determine the movement path, movement speed, movement acceleration, grasping posture, and grasping force of the robotic arm; When the task type is low-power strategy or fixed task strategy, call the instructions predefined by the instruction set module to determine the movement path, movement speed, movement acceleration, grasping posture, and grasping force of the robotic arm; S5. Modal optimization and learning: Store the data and task feedback information of the current execution, analyze the operation effect, and optimize the modal selection and parameter settings for the next task.

2. The multi-arm collaborative control method for multi-modal perception according to claim 1, wherein, In step S2, the module selection according to the task type is specifically: The low-power strategy selects to enable the instruction set module; The complex path optimization strategy selects to enable the 3D vision module; The simple path optimization strategy selects to enable the 2D vision module; The fixed task strategy selects to enable the instruction set module; The complex task strategy selects to enable the 3D vision module; The simple task strategy selects to enable the 2D vision module.

3. The multi-arm collaborative control method for multi-modal perception according to claim 1, characterized in that The 2D vision module is used to output the planar coordinates of the material.

4. The multi-arm collaborative control method for multi-modal perception according to claim 1, characterized in that The 3D vision module is used to output the spatial coordinates of the material.

5. The multi-arm collaborative control method for multi-modal perception according to claim 1, characterized in that In step S4, the fusion of the real-time data collected by the sensing module is specifically: The fusion process adopts the weighted average algorithm: D fusion = λ 2D · D 2D + λ 3D · D 3D Among them, λ 2D is the weight coefficient of 2D visual data; λ 3D is the weight coefficient of 3D visual data; D 2D is the two-dimensional visual data; D 3D is the three-dimensional visual data; D fusion is the fused perception data.

6. The multi-arm collaborative control method for multi-modal perception according to claim 5, characterized in that, The weighted average algorithm: When the surface texture of the material is poor, increase the weight coefficient of the 3D vision data; When the material is partially blocked or the texture information is lost, increase the weight coefficient of the 2D vision data.

7. The multi-arm collaborative control method for multi-modal perception according to claim 1, characterized in that The method for determining the grasping posture in step S4 is: θ adjust = θ desire + K·(θ target - θ actual ) Where: θ adjust is the adjusted joint angle of the robotic arm; θ desire is the desired angle; θ target , θ actual are the target pose and the current pose respectively; K is the gain coefficient, which determines the speed and amplitude of the adjustment.

8. The multi-arm collaborative control method for multi-modal perception according to claim 1, wherein In step S5, the optimization of the modal selection for the next task is specifically: Where: M is the optimized mode; f(T, C, R) is the optimized mode selection function; T is the task type; C is the resource condition; R is the feedback result; Ω is the set of all available modalities; M k is the currently evaluated modality; E accuracy (M k ),E time (M k ), E resource (M k ) represent the precision error, time consumption, and resource occupancy of the modality respectively; α, β, γ are adjustable weight coefficients, and α + β + γ = 1.

9. A multi-robot arm collaborative control system with multi-modal perception, characterized in that, For implementing the multi-robotic-arm collaborative control method with multi-modal perception as described in any one of claims 1-8, including: A task type judgment module, used to judge the task type according to the power consumption and path requirements of the target task. The task types include: low-power strategy, complex path optimization strategy, simple path optimization strategy, fixed task strategy, complex task strategy, simple task strategy; A modal selection module, used to select modules according to the task type. The modules include a sensing module and an instruction set module; The sensing module includes: 2D vision module, 3D vision module; A data acquisition module, which is used to collect data such as the position of the material, the shape of the material, the position of the obstacle, and the shape of the obstacle in real time; A data fusion and path planning module, which is used to fuse the collected real-time data to determine the movement path, movement speed, movement acceleration, grasping posture, and grasping force of the robotic arm; A modal optimization module, which is used to store the data executed this time and task feedback information, analyze the operation effect, and optimize the modal selection and parameter setting of the next task.

10. A multi-robot collaborative control robot with multi-modal perception, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor, which is used to execute the multi-robotic arm cooperative control method with multi-modal perception according to any one of claims 1-8.

Citation Information

Cited By

  • Intelligent sorting management system for industrial robot

    CN120790549A

  • Double-arm robot autonomous control system and method based on remote operation and visual features

    CN120816484A

  • Dexterous hand multi-mode sensing and control method, system and equipment and medium

    CN121552386A

  • Double-mechanical-arm cooperative grabbing method based on drug traceability code and multi-source vision fusion

    CN121733545A