A collaborative carrying method based on multimodal large model mechanical arm and AGV integration

By integrating a multimodal large model with an AGV for collaborative handling, an environmental map is constructed in real time and image-enhanced localization is performed. This solves the problems of insufficient automation and environmental perception in traditional automated guided vehicles, and achieves efficient and accurate object handling and human-machine interaction.

CN121120764BActive Publication Date: 2026-04-07TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional automated guided vehicles (AGVs) robots have low levels of automation in cargo handling, insufficient environmental perception and visual positioning accuracy, lack efficient human-machine interaction capabilities, and are unable to adapt to dynamic and complex environments or understand users' natural commands.

Method used

By constructing a multimodal large-model robotic arm and AGV integrated collaborative handling method, an environmental map is built in real time. The multimodal large-model is used for image enhancement and localization. Combined with the robotic arm's grasping and placing of objects, autonomous navigation and high-precision object positioning are achieved.

Benefits of technology

It improves the automation level of automated guided vehicles, enhances positioning accuracy and grasping success rate in complex environments, strengthens human-machine interaction capabilities, and meets the needs of intelligent, efficient, and precise handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120764B_ABST
    Figure CN121120764B_ABST
Patent Text Reader

Abstract

The application provides a collaborative handling method based on multimodal large model mechanical arm and AGV integration, by constructing a novel collaborative handling mechanism, effectively solving the problems of low automation degree of the existing automatic guided vehicle handling strategy, high task execution time cost and low system response efficiency, and effectively solving the problems of insufficient environment perception and visual positioning accuracy of the existing automatic guided vehicle handling strategy, leading to positioning difficulties in complex or dynamic scenes and low target grasping success rate, and effectively solving the problems of lack of efficient human-computer interaction mechanism, limited dynamic interaction capability and inability to flexibly adapt to changing working conditions of the existing automatic guided vehicle handling strategy, and effectively solving the problem of poor understanding ability of the existing automatic guided vehicle handling strategy for natural language instructions of users, so that the intelligent, efficient and accurate handling requirements can be met in actual application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mobile mechanical operation, in particular to a collaborative carrying method based on integration of a multi-modal large model mechanical arm and an AGV. BACKGROUND

[0002] Robots have been widely used in various scenes in life due to the development of related technologies, especially the AGV robot (the robot generally adopts the abbreviation of the automated guided vehicle or AGV).

[0003] However, the present inventors have found that the conventional AGV robot generally relies on fixed algorithm modes in the process of realizing cargo carrying, and lacks environmental perception and flexible interaction capabilities.

[0004] Specifically, the carrying mode of the conventional AGV robot can be mainly divided into the following two modes:

[0005] (1) An AGV carrying strategy based on human cooperation, which realizes real-time mapping and positioning through SLAM (Simultaneous Localization And Mapping) technology and laser radar scanning of environmental features, combines a dynamic path planning algorithm for autonomous navigation and obstacle avoidance, and intelligently schedules tasks by a central system. When the AGV reaches the designated location, the user is reminded to manually put in or take out the carried object through interactive methods such as voice or light prompting, which is difficult to realize fully automated whole-process carrying.

[0006] (2) An AGV carrying strategy based on traditional extension, which realizes integrated operation of grabbing, carrying and assembling by fusing a traditional mechanical arm with an AGV, so that it is upgraded from a single transportation device to an intelligent composite robot with operation capabilities. The traditional mechanical arm used in this strategy lacks flexible strategy functions such as visual perception and instruction understanding, and is difficult to adapt to dynamically changing physical environments.

[0007] In summary, the conventional AGV carrying robot has the following limitations:

[0008] (1) The AGV carrying strategy based on human cooperation has low automation level, and the AGV carrying strategy based on traditional extension has weak environmental understanding and interaction capabilities. Both strategies cannot realize highly automated carrying tasks in dynamically complex environments.

[0009] (2) Lack of automatic, high-precision visual positioning function for the object being transported, making it difficult to accurately and automatically grasp the object, so in scenarios with high repeatability and large environmental changes, there may be problems of failed grasping and large time overhead;

[0010] (3) Lack of ability to understand natural user instructions, unable to truly understand and execute tasks, relying only on pre-set fixed paths and manual intervention, so in complex and changing scenarios, data needs to be collected and programs need to be set, increasing the complexity of the task. SUMMARY

[0011] The present application provides a collaborative transportation method based on the integration of multi-modal large model robot arm and AGV, by constructing a novel collaborative transportation mechanism, effectively solving the problem of low automation of existing automated guided vehicle transportation strategy, resulting in high time cost and low system response efficiency, and effectively solving the problem of insufficient environmental perception and visual positioning accuracy of existing automated guided vehicle transportation strategy, resulting in difficulty in positioning in complex or dynamic scenarios and low target grasping success rate, and effectively solving the problem of lack of efficient human-computer interaction mechanism and limited dynamic interaction capability of existing automated guided vehicle transportation strategy, unable to flexibly adapt to changing working conditions, and effectively solving the problem of poor understanding of natural language instructions by users of existing automated guided vehicle transportation strategy, so as to meet the intelligent, efficient and accurate transportation needs in practical applications.

[0012] In a first aspect, the present application provides a collaborative transportation method based on the integration of multi-modal large model robot arm and AGV, the method comprising:

[0013] An environmental map of the preset running environment is constructed in real time by the automated guided vehicle during driving in the preset running environment;

[0014] An operating route passing through the starting point and the target point is selected according to the environmental map;

[0015] Based on the structural analysis of the operating route, the automated guided vehicle autonomously navigates from the starting point to the target point;

[0016] The parameters of the robot arm driven based on the multi-modal large model are initialized for the environment where the target point is located, wherein the parameters of the robot arm include robot arm parameters and robot arm position;

[0017] An environment image of the object to be transported at the target point is obtained by using a camera mounted on the robot arm to take pictures of the environment of the object to be transported;

[0018] The environment image is data enhanced in combination with a cross-attention mechanism algorithm and an object name of the object to be transported, to obtain an enhanced environment image;

[0019] input the enhanced environment image and the predefined system prompt word into the multi-modal large model to obtain pixel position coordinates of the required carrying object;

[0020] convert the pixel position coordinates into relative position coordinates of the robot in a physical world coordinate system by using a hand-eye calibration algorithm;

[0021] According to the relative position coordinates, call the action execution module of the robot to the specified position to grasp the required carrying object, and then place the required carrying object;

[0022] Let the automated guided vehicle autonomously navigate back to the starting point according to the running route.

[0023] In a second aspect, the application provides a collaborative carrying device integrating a robot and an AGV based on a multi-modal large model, the device comprising:

[0024] a mapping unit configured to construct an environment map of a preset running environment in real time by an automated guided vehicle during driving in the preset running environment;

[0025] a selection unit configured to select a running route passing through a starting point and a target point according to the environment map;

[0026] a navigation unit configured to autonomously navigate the automated guided vehicle from the starting point to the target point based on structured analysis of the running route;

[0027] an initialization unit configured to initialize parameters of a robot driven based on a multi-modal large model for an environment where the target point is located, wherein the parameters of the robot include robot parameters and robot position;

[0028] a shooting unit configured to shoot an environment of a required carrying object of the target point by using a camera mounted on the robot to obtain an environment image;

[0029] an enhancement unit configured to perform data enhancement on the environment image in combination with a cross-attention mechanism algorithm and an object name of the required carrying object to obtain an enhanced environment image;

[0030] a positioning unit configured to input the enhanced environment image and a predefined system prompt word into the multi-modal large model to obtain pixel position coordinates of the required carrying object;

[0031] a conversion unit configured to convert the pixel position coordinates into relative position coordinates of the robot in a physical world coordinate system by using a hand-eye calibration algorithm;

[0032] a grasping and placing unit configured to call an action execution module of the robot to a specified position to grasp the required carrying object according to the relative position coordinates, and then place the required carrying object;

[0033] A return unit is configured to enable the AGV to autonomously navigate back to the starting point according to the operation route.

[0034] In a third aspect, the present application provides a processing device, comprising a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program in the memory to implement the method in the first aspect.

[0035] In a fourth aspect, the present application provides a computer readable storage medium, which stores a plurality of instructions, and the instructions are adapted to be loaded by a processor to implement the method in the first aspect.

[0036] From the above, the present application has the following beneficial effects:

[0037] For the object carrying target based on the AGV, the present application constructs a novel collaborative carrying mechanism, effectively solves the problems of low automation degree of the existing AGV carrying strategy, which leads to high time cost of task execution and low system response efficiency, and effectively solves the problems of insufficient environmental perception and visual positioning accuracy of the existing AGV carrying strategy, which leads to positioning difficulty in complex or dynamic scenes and low target grasping success rate, and effectively solves the problems of lack of efficient human-computer interaction mechanism, limited dynamic interaction capability and inability to flexibly adapt to changing working conditions of the existing AGV carrying strategy, and effectively solves the problem of poor understanding ability of the existing AGV carrying strategy for natural language instructions of users, so as to meet the intelligent, efficient and accurate carrying demand in practical application. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0039] Figure 1 FIG. 1 is a flowchart of a collaborative carrying method based on a multimodal large model mechanical arm and AGV integration according to an embodiment of the present application. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0041] The terms "first", "second", and the like in the description and in the claims of the present application and above drawings are used for distinguishing between similar objects and not necessarily for describing a specific sequential or chronological order. It is to be understood that the use of the terms so construed can be interchanged, such that the embodiments described herein can be carried out in other than the order discussed herein without departing from the scope of the application. Further, the terms "comprise" and "have" and any variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises a list of steps or modules not only comprises those steps or modules but can also comprise other steps or modules not expressly listed or inherent to such process, method, system, product or apparatus. The naming or numbering of steps appearing in the present application does not mean that the steps must be executed in the time / logical order indicated by the naming or numbering, and the execution order of the steps that have been named or numbered can be changed according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.

[0042] The division of modules appearing in the present application is a logical division, and in actual application, there can be another division manner, for example, multiple modules can be combined or integrated in another system, or some features can be ignored or not executed, in addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be through some interfaces, the indirect coupling or communication connection between the modules can be electrical or other similar forms, which are not limited in the present application. Moreover, the modules or sub-modules described as separate components can or can not be physically separate, can or can not be physical modules, or can be distributed into multiple circuit modules, and part or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme.

[0043] Before introducing the collaborative handling method based on the integration of multi-modal large model mechanical arm and AGV provided by the present application, first introduce the background content involved in the present application.

[0044] The multi-modal large model-based mechanical arm and AGV integrated collaborative handling method, device and computer readable storage medium provided by the application can be applied to a processing device, and is used for effectively solving the problems of low automation degree of the existing AGV handling strategy, high task execution time cost and low system response efficiency, effectively solving the problems of insufficient environmental perception and visual positioning accuracy of the existing AGV handling strategy, low positioning difficulty in complex or dynamic scenes and low target grasping success rate, effectively solving the problems of lack of efficient human-computer interaction mechanism of the existing AGV handling strategy, limited dynamic interaction capability and inability to flexibly adapt to changing working conditions, and effectively solving the problem of poor understanding of natural language instructions of the existing AGV handling strategy. Thus, the intelligent, efficient and accurate handling requirements can be met in actual application.

[0045] The multi-modal large model-based mechanical arm and AGV integrated collaborative handling method mentioned in the application can be a multi-modal large model-based mechanical arm and AGV integrated collaborative handling device, or a server, a physical host or a user equipment (UE) of different types of processing devices integrated with the multi-modal large model-based mechanical arm and AGV integrated collaborative handling device. The multi-modal large model-based mechanical arm and AGV integrated collaborative handling device can be realized in hardware or software, the UE can be a terminal device such as a smart phone, a tablet computer, a notebook computer, a desktop computer or a personal digital assistant (PDA), and the processing device can be set in the form of a device cluster.

[0046] Specifically, if the application scheme mainly provides a control center function service in actual application, the processing device executing the multi-modal large model-based mechanical arm and AGV integrated collaborative handling method or carrying the corresponding application service of the multi-modal large model-based mechanical arm and AGV integrated collaborative handling method has highly flexible specific device types and device deployment forms.

[0047] If the automatic guided vehicle is included in the category of the processing device, obviously, the processing device needs to be further adjusted in terms of software and hardware. For example, the processing device can include an automatic guided vehicle and a data processing part, and the data processing device can be further divided into a basic automatic guided vehicle control device and a data processing center for executing the data processing involved in the application scheme.

[0048] In addition, if the result display is involved, the processing device also needs to be configured with a corresponding display screen (touch screen), or connected with a corresponding display device or other device with a display screen, which can be flexibly configured as needed.

[0049] Next, the method for collaborative handling of the integrated multi-modal large model robot arm and AGV provided by the present application will be introduced.

[0050] First, refer to Figure 1 , Figure 1 The present application shows a flowchart of the method for collaborative handling of the integrated multi-modal large model robot arm and AGV. The method for collaborative handling of the integrated multi-modal large model robot arm and AGV provided by the present application can specifically include the following steps S101 to S10:

[0051] Step S101, the automatic guided vehicle constructs an environment map of the preset running environment in real time during the driving process of the preset running environment;

[0052] It can be understood that this step mainly uses SLAM technology to build a map of the current running environment or running area.

[0053] In specific operation, the automatic guided vehicle (AGV) usually only needs to drive once, and the laser radar, camera and other types of sensors configured on the automatic guided vehicle can be used to complete the collection of environmental scanning data, and then real-time SLAM mapping is performed.

[0054] As an example, the automatic guided vehicle can use a SLAM laser navigation AGV car, the car model is Weizhan S1-H, the self-weight is 67 kg, the load is 60 kg, and the running mode supports free navigation and straight-line navigation.

[0055] In addition, it can be understood that the present application is described from the overall level, and the SLAM mapping of the environment map is usually only done once at the initial stage when the present application is applied in the current running environment. Of course, it is not ruled out that in some cases, a mapping process is performed at a preset time point, at a preset time interval, randomly selected work tasks, specified work tasks, or each work task, which can be adaptively configured as needed.

[0056] Further, as an exemplary embodiment, the automatic guided vehicle constructs an environment map of the preset running environment in real time during the driving process of the preset running environment, which can specifically include:

[0057] 1.1) Control the automated guided vehicle to follow the initial path within the preset operating environment. During the lap, lidar data was continuously collected. With camera image data To form a set of environmental characteristic information , ;

[0058] As you can see, this also involves the initialization path. The path can be configured manually (including manually preset paths or manually planned paths in real time), or it can be generated in advance or in real time by the system according to the corresponding autonomous generation strategy.

[0059] 1.2) Set up environmental characteristic information Time synchronization and spatial registration are performed to generate a continuous sequence of sensor observations. , ,in, This is a sensor data synchronization and registration function used to ensure the consistency of multi-source data under the same time axis and coordinate system;

[0060] As can be seen, this involves temporal and spatial alignment / registration processes to form a standardized time series, i.e., a continuous sequence of sensor observations. This will facilitate further processing.

[0061] 1.3) Sensor observation sequence Input SLAM mapping module Output the current pose estimate Environmental Map , ,in, For automated guided vehicles in Pose information at any given moment SLAM mapping module Output a local or global environment map;

[0062] Understandably, this specifically involves the application of SLAM technology, which uses the previously obtained continuous sensor observation sequences... As input to the algorithm, it is fed into the SLAM algorithm and the SLAM mapping module. The system performs mapping processing to achieve simultaneous localization and mapping, obtaining information on both pose and map, i.e., current pose estimation. Environmental Map .

[0063] In addition, it can be noted that in the specific mapping process, both local mapping level processing and overall mapping level processing can be involved, and the specific processing scale can be configured as needed.

[0064] 1.4) Post-processing of the environment map to obtain the running environment map , wherein the post-processing includes map boundary clipping, noise filtering, and coordinate standardization for subsequent path navigation.

[0065] After obtaining the environment map through SLAM mapping processing, further post-processing or secondary optimization processing can be performed to form an environment map with higher map quality (mainly in terms of precision and structural clarity) and better usability through specific operations such as map boundary clipping, noise filtering, and coordinate standardization.

[0066] Step S102: Select a running route passing through the starting point and the target point according to the environment map.

[0067] It can be understood that, in the case of obtaining the environment map obtained through mapping processing, the starting point (starting position) and the target point (target position) can be selected in the environment map according to the current object carrying requirements, and then the running route passing through the starting point and the target point or the running route from the starting point to the target point can be selected.

[0068] Here, the running route planning process can be either manual selection or system autonomous planning, depending on the application requirements.

[0069] Preferably, the starting point and the target point can be manually selected, and the system can autonomously plan the specific driving route corresponding to the route planning goals of the shortest route, the shortest time, the least obstacles, or the least unexpected routes.

[0070] As an example, the starting point and the target point of the automated carrying task are selected by the user or the system automatically, S is the starting position of the task, such as a material storage area or a loading point, is the destination position of the task, such as a processing station, an assembly area, or a delivery outlet. , the system retrieves and selects a running route from to , which should avoid obstacles and meet the safety and efficiency requirements of path planning. The position of the target point can be flexibly set in the environment map through the system interface according to the actual needs of specific operation tasks, so as to flexibly support various carrying task scenes.

[0071] In step S103, on the basis of the structured analysis of the running route, the automatic guided vehicle is autonomously navigated from the starting point to the target point.

[0072] In the case where the running route is planned according to the current situation, the specific vehicle driving control can be promoted, and in this process, on the basis of the structured analysis of the running route, the local path correction and obstacle avoidance control can be performed in combination with the continuous updating of the positioning in the environment map and the real-time sensor data, so that the automatic guided vehicle is autonomously navigated from the starting point to the target point.

[0073] It can be understood that this is a mature field of vehicle control, and existing related processing algorithms can generally be used, of course, it is not excluded that the algorithm obtained by referring to the existing algorithm for further optimization or the novel algorithm researched by oneself is used when applying the scheme of the present application.

[0074] In step S104, the parameters of the mechanical arm driven by the multi-modal large model are initialized for the environment where the target point is located.

[0075] It can be understood that when the automatic guided vehicle arrives at the destination point of the current running route, the object grabbing process needs to be continued for the object to be carried at this position.

[0076] According to the current situation, the specific parameters of the mechanical arm driven by the multi-modal large model specially configured by the present application can be initialized for the mechanical arm deployed on the automatic guided vehicle and the environment (within a preset range) where the target point is located at this time.

[0077] It can be seen that the initialized parameters can specifically relate to the two aspects of mechanical arm parameters and mechanical arm positions, and in addition, other specific parameter types can also be involved as the case may be.

[0078] As an example, the mechanical arm involved in the present application can use MyArm-2 mechanical arm, model MV4 H7 Plus, equipped with STM32H7 processor with a frequency of 480MHZ, and a 5 million pixel camera. The multi-modal large model used in the mechanical arm can be Qwen-vl-max visual language large model, which can be used in a local model mode or an external calling mode, that is, a third-party trusted application programming interface can be called to perform the model inference required by the scheme of the present application.

[0079] Step S105: Using the camera mounted on the robotic arm, take pictures of the environment of the object to be moved at the target point to obtain an environmental image;

[0080] Understandably, in addition to cameras that provide the scanned images needed for path planning, automated guided transport units can also be equipped with cameras that provide the images needed for robotic arm control, specifically mounted on the robotic arm.

[0081] In this case, the camera is specifically used to capture the surrounding environment of the object to be transported at the current target point. The resulting environmental image contains the image content of the subsequent multimodal analysis and processing carried out by the multimodal large model.

[0082] As an example, it can be specifically utilized by mounting it on a robotic arm. The camera on the target point The camera performs image acquisition operations on the environment where the object being transported is located, and is mounted on the robotic arm. Under the control of the camera, the angle is flexibly adjusted until the lens is parallel to the plane of the object to ensure clear visual information is obtained. Through this shooting operation, an image data reflecting the target object is obtained, which can be recorded as an image. The data is then saved to the robotic arm's memory for use in subsequent programs.

[0083] Step S106: Combine the cross-attention mechanism algorithm and the name of the object to be moved to perform data augmentation on the environmental image to obtain an enhanced environmental image.

[0084] Understandably, the environmental images captured by the camera on the robotic arm can be further enhanced (which can be understood as a kind of preprocessing) in the design of this application to improve their image quality.

[0085] In practice, in addition to the cross-attention mechanism algorithm specially configured in this application, it is also necessary to combine the name of the object to be moved (an additional input to the algorithm) to achieve a highly targeted and fine-grained image enhancement effect.

[0086] Among them, the cross-attention mechanism algorithm can be understood as an attention mechanism that calculates the dependency relationship between two different input sequences under the condition of information interaction between two sequences, so as to achieve subtle image enhancement.

[0087] Specifically, as an exemplary embodiment, this method combines a cross-attention mechanism algorithm with the name of the object to be moved to perform data augmentation on the image, resulting in an enhanced environmental image. This may specifically include:

[0088] 6.1) Environmental images and a name text of a desired object of carrying selected by the user Both are input into a multi-modal model In which, an environmental image is extracted and the name text Corresponding cross-attention map representation , Wherein, the multi-modal model is used to align the spatial and semantic feature information of the input, the cross-attention map representation The size of the rows and columns of the cross-attention map representation is the number of rows and columns after the environmental image is divided into fixed-size blocks.

[0089] As can be seen, the use of a multi-modal model is specifically involved here, and the output thereof corresponds to the image form of the input environmental image , and the cross-attention processing result is in the form of a graph, and the size of the rows and columns of the cross-attention map representation is specifically the number of rows and columns after the environmental image is divided into fixed-size blocks.

[0090] 6.2) The cross-attention map representation is used as a weight to generate a mask for the environmental image through interpolation , Wherein, is the number of columns of the environmental image , is the number of rows of the environmental image , represents the conversion of the matrix size of the cross-attention map representation to and using an interpolation method;

[0091] After obtaining the cross-attention processing result, i.e. , it can be used as a weight factor that can be focused on in the interpolation process, and interpolation is used to generate a mask on the basis of the environmental image , which can be understood as a reference information used to promote specific image enhancement operations.

[0092] As can be seen, in the processing here, it is specifically developed from the number characteristics of the matrix rows / columns.

[0093] 6.3) The mask is fused with the environmental image to generate an enhanced environmental image , wherein, is the selected image minimum visibility preventing complete occlusion, preferably 15%.

[0094] It can be understood that, after the mask image is obtained , the key reference information required by this specific data enhancement operation can be fused with the environment image itself to generate an enhanced environment image .

[0095] In addition, it can be seen that, starting from the specific quantitative formula, a more specific fusion scheme is given to achieve a more convenient application and better enhanced image enhancement goal.

[0096] Step S107, input the enhanced environment image and the predefined system prompt word into the multi-modal large model to obtain the pixel position coordinates of the required object to be carried;

[0097] After obtaining the enhanced environment image, it can be easily understood that the multi-modal large model mentioned in step S104 can be used for specific image analysis to obtain the pixel position coordinates of the required object to be carried in the pixel coordinate system under the current situation.

[0098] The model input of the multi-modal large model also involves a predefined system prompt word, which can be understood as a prompt word for image object pixel coordinate positioning designed by the present application according to the context instruction learning (In-context Instruction Learning) carefully. Through the associated information of the context, it guides the multi-modal large model to identify and locate the target object.

[0099] Further, as an exemplary embodiment, inputting the enhanced environment image and the predefined system prompt word into the multi-modal large model to obtain the pixel position coordinates of the required object to be carried can specifically include:

[0100] 7.1) according to the enhanced environment image , joint the predefined system prompt word , construct an image-text pair input , wherein, the predefined system prompt word is a prompt word for image object pixel coordinate positioning designed according to the context instruction learning;

[0101] It can be understood that this link can be understood as inputting the enhanced environment image and the predefined system prompt word , do a simplified dataset configuration process to facilitate subsequent multimodal large model .

[0102] 7.2) input the image-text pair into the multimodal large model , to obtain the pixel position coordinates of the target object in the enhanced environment image . , , wherein the multimodal large model with image-language joint inference capability , internally contains a visual encoder, a language encoder and a fusion module;

[0103] For this multimodal large model , it internally contains a visual encoder, a language encoder and a fusion module, so as to infer the pixel position of the target object in the image from the joint information of image and text.

[0104] In this way, by inputting the image-text pair into the multimodal large model , it is converted or mapped into the pixel position coordinates of the target object in the enhanced environment image .

[0105] 7.3) range rationality check is performed on the pixel position coordinates , if it does not meet the image pixel range, the multimodal large model is called again to generate the pixel position coordinates.

[0106] After obtaining the pixel position coordinates of the target object in the enhanced environment image by the multimodal large model , it can be understood that the present application is based on insurance purposes and continues to configure the range rationality check mechanism at the detail level, so as to screen out the rare case that the pixel position coordinates do not meet the image pixel range after range rationality check, to promote the multimodal large model to reprocess, eliminate unexpected situations and stably output high-quality pixel position coordinates .

[0107] Further, as an exemplary embodiment, in the range rationality check, for the environment image , i.e. the image taken by the camera mounted on the mechanical arm, the pixel size is , then the pixel position coordinates can be specifically designed to satisfy the following inequality relationship:​​​​

[0108] ,

[0109] .

[0110] Thus, by the specific horizontal and vertical coordinate threshold limit, the range rationality check is efficiently and accurately completed.

[0111] As an example, if the pixel size of the environment image is specifically 480x720, the inequality relationship that needs to be satisfied is:

[0112] ,

[0113] .

[0114] Step S108, the pixel position coordinates are converted into the relative position coordinates of the robot in the physical world coordinate system by using the hand-eye calibration algorithm.

[0115] It can be understood that the hand-eye calibration algorithm is a calibration algorithm specially configured for the camera / camera deployed at the end position of the robot, which is used to calibrate between the pixel coordinate system and the physical world coordinate system of the robot.

[0116] Thus, the pixel coordinates obtained in the foregoing, i.e. , can be converted into relative position coordinates.

[0117] Further, the pixel position coordinates are converted into the relative position coordinates of the robot in the physical world by using the hand-eye calibration algorithm, which specifically can include:

[0118] 8.1) the mapping model of the pixel and the physical coordinate is established by using the robot parameter in the initialized parameter, and the affine transformation model is obtained through the offline hand-eye calibration process , which is used to map the pixel position in the image coordinate system to the position in the physical coordinate system in the working plane of the robot ,

[0119] It can be noted that the robot parameter initialized in the foregoing step S104 is applied, the mapping model of the pixel and the physical coordinate is established by using the parameter, and the affine transformation model is obtained through the offline hand-eye calibration process, which is used for subsequent conversion of the pixel position coordinates .

[0120] 8.2) the pixel position coordinates Input to the affine transformation model , output corresponding to the relative position coordinates in the physical world coordinate system , , wherein, is the conversion function corresponding to the affine transformation model .

[0121] As introduced before, the affine transformation model adapted to the current situation is constructed , then the pixel position coordinates required to be converted can be input, and the corresponding relative position coordinates in the physical world coordinate system are obtained.

[0122] Step S109, according to the relative position coordinates, the action execution module of the robot arm is called to grasp the object to be carried at the specified position, and then the object to be carried is placed.

[0123] After obtaining the relative position coordinates of the current object to be carried in the physical world coordinate system, it can be understood that the specific object grasping and placing can be promoted based on the relative position coordinates, which involves the calling of the action execution module of the robot arm (which can be understood as a conventional hardware configuration for grasping objects by the robot arm).

[0124] Further, as an exemplary embodiment herein, according to the relative position coordinates, the action execution module of the robot arm is called to grasp the object to be carried at the specified position, and then the object to be carried is placed, which can specifically include:

[0125] 9.1) According to the relative position coordinates , combined with the depth information of the object to be carried, the action instruction sequence for controlling the motion of the end effector of the robot arm is generated, wherein, represents that the robot arm moves to the set grasping position, represents that the gripper is closed to complete the object grasping;

[0126] It can be noted that for the camera configured on the robot arm, which is a depth camera, depth information can also be collected, so that in the setting of the embodiment herein, the action instruction sequence can be generated in combination with the relative position coordinates obtained before.

[0127] 9.2) According to the action instruction sequence , the action execution module of the robot arm is called to execute the action instruction sequence in sequence to complete the grasping process of the target object, and wherein, indicates whether the grasping operation is successful, confirmed by the end effector feedback signal;

[0128] In the case where the action instruction sequence is generated, the grasping of the current object, i.e. the object to be carried, can be completed by using and invoking the action instruction sequence.

[0129] 9.3) According to the initial robot arm position parameter in the previously initialized parameters, the placement action instruction sequence is constructed, wherein, indicates that the robot arm is transferred back to the initial position, indicates that the gripper is opened to complete the object placement;

[0130] In a general sense, the entire collaborative carrying process of the object, as mentioned herein, also involves the object placement action.

[0131] Correspondingly, the generation process of the placement action instruction sequence can be performed, and it can be noted that the specific robot arm position parameter of the initial robot arm position parameter initialized in the previous step S104 is also applied here.

[0132] 9.4) According to the placement action instruction sequence , the action execution module of the robot arm is invoked to place the object to be carried at the target position and release it, and wherein, indicates whether the placement operation is successful.

[0133] Similar to the previous action instruction sequence , in the case where the placement action instruction sequence is generated, the placement of the current object, i.e. the object to be carried, can be completed by using and invoking the action instruction sequence.

[0134] Step S1010, the automated guided vehicle is returned to the starting point according to the running route and autonomously navigates.

[0135] After completing the carrying of the current object to be carried, the automated guided vehicle can continue to autonomously navigate back to the starting point according to the previously planned running route, thereby completing a collaborative carrying task of the object.

[0136] It can be understood that, similar to the previous destination, during the navigation process, the automated guided vehicle also combines environmental perception and path tracking algorithms to avoid obstacles and maintain stable driving, ensuring that the carrying task is completed safely and efficiently.

[0137] In addition, it is easy to understand that the system can also perform real-time monitoring, dynamic evidence storage and post-tracing for the processing dynamics involved in the above collaborative handling process to meet diversified application requirements.

[0138] And specifically corresponding to some cases of staff-oriented application processing, the method of the present application can also involve corresponding real-time monitoring result display, dynamic evidence storage result display and post-tracing result display, and among them also involves corresponding human-computer interaction to further control, such as specific display content / object adjustment.

[0139] Finally, for the above scheme content, in general, for the object handling target based on the automated guided vehicle, the present application effectively solves the problems of low automation degree of existing automated guided vehicle handling strategy, which leads to high time cost of task execution and low system response efficiency, and effectively solves the problems of insufficient environmental perception and visual positioning accuracy of existing automated guided vehicle handling strategy, which leads to positioning difficulty in complex or dynamic scenes and low target grasping success rate, and effectively solves the problems of lack of efficient human-computer interaction mechanism of existing automated guided vehicle handling strategy, limited dynamic interaction ability, and inability to flexibly adapt to changing working conditions, and effectively solves the problem of poor understanding ability of existing automated guided vehicle handling strategy for natural language instructions of users, so as to meet the intelligent, efficient and accurate handling requirements in actual application.

[0140] Specifically, in the above scheme setting, the following more detailed effect explanation is also provided:

[0141] 1) Since the above steps S101 to S103 construct a map in the running environment, accept user-defined paths, and combine real-time positioning technology, not only high-precision perception and mapping of the working environment are realized, but also handling paths can be flexibly set according to task requirements, and through dynamic obstacle avoidance control, the autonomous navigation ability of the automated guided vehicle in complex environments is improved, thereby effectively solving the technical problem of poor autonomous navigation ability in handling tasks;

[0142] 2) Since step S106 extracts object feature information in the image through the cross-attention mechanism algorithm, it can effectively improve the visual positioning accuracy of the multi-modal large model, so as to more accurately locate the pixel position of the object, effectively solving the technical problem that existing handling robots cannot accurately position the environmental objects;

[0143] 3) Since steps S107 and S108 enable the multimodal large model to understand the visual information of the target object in the image through contextual instruction learning, the robot arm can accurately obtain the relative position of the target object in the physical world, thereby effectively solving the technical problem that existing transport robots are difficult to achieve high-precision object grasping in the physical world;

[0144] 4) Since step S109 enables the robot arm to sequentially complete the specified object grasping task by analyzing the mechanical instructions such as grasping and placing of the robot arm, the technical problem that existing transport robots are difficult to achieve automatic transport tasks after reaching the destination is effectively solved.

[0145] The above is an introduction to the collaborative transport method based on the integration of the multimodal large model robot arm and AGV provided by the present application. In order to better implement the collaborative transport method based on the integration of the multimodal large model robot arm and AGV provided by the present application, the present application also provides a collaborative transport device based on the integration of the multimodal large model robot arm and AGV from the functional module perspective.

[0146] Specifically, in the present application, the collaborative transport device based on the integration of the multimodal large model robot arm and AGV can specifically include the following structure:

[0147] The mapping unit is configured to construct an environment map of the preset running environment in real time by the AGV during the driving process of the preset running environment;

[0148] The selection unit is configured to select a running route passing through the starting point and the target point according to the environment map;

[0149] The navigation unit is configured to enable the AGV to autonomously navigate from the starting point to the target point on the basis of the structured analysis of the running route;

[0150] The initialization unit is configured to initialize the parameters of the robot arm driven by the multimodal large model for the environment where the target point is located;

[0151] The shooting unit is configured to use the camera mounted on the robot arm to shoot the environment of the object to be transported at the target point, and obtain an environment image;

[0152] The enhancement unit is configured to perform data enhancement on the environment image in combination with the cross-attention mechanism algorithm and the object name of the object to be transported, and obtain an enhanced environment image;

[0153] The positioning unit is configured to input the enhanced environment image and the predefined system prompt word into the multimodal large model to obtain the pixel position coordinates of the object to be transported;

[0154] A conversion unit is configured to convert the pixel position coordinates into relative position coordinates of the robot arm in a physical world coordinate system by using a hand-eye calibration algorithm.

[0155] A pick-and-place unit is configured to call an action execution module of the robot arm to a specified position to pick up a desired object to be carried according to the relative position coordinates, and then place the desired object to be carried.

[0156] A return unit is configured to let the automated guided vehicle autonomously navigate back to the starting point according to the running route.

[0157] As an exemplary embodiment, the mapping unit is specifically configured to:

[0158] control the automated guided vehicle to run along the initialized path in the preset running environment for one round, and continuously collect laser radar data in the process and camera image data to form an environment feature information set , ;

[0159] synchronize the environment feature information set in time and space, and generate a continuous sensor observation sequence , wherein, is a sensor data synchronization and registration function;

[0160] input the sensor observation sequence into a SLAM mapping module to output a current pose estimate and an environment map , wherein, is the pose information of the automated guided vehicle at moment, is a local environment map or a global environment map output by the SLAM mapping module ;

[0161] post-process the environment map to obtain a running environment map , wherein the post-processing includes map boundary clipping, noise filtering and coordinate standardization, for subsequent path navigation.

[0162] As another exemplary embodiment, the enhancement unit is specifically configured to:

[0163] input the environment image and the name text of the desired object to be carried selected by the user into a multi-modal model to extract environment image and name text The corresponding cross-attention map representation , Among them, multimodal model Cross-attention map representation is used to align spatial and semantic feature information of the input. The size of the rows and columns is the environmental image. The number of rows and columns after dividing into fixed-size blocks;

[0164] Representing the cross-attention map As weights, the environmental image is generated through interpolation. Masks for data augmentation , ,in, For environmental images The number of columns, For environmental images The number of rows, This indicates that the cross-attention map is represented using an interpolation method. Convert the matrix size to and ;

[0165] cover face With environmental images Fusion to generate enhanced environment images , ,in, Set the minimum visibility for the selected image.

[0166] As another exemplary embodiment, the positioning unit is specifically used for:

[0167] Based on enhanced environmental images Combined with predefined system prompts Construct image-text pairs for input , Among them, predefined system prompt words Prompt words designed for locating pixel coordinates of objects in an image, based on contextual instructions;

[0168] Input image-text pairs Input up to a multimodal large model To obtain the image of the object to be moved in the enhanced environment. pixel position coordinates , Among them, the multimodal large model with joint reasoning capabilities of image and language It contains a visual encoder, a language encoder, and a fusion module.

[0169] pixel position coordinates range rationality check is performed, and if the image pixel range is not met, the multi-modal large model is called again The pixel position coordinate generation process is re-performed.

[0170] As another exemplary embodiment, in the range rationality check, for the environmental image , assuming the pixel size is , the pixel position coordinates satisfy the following inequality relationship:

[0171] ,

[0172] .

[0173] As another exemplary embodiment, the conversion unit is specifically used for:

[0174] Using the robot arm parameters in the previous initialization parameters, a mapping model of pixels and physical coordinates is established, and through an offline hand-eye calibration process, an affine transformation model is obtained The affine transformation model is used to map the pixel position in the image coordinate system to the position in the physical coordinate system in the robot arm working plane ;

[0175] The pixel position coordinates are input into the affine transformation model , and the corresponding relative position coordinates in the physical world coordinate system are output , wherein is the conversion function corresponding to the affine transformation model .

[0176] As another exemplary embodiment, the grasping and placing unit is specifically used for:

[0177] According to the relative position coordinates , combined with the depth information of the object to be carried , a motion instruction sequence for controlling the motion of the robot arm end effector is generated , , wherein represents that the robot arm moves to the set grasping position, represents that the gripper is closed to complete the object grasping;

[0178] According to the motion instruction sequence , the motion execution module of the robot arm is called to execute the motion instruction sequence in sequence to complete the grasping process of the target object, and have wherein, indicates whether the grasping operation is successful, confirmed by the end feedback signal;

[0179] According to the initial mechanical arm position parameter in the previously initialized parameter , a placing action instruction sequence is constructed , wherein, indicates that the mechanical arm is transferred back to the starting position, indicates that the gripper is opened to complete the object placing;

[0180] According to the placing action instruction sequence , the action execution module of the mechanical arm is called to place the required carrying object at the target position and release, and have wherein, indicates whether the placing operation is successful.

[0181] The application also provides a processing device from the hardware structure, specifically, the processing device of the application can include a processor, a memory and an input and output device (if a device cluster is involved, the devices in the device cluster are also configured according to a single processing device), the processor is used to execute the computer program stored in the memory to realize the functions of Figure 1 the steps of the collaborative carrying method based on the integrated multi-modal large model mechanical arm and AGV in the corresponding embodiments; or the processor is used to execute the computer program stored in the memory to realize the functions of the units in the above embodiments, the memory is used to store the computer program required by the collaborative carrying method based on the integrated multi-modal large model mechanical arm and AGV in the corresponding embodiments. Figure 1

[0182] For example, the computer program can be divided into one or more modules / units, one or more modules / units are stored in the memory and executed by the processor to complete the application. One or more modules / units can be a series of computer program instructions that can complete a specific function, which is used to describe the execution process of the computer program in the computer device.

[0183] The processing device can include, but is not limited to, a processor, a memory, an input and output device. Those skilled in the art can understand that the schematic diagram is only an example of the processing device and does not constitute a limitation on the processing device, which can include more or fewer components than the schematic diagram, or combine certain components, or different components, for example, the processing device can also include a network access device, a bus, etc., and the processor, the memory, the input and output device are connected through the bus. ​

[0184] A processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the processing device, connecting all parts of the device through various interfaces and lines.

[0185] Memory can be used to store computer programs and / or modules. The processor performs various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can primarily include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the processing device, etc. Furthermore, memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0186] When a processor executes a computer program stored in memory, it can specifically perform the following functions:

[0187] The automated guided vehicle (AGV) constructs an environmental map of the preset operating environment in real time as it travels through the preset operating environment.

[0188] Select the route that passes through the starting point and the destination point based on the environmental map;

[0189] Based on the structured analysis of the route, the automated guided vehicle (AGV) can autonomously navigate from the starting point to the target point.

[0190] The parameters of the robotic arm, driven by a multimodal large model, are initialized for the environment of the target point. These parameters include the robotic arm parameters and its position.

[0191] Using a camera mounted on a robotic arm, the environment of the target point where the object needs to be moved is photographed to obtain an environmental image;

[0192] In combination with the cross-attention mechanism algorithm and the object name of the required carrying object, data augmentation is performed on the environment image to obtain an augmented environment image.

[0193] The augmented environment image and the predefined system prompt word are input into the multi-modal large model to obtain the pixel position coordinates of the required carrying object.

[0194] The pixel position coordinates are converted into relative position coordinates of the robot in the physical world coordinate system by using a hand-eye calibration algorithm.

[0195] According to the relative position coordinates, the action execution module of the robot is called to the specified position to grasp the required carrying object, and then the required carrying object is placed.

[0196] The automated guided vehicle is navigated back to the starting point according to the running route.

[0197] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described multi-modal large model-based robot and AGV integrated collaborative carrying device, processing device and corresponding units can be referred to as Figure 1 The description of the multi-modal large model-based robot and AGV integrated collaborative carrying method in the corresponding embodiments will not be repeated here.

[0198] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by related hardware controlled by instructions, which can be stored in a computer readable storage medium and loaded and executed by a processor.

[0199] To this end, the present application provides a computer readable storage medium, which stores a plurality of instructions capable of being loaded by a processor to execute the present application as Figure 1 The steps of the multi-modal large model-based robot and AGV integrated collaborative carrying method in the corresponding embodiments can be referred to as Figure 1 The description of the multi-modal large model-based robot and AGV integrated collaborative carrying method in the corresponding embodiments will not be repeated here.

[0200] The computer readable storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0201] Due to the instructions stored in the computer readable storage medium, the present application as Figure 1Corresponding to the steps of the collaborative handling method of the multi-modal large model mechanical arm and AGV integration in the embodiment, therefore, the application can be realized as Figure 1 Corresponding to the beneficial effects that can be achieved by the collaborative handling method of the multi-modal large model mechanical arm and AGV integration in the embodiment, details are given in the foregoing description, which will not be repeated here.

[0202] The above provides a collaborative handling method, device, processing equipment and computer readable storage medium based on the integration of multi-modal large model mechanical arm and AGV, which are described in detail in this application. The principles and implementation methods of the application are described in this paper. The above description of the embodiments is only used to help understand the core idea of the application. At the same time, for those skilled in the art, according to the idea of the application, the specific implementation and application range will be changed. In view of the above, the content of the specification should not be understood as a limitation of the application.

Claims

1. A collaborative handling method based on the integration of a multimodal large-model robotic arm and an AGV, characterized in that, The method includes: During the operation of the automated guided vehicle in the preset operating environment, an environmental map of the preset operating environment is constructed in real time. Select a route that passes through the starting point and the target point based on the environmental map; Based on the structured analysis of the operating route, the automated guided vehicle (AGV) can autonomously navigate from the starting point to the target point. Initialize the parameters of the robotic arm driven by a multimodal large model for the environment of the target point; Using a camera mounted on the robotic arm, the environment of the object to be moved at the target point is photographed to obtain an environmental image; By combining the cross-attention mechanism algorithm and the name of the object to be moved, the environmental image is augmented to obtain an enhanced environmental image; The enhanced environment image and predefined system prompts are input into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The pixel position coordinates are converted into the relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; Based on the relative position coordinates, the motion execution module of the robotic arm is invoked to move to the designated position to grab the object to be transported, and then the object to be transported is placed. The automated guided vehicle (AGV) will autonomously navigate back to the starting point according to the established route. The method of combining the cross-attention mechanism algorithm and the name of the object to be moved to perform data augmentation on the image to obtain an enhanced environment image includes: Environmental images and the name text of the object to be moved selected by the user. Both are input into the multimodal model. Extract the environmental image. and the name text The corresponding cross-attention map representation , The multimodal model The cross-attention map is used to align the spatial and semantic features of the input. The size of the rows and columns is that of the environmental image. The number of rows and columns after dividing into fixed-size blocks; Represent the cross attention map As weights, the environmental image is generated through interpolation. Masks for data augmentation , ,in, For the environmental image The number of columns, For the environmental image The number of rows, This indicates that the cross-attention map is represented using an interpolation method. Convert the matrix size to and ; The mask With the environmental image Fusion to generate enhanced environment images , ,in, The lowest visibility level for the selected image; The step of inputting the enhanced environment image and predefined system prompts into the multimodal large model to obtain the pixel position coordinates of the object to be transported includes: Based on enhanced environmental images Combined with predefined system prompts Construct image-text pairs for input , The predefined system prompt words Prompt words designed for locating pixel coordinates of objects in an image, based on contextual instructions; Input the image-text pair Input up to a multimodal large model The desired transported object is obtained in the enhanced environment image. pixel position coordinates , Among them, the multimodal large model with joint reasoning capabilities of image and language It contains a visual encoder, a language encoder, and a fusion module. pixel position coordinates A range validity check is performed. If the range does not conform to the image pixel range, the multimodal large model is invoked again. Regenerate the pixel position coordinates.

2. The method according to claim 1, characterized in that, The process of constructing an environmental map of the preset operating environment in real time during the operation of the automated guided vehicle includes: The automated guided vehicle is controlled to follow the initial path within the preset operating environment. During the lap, lidar data was continuously collected. With camera image data To form a set of environmental characteristic information , ; The environmental feature information set Time synchronization and spatial registration are performed to generate a continuous sequence of sensor observations. , ,in, For sensor data synchronization and registration functions; The sensor observation sequence Input SLAM mapping module Output the current pose estimate Environmental Map , ,in, For the automated guided vehicle in Pose information at any given moment For the SLAM mapping module Output a local or global environment map; For the environmental map Post-processing is performed to obtain the runtime environment map. The post-processing includes map boundary clipping, noise filtering, and coordinate standardization for use in subsequent path navigation.

3. The method according to claim 1, characterized in that, In the range reasonableness check, for the environmental image Let the pixel size be Then the pixel position coordinates The following inequalities must be satisfied: , 。 4. The method according to claim 1, characterized in that, The step of converting the pixel position coordinates into the relative position coordinates of the robotic arm in the physical world using a hand-eye calibration algorithm includes: Using the robotic arm parameters from the previously initialized parameters, a mapping model between pixels and physical coordinates is established, and an affine transformation model is obtained through an offline hand-eye calibration process. The affine transformation model Used to determine the pixel position in the image coordinate system Mapped to position in the physical coordinate system below the working plane of the robotic arm , ; pixel position coordinates Input to the affine transformation model Output the relative position coordinates corresponding to the physical world coordinate system. , ,in, It corresponds to the affine transformation model. The conversion function.

5. The method according to claim 1, characterized in that, The step of calling the robotic arm's motion execution module to a designated position to grasp the object to be transported based on the relative position coordinates, and then placing the object to be transported, includes: Based on relative position coordinates Combined with the depth information of the object to be transported Generate a sequence of motion commands to control the movement of the robotic arm's end effector. , ,in, This indicates that the robotic arm has moved to the set grasping position. This indicates that the gripper is closed to complete the object grasping process; According to the action instruction sequence Call the motion execution module of the robotic arm. The sequence of action instructions is executed sequentially. To complete the process of grasping the target object, and has ,in, This indicates whether the grabbing operation was successful, confirmed by the terminal feedback signal; Based on the initial robotic arm position parameters from the previously initialized parameters. Construct a sequence of placement action instructions , ,in, This indicates that the robotic arm has returned to its starting position. This indicates that the grippers are opened to complete the object placement; According to the placement action instruction sequence Call the motion execution module of the robotic arm. Place the object to be transported at the target location and release it, and have ,in, This indicates whether the placement operation was successful.

6. A collaborative handling device based on the integration of a multimodal large-model robotic arm and an AGV, characterized in that, The device includes: The mapping unit is used to construct an environmental map of the preset operating environment in real time during the driving of the automated guided vehicle in the preset operating environment; The selection unit is used to select a route that passes through the starting point and the target point based on the environmental map. The navigation unit is used to enable the automated guided vehicle to autonomously navigate from the starting point to the target point based on the structured analysis of the running route; An initialization unit is used to initialize the parameters of a robotic arm driven by a multimodal large model for the environment of the target point, wherein the parameters of the robotic arm include robotic arm parameters and robotic arm position. The imaging unit is used to capture images of the environment of the object to be moved at the target point using a camera mounted on the robotic arm, thereby obtaining an environmental image. An enhancement unit is used to combine a cross-attention mechanism algorithm and the name of the object to be moved to perform data enhancement on the environmental image, thereby obtaining an enhanced environmental image. The positioning unit is used to input the enhanced environment image and predefined system prompts into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The conversion unit is used to convert the pixel position coordinates into the relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; The gripping and placing unit is used to call the motion execution module of the robotic arm to a designated position to grip the object to be transported according to the relative position coordinates, and then place the object to be transported. The return unit is used to enable the automated guided vehicle to autonomously navigate back to the starting point according to the operating route; The enhancement unit is specifically used for: Environmental images and the name text of the object to be moved selected by the user. Both are input into the multimodal model. Extract the environmental image. and the name text The corresponding cross-attention map representation , The multimodal model The cross-attention map is used to align the spatial and semantic features of the input. The size of the rows and columns is that of the environmental image. The number of rows and columns after dividing into fixed-size blocks; Represent the cross attention map As weights, the environmental image is generated through interpolation. Masks for data augmentation , ,in, For the environmental image The number of columns, For the environmental image The number of rows, This indicates that the cross-attention map is represented using an interpolation method. Convert the matrix size to and ; The mask With the environmental image Fusion to generate enhanced environment images , ,in, The lowest visibility level for the selected image; The positioning unit is specifically used for: Based on enhanced environmental images Combined with predefined system prompts Construct image-text pairs for input , The predefined system prompt words Prompt words designed for locating pixel coordinates of objects in an image, based on contextual instructions; Input the image-text pair Input up to a multimodal large model The desired transported object is obtained in the enhanced environment image. pixel position coordinates , Among them, the multimodal large model with joint reasoning capabilities of image and language It contains a visual encoder, a language encoder, and a fusion module. pixel position coordinates A range validity check is performed. If the range does not conform to the image pixel range, the multimodal large model is invoked again. Regenerate the pixel position coordinates.

7. A processing device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the method as described in any one of claims 1 to 5 when it invokes the computer program in the memory.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Clouded AGV application system of 5G smart factory

    CN112731914A

  • Path planning method and system based on four-AGV cooperative carrying

    CN120778132A