Collaborative carrying method based on integration of multi-modal large model mechanical arm and AGV (Automatic Guided Vehicle)
By integrating a multimodal large-scale robotic arm with an AGV for collaborative handling, the problems of low automation and insufficient environmental perception in traditional automated guided vehicle robots have been solved. This has enabled high-precision object positioning and grasping, improving handling efficiency and success rate.
Patent Information
- Application Number
- CN202511651186.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-11-12
AI Technical Summary
Traditional automated guided vehicles (AGVs) robots suffer from low automation, insufficient environmental perception and visual positioning accuracy, and a lack of efficient human-machine interaction capabilities during cargo handling. This results in difficulties in positioning, grasping failures, and high time consumption in complex or dynamic scenarios.
A collaborative handling method integrating a multimodal large-scale robotic arm and an AGV is adopted. By constructing an environmental map in real time and combining a cross-attention mechanism and a hand-eye calibration algorithm, high-precision object positioning and grasping are achieved, supporting autonomous navigation and human-computer interaction.
It improves the autonomous navigation capability of automated guided vehicles in complex environments, enhances visual positioning accuracy and grasping success rate, and meets the needs of intelligent, efficient and precise handling.
Smart Images

Figure CN121120764A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mobile mechanical operation, in particular to a collaborative carrying method based on integration of a multi-modal large model mechanical arm and an AGV. BACKGROUND
[0002] Robots have been widely used in various scenes in life due to the development of related technologies, especially the AGV robot (the robot generally adopts the abbreviation of the automated guided vehicle or AGV).
[0003] However, the present application inventors found that the traditional AGV robot generally relies on fixed algorithm mode in the process of realizing cargo carrying, and lacks environmental perception and flexible interaction ability.
[0004] Specifically, the carrying mode of the traditional AGV robot can be mainly divided into the following two kinds: (1) The AGV carrying strategy based on human cooperation, which realizes real-time mapping and positioning through SLAM (Simultaneous Localization And Mapping) technology and laser radar scanning of environmental features, combines dynamic path planning algorithm for autonomous navigation and obstacle avoidance, and intelligently schedules tasks by a central system. When the AGV arrives at the designated location, it reminds the user to manually put in or take out the carried object through voice or light prompting interaction, which is difficult to realize fully automated whole-process carrying. (2) The AGV carrying strategy based on traditional extension, which realizes the strategy of integrated operation of grabbing, carrying and assembling by fusing traditional mechanical arms with AGV, so that it upgrades from a single transportation device to an intelligent composite robot with operation ability. The traditional mechanical arms used in this strategy lack flexible strategy functions such as visual perception and instruction understanding, and are difficult to adapt to dynamic changes in the physical environment.
[0005] In summary, the above-mentioned traditional AGV carrying robot has the following limitations: (1) The AGV carrying strategy based on human cooperation has low automation level, and the AGV carrying strategy based on traditional extension has weak environmental understanding and interaction ability. Both strategies cannot realize highly automated carrying tasks in dynamic and complex environmental conditions. (2) Lack of automatic and high-precision visual positioning function for the carried objects, making it difficult to accurately and automatically grab the objects, so in the scenes with high repeatability and large environmental changes, there will be problems of grabbing failure and large time consumption. (3) Lacking the ability to understand users’ natural instructions, it is impossible to truly understand and execute tasks, relying only on preset fixed paths and manual intervention. Therefore, in complex and changing scenarios, it is necessary to collect data and set up programs again, which increases the complexity of tasks. Summary of the Invention
[0006] This application provides a collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV. By constructing a novel collaborative handling mechanism, it effectively solves the problems of low automation level in existing automated guided vehicle (AGV) handling strategies, which leads to high task execution time costs and low system response efficiency; it also effectively solves the problems of insufficient environmental perception and visual positioning accuracy in existing AGV handling strategies, which leads to difficulties in positioning and low target grasping success rate in complex or dynamic scenes; it also effectively solves the problems of lack of efficient human-computer interaction mechanism, limited dynamic interaction capability, and inability to flexibly adapt to changing working conditions in existing AGV handling strategies; and it effectively solves the problem of poor understanding of user natural language commands in existing AGV handling strategies. Thus, it can meet the needs of intelligent, efficient, and precise handling in practical applications.
[0007] Firstly, this application provides a collaborative handling method based on the integration of a multimodal large-model robotic arm and an AGV, the method comprising: The automated guided vehicle (AGV) constructs an environmental map of the preset operating environment in real time as it travels through the preset operating environment. Select the route that passes through the starting point and the destination point based on the environmental map; Based on the structured analysis of the route, the automated guided vehicle (AGV) can autonomously navigate from the starting point to the target point. Initialize the parameters of the robotic arm driven by a multimodal large model for the environment where the target point is located. The parameters of the robotic arm include the robotic arm parameters and the robotic arm position. Using a camera mounted on a robotic arm, the environment of the target point where the object needs to be moved is photographed to obtain an environmental image; By combining the cross-attention mechanism algorithm and the name of the object to be moved, data augmentation is performed on the environmental image to obtain an enhanced environmental image; The enhanced environmental image and predefined system prompts are input into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The pixel position coordinates are converted into the relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; Based on the relative position coordinates, the robot arm's motion execution module is invoked to move to the designated position to grab the object to be transported, and then the object to be transported is placed. The automated guided vehicle (AGV) will autonomously navigate back to its starting point based on the route it takes.
[0008] Secondly, this application provides a collaborative handling device based on the integration of a multimodal large-scale robotic arm and an AGV, the device comprising: The mapping unit is used to build an environmental map of the preset operating environment in real time during the driving of the automated guided vehicle in the preset operating environment; The selection unit is used to select a route that passes through the starting point and the destination point based on the environment map. The navigation unit is used to enable the automated guided vehicle to navigate autonomously from the starting point to the target point based on the structured analysis of the running route; The initialization unit is used to initialize the parameters of the robotic arm driven by a multimodal large model for the environment where the target point is located. The parameters of the robotic arm include the robotic arm parameters and the robotic arm position. The imaging unit is used to capture images of the environment of the object to be moved at the target point using a camera mounted on the robotic arm. The enhancement unit is used to combine the cross-attention mechanism algorithm and the name of the object to be moved to perform data enhancement on the environmental image, resulting in an enhanced environmental image. The localization unit is used to input the enhanced environment image and predefined system prompts into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The conversion unit is used to convert pixel position coordinates into relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; The gripping and placing unit is used to call the motion execution module of the robotic arm to the designated position to grip the object to be transported based on the relative position coordinates, and then place the object to be transported. The return unit is used to allow the automated guided vehicle to autonomously navigate back to the starting point according to the operating route.
[0009] Thirdly, this application provides a processing device, including a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the method provided in the first aspect of this application when it invokes the computer program in the memory.
[0010] Fourthly, this application provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the method provided in the first aspect of this application.
[0011] From the above, it can be concluded that this application has the following beneficial effects: For object handling targets based on automated guided vehicles (AGVs), this application constructs a novel collaborative handling mechanism. This mechanism effectively addresses the problems of low automation levels in existing AGV handling strategies, leading to high task execution time costs and low system response efficiency; insufficient environmental perception and visual positioning accuracy in existing AGV handling strategies, resulting in difficulties in positioning and low target acquisition success rates in complex or dynamic scenarios; lack of efficient human-computer interaction mechanisms, limited dynamic interaction capabilities, and inability to flexibly adapt to changing working conditions; and poor understanding of user natural language commands in existing AGV handling strategies. Therefore, in practical applications, this mechanism can meet the needs of intelligent, efficient, and precise handling. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of a collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV, as described in this application. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved.
[0016] The module division described in this application is a logical division. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the shown or discussed mutual coupling, direct coupling, or communication connection may be through some interface, and the indirect coupling or communication connection between modules may be electrical or other similar forms, none of which are limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules can be selected to achieve the purpose of the solution in this application according to actual needs.
[0017] Before introducing the collaborative handling method based on the integration of multimodal large-scale robotic arm and AGV provided in this application, we will first introduce the background content involved in this application.
[0018] The collaborative handling method, apparatus, and computer-readable storage medium based on the integration of a multimodal large-model robotic arm and AGV provided in this application can be applied to processing equipment. By constructing a novel collaborative handling mechanism, it effectively solves the problems of low automation level in existing automated guided vehicle (AGV) handling strategies, which leads to high task execution time costs and low system response efficiency; it also effectively solves the problems of insufficient environmental perception and visual positioning accuracy in existing AGV handling strategies, which leads to difficulties in positioning and low target grasping success rate in complex or dynamic scenes; it further solves the problems of lack of efficient human-machine interaction mechanism, limited dynamic interaction capability, and inability to flexibly adapt to changing working conditions in existing AGV handling strategies; and it effectively solves the problem of poor understanding of user natural language commands in existing AGV handling strategies. Thus, it can meet the intelligent, efficient, and precise handling requirements in practical applications.
[0019] The collaborative transport method based on the integration of a multimodal large-scale robotic arm and an AGV mentioned in this application can be executed by a collaborative transport device based on the integration of a multimodal large-scale robotic arm and an AGV, or by different types of processing devices such as servers, physical hosts, or user equipment (UEs) that integrate such a collaborative transport device. The collaborative transport device based on the integration of a multimodal large-scale robotic arm and an AGV can be implemented in hardware or software. The UE can specifically be a terminal device such as a smartphone, tablet, laptop, desktop computer, or personal digital assistant (PDA). The processing devices can be configured in a device cluster.
[0020] Specifically, if the solution of this application mainly provides the functional services of the control center in practical applications, then the specific equipment type and equipment deployment form of the processing equipment that executes the collaborative handling method based on the integration of multimodal large-scale robotic arm and AGV of this application, or the processing equipment equipped with the corresponding application services of the collaborative handling method based on the integration of multimodal large-scale robotic arm and AGV of this application, are highly flexible.
[0021] If automated guided vehicles (AGVs) are included in the category of processing equipment, then the processing equipment obviously needs to be further adapted in terms of both hardware and software. For example, the processing equipment may include AGVs and a data processing unit. The data processing unit may be further divided into a basic AGV control unit and a data processing hub that performs the data processing involved in the present application.
[0022] In addition, if results are to be displayed, the processing device also needs to be equipped with a corresponding display screen (control touch screen), or be connected to a corresponding display device or other devices with a display screen. These can be flexibly configured according to actual needs.
[0023] The following section introduces the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV, as provided in this application.
[0024] First, refer to Figure 1 , Figure 1 This paper illustrates a flowchart of a collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV, as described in this application. The collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV provided in this application may specifically include the following steps S101 to S10: Step S101: During the operation of the automated guided vehicle in the preset operating environment, an environmental map of the preset operating environment is constructed in real time. Understandably, this step mainly uses SLAM technology to perform mapping of the current operating environment or operating area.
[0025] In practice, the automated guided vehicle (AGV) usually only needs to make one trip or one loop to collect environmental scanning data through sensors such as LiDAR and cameras configured on the AGV, and then perform SLAM mapping in real time.
[0026] As an example, the automated guided vehicle can be a SLAM laser navigation AGV, model VIGAP S1-H, with a tare weight of 67 kg and a payload of 60 kg. It supports both free navigation and straight-line navigation.
[0027] Furthermore, it is understood that the solution in this application is explained from an overall perspective. The SLAM mapping process for the environment map is usually performed only once in the initial stage when the solution is applied in the current operating environment. Of course, it is not excluded that in some cases, mapping may be performed once under triggering strategies such as preset time points, preset time intervals, random sampling of work tasks, specified work tasks, or each work task. This can be adaptively configured according to actual needs.
[0028] Furthermore, as an exemplary embodiment, the automated guided vehicle (AGV) constructs an environmental map of the preset operating environment in real time during its journey within that environment. Specifically, this may include: 1.1) Control the automated guided vehicle to follow the initial path within the preset operating environment. During the lap, lidar data was continuously collected. With camera image data To form a set of environmental characteristic information , ; As you can see, this also involves the initialization path. The path can be configured manually (including manually preset paths or manually planned paths in real time), or it can be generated in advance or in real time by the system according to the corresponding autonomous generation strategy.
[0029] 1.2) Set up environmental characteristic information Time synchronization and spatial registration are performed to generate a continuous sequence of sensor observations. , ,in, This is a sensor data synchronization and registration function used to ensure the consistency of multi-source data under the same time axis and coordinate system; As can be seen, this involves temporal and spatial alignment / registration processes to form a standardized time series, i.e., a continuous sequence of sensor observations. This will facilitate further processing.
[0030] 1.3) Sensor observation sequence Input SLAM mapping module Output the current pose estimate Environmental Map , ,in, For automated guided vehicles in Pose information at any given moment SLAM mapping module Output a local or global environment map; Understandably, this specifically involves the application of SLAM technology, which uses the previously obtained continuous sensor observation sequences... As input to the algorithm, it is fed into the SLAM algorithm and the SLAM mapping module. The system performs mapping processing to achieve simultaneous localization and mapping, obtaining information on both pose and map, i.e., current pose estimation. Environmental Map .
[0031] Furthermore, it can be noted that the specific mapping process can involve both local mapping level processing and overall mapping level processing, and the specific processing scale can be configured according to actual needs.
[0032] 1.4) Environmental Map Post-processing is performed to obtain the runtime environment map. The post-processing includes map boundary clipping, noise filtering, and coordinate standardization for use in subsequent path navigation.
[0033] An environmental map was obtained through SLAM mapping. Afterwards, post-processing, or secondary optimization, can be carried out. Through specific operations such as map boundary clipping, noise filtering, and coordinate standardization, a higher quality (mainly in terms of accuracy and structural clarity) and more user-friendly environmental map can be formed.
[0034] Step S102: Select the route passing through the starting point and the target point based on the environmental map; Understandably, once the environmental map obtained from the mapping process is acquired, i.e., the runtime environment map... In this case, you can select the corresponding starting point (starting position) and target point (target position) in the environment map according to the transportation needs of the objects to be transported in the current situation, and then select the route that passes through the starting point and target point or travels from the starting point to the target point.
[0035] The route planning process here can be either manually selected or planned autonomously by the system, depending on the application requirements of the solution.
[0036] Preferably, the starting point and the destination point can be selected manually, and then the system can autonomously plan the specific driving route with the corresponding shortest route, shortest time, fewest obstacles, or fewest unexpected routes.
[0037] As an example, the starting point of the automated material handling task is specifically selected by the user or the system automatically. and target point S represents the starting point of the task, such as the material storage area or loading point. Given the destination location of the task, such as a processing station, assembly area, or shipping port, the system then searches the environment map and selects a route from... Departure and Arrival Operating route The route Obstacles should be avoided, and the path planning should meet the requirements of safety and efficiency. (Starting point) and target point The location can be flexibly set in the environment map through the system interface according to the actual needs of specific tasks, thus flexibly supporting a variety of handling task scenarios.
[0038] Step S103: Based on the structured analysis of the route, the automated guided vehicle (AGV) is enabled to autonomously navigate from the starting point to the target point. With the operational route planned according to the current situation, specific vehicle driving control can be implemented here. In this process, based on the structured analysis of the operational route, combined with continuous updates to its own location in the environmental map and real-time sensor data, local path correction and obstacle avoidance control can be performed, enabling the automated guided vehicle to autonomously navigate from the starting point to the target point.
[0039] It is understandable that this is a mature field of vehicle control, and existing related processing algorithms can usually be used. Of course, it is also possible that when applying the solution of this application, existing algorithms may be used to further optimize the algorithm or even a novel algorithm developed in-house.
[0040] Step S104: Initialize the parameters of the robotic arm driven by a multimodal large model for the environment where the target point is located; Understandably, when the automated guided vehicle reaches the destination of the current route, it needs to continue to grasp the object to be transported at that location.
[0041] In response to the current situation, the specific parameters of the robotic arm deployed on the automated guided vehicle can be initialized based on the multimodal large model specially configured in this application, corresponding to the environment of the target point (within the preset range).
[0042] As can be seen, the initialization parameters can specifically involve two main aspects: robotic arm parameters and robotic arm position. In addition, other specific parameter types may also be involved depending on the situation.
[0043] As an example, the robotic arm involved in this application can specifically be the MyArm-2 robotic arm, model MV4 H7 Plus, equipped with an STM32H7 processor with a main frequency of 480MHz and a 5-megapixel camera. The multimodal large model used in the robotic arm can specifically be the Qwen-vl-max visual language large model. In addition to using a local model method, it can also use an external calling method, that is, it can call a third-party trusted application programming interface to perform the model inference required by the solution of this application.
[0044] Step S105: Using the camera mounted on the robotic arm, take pictures of the environment of the object to be moved at the target point to obtain an environmental image; Understandably, in addition to cameras that provide the scanned images needed for path planning, automated guided transport units can also be equipped with cameras that provide the images needed for robotic arm control, specifically mounted on the robotic arm.
[0045] In this case, the camera is specifically used to capture the surrounding environment of the object to be transported at the current target point. The resulting environmental image contains the image content of the subsequent multimodal analysis and processing carried out by the multimodal large model.
[0046] As an example, it can be specifically utilized by mounting it on a robotic arm. The camera on the target point The camera performs image acquisition operations on the environment where the object being transported is located, and is mounted on the robotic arm. Under the control of the camera, the angle is flexibly adjusted until the lens is parallel to the plane of the object to ensure clear visual information is obtained. Through this shooting operation, an image data reflecting the target object is obtained, which can be recorded as an image. The data is then saved to the robotic arm's memory for use in subsequent programs.
[0047] Step S106: Combine the cross-attention mechanism algorithm and the name of the object to be moved to perform data augmentation on the environmental image to obtain an enhanced environmental image. Understandably, the environmental images captured by the camera on the robotic arm can be further enhanced (which can be understood as a kind of preprocessing) in the design of this application to improve their image quality.
[0048] In practice, in addition to the cross-attention mechanism algorithm specially configured in this application, it is also necessary to combine the name of the object to be moved (an additional input to the algorithm) to achieve a highly targeted and fine-grained image enhancement effect.
[0049] Among them, the cross-attention mechanism algorithm can be understood as an attention mechanism that calculates the dependency relationship between two different input sequences under the condition of information interaction between two sequences, so as to achieve subtle image enhancement.
[0050] Specifically, as an exemplary embodiment, this method combines a cross-attention mechanism algorithm with the name of the object to be moved to perform data augmentation on the image, resulting in an enhanced environmental image. This may specifically include: 6.1) Environmental images and the name text of the object to be moved selected by the user. Both are input into the multimodal model. Extracting environmental images and name text The corresponding cross-attention map representation , Among them, multimodal model Cross-attention map representation is used to align spatial and semantic feature information of the input. The size of the rows and columns is the environmental image. The number of rows and columns after dividing into fixed-size blocks; As can be seen, this specifically involves the use of a multimodal model, and its output corresponds to the input environmental image. The image format is a graph representation of the cross-attention processing result, and the cross-attention graph representation... The size of the rows and columns, specifically the environmental image. The number of rows and columns after dividing into blocks of fixed size.
[0051] 6.2) Represent the cross-attention map As weights, the environmental image is generated through interpolation. Masks for data augmentation , ,in, For environmental images The number of columns, For environmental images The number of rows, This indicates that the cross-attention map is represented using an interpolation method. Convert the matrix size to and ; After obtaining the cross-attention processing results, i.e. Then, it can be used as a key weighting factor in the interpolation process, and interpolation can be applied to the environmental image. Based on this, interpolation is used to generate the mask. The face covering This can be understood as a reference information used to advance specific image enhancement operations.
[0052] Furthermore, it can be seen that the processing here is specifically based on the quantitative characteristics of the matrix rows / columns.
[0053] 6.3) Cover your face With environmental images Fusion to generate enhanced environment images , ,in, The minimum visibility of the selected image is set to prevent complete obscuration, preferably 15%.
[0054] Understandably, after obtaining cover After obtaining the key reference information required for this specific data augmentation operation, it can be compared with the environmental image it is targeting. The images are fused together to generate an enhanced environmental image. .
[0055] Furthermore, it can be seen that, starting from the specific quantization formula, a more specific fusion scheme is given here to achieve the image enhancement goal that is easier to apply and has a better enhancement effect.
[0056] Step S107: Input the enhanced environment image and predefined system prompts into the multimodal large model to obtain the pixel position coordinates of the object to be transported; After obtaining the enhanced environmental image, it is easy to understand that the specific image analysis can be performed using the multimodal large model mentioned in step S104 to obtain the pixel position coordinates of the object to be transported in the pixel coordinate system under the current situation.
[0057] Among them, the model input of the multimodal large model also involves predefined system prompt words, which can be understood as prompt words carefully designed by this application based on in-context instruction learning for the localization of pixel coordinates of objects in images. Through the contextual information, the multimodal large model is guided to identify and locate the target object.
[0058] Furthermore, as an exemplary embodiment, the enhanced environmental image and predefined system prompts are input into the multimodal large model to obtain the pixel position coordinates of the object to be transported, which may specifically include: 7.1) Based on enhanced environmental images Combined with predefined system prompts Construct image-text pairs for input , Among them, predefined system prompt words Prompt words designed for locating pixel coordinates of objects in an image, based on contextual instructions; This step can be understood as enhancing the environmental image. and predefined system prompts We will perform simplified dataset configuration to facilitate subsequent multimodal large-scale models. Process it.
[0059] 7.2) Input image-text pairs Input up to a multimodal large model To obtain the image of the object to be moved in the enhanced environment. pixel position coordinates , Among them, the multimodal large model with joint reasoning capabilities of image and language It contains a visual encoder, a language encoder, and a fusion module. For this multimodal large model It contains a visual encoder, a speech encoder, and a fusion module, which enables it to infer the pixel position of a target object in an image from the combined information of the image and text.
[0060] Thus, through this multimodal large model Input image and text pairs Transformation, or mapping, into the enhanced environment image of the object to be moved. pixel position coordinates .
[0061] 7.3) Pixel position coordinates Perform a range validity check; if it does not conform to the image pixel range, then call the multimodal large model again. Regenerate the pixel position coordinates.
[0062] Through multimodal large model To obtain the required transported object in the enhanced environment image with high precision reasoning pixel position coordinates Subsequently, it is understandable that this application, for insurance purposes, continued to configure a range reasonableness check mechanism at a detailed level in order to filter out pixel position coordinates after the range reasonableness check was performed. This rare case, which does not conform to the image pixel range, drives the multimodal large model. The process was reprocessed to eliminate unexpected cases and ensure stable, high-quality output of pixel position coordinates. .
[0063] Furthermore, as an exemplary embodiment, in the scope reasonableness check, for environmental images That is, the image captured by a camera mounted on a robotic arm, with a pixel size of [missing information]. Then the pixel position coordinates Specifically, it can be designed to satisfy the following inequalities: , .
[0064] In this way, by setting specific thresholds for the horizontal and vertical axes, the range rationality check can be completed efficiently and accurately.
[0065] As an example, if the environment image If the pixel size is specifically 480×720, then the following inequalities need to be satisfied: , .
[0066] Step S108: Use a hand-eye calibration algorithm to convert the pixel position coordinates into the relative position coordinates of the robotic arm in the physical world coordinate system; Understandably, the hand-eye calibration algorithm is a calibration algorithm specifically configured for cameras / cameras deployed at the end of a robotic arm, used to perform calibration processing between the pixel coordinate system and the physical world coordinate system of the robotic arm.
[0067] Thus, the previously obtained pixel coordinates, i.e. This is converted into relative position coordinates.
[0068] Furthermore, this section utilizes a hand-eye calibration algorithm to convert pixel position coordinates into relative position coordinates of the robotic arm in the physical world, which can specifically include: 8.1) Using the robotic arm parameters from the previously initialized parameters, establish a mapping model between pixels and physical coordinates, and obtain an affine transformation model through an offline hand-eye calibration process. Affine transformation model Used to determine the pixel position in the image coordinate system Mapped to position in the physical coordinate system below the working plane of the robotic arm , ; It can be noted that the robotic arm parameters obtained in the previous step S104 are used here. Using these parameters, a mapping model between pixels and physical coordinates is established, and an affine transformation model is obtained through an offline hand-eye calibration process. Provides subsequent pixel position coordinates The conversion and use of.
[0069] 8.2) Pixel position coordinates Input to affine transformation model Outputs the relative position coordinates in the physical world coordinate system. , ,in, It is a corresponding affine transformation model The conversion function.
[0070] As introduced earlier, an affine transformation model adapted to the current situation has been constructed. Then, you can input the pixel coordinates to be converted. And obtain the relative position coordinates corresponding to the physical world coordinate system. .
[0071] Step S109: Based on the relative position coordinates, call the motion execution module of the robotic arm to the designated position to grab the object to be transported, and then place the object to be transported. After obtaining the relative position coordinates of the object to be moved in the physical world coordinate system, it is understandable that the specific object can be grasped and placed based on these relative position coordinates. This involves calling the motion execution module of the robotic arm (which can be understood as a conventional hardware configuration that uses a robotic arm to grasp objects).
[0072] Furthermore, as an exemplary embodiment here, based on relative position coordinates, the robotic arm's motion execution module is invoked to a designated position to grasp the object to be transported, and then the object is placed. Specifically, this may include: 9.1) Based on relative position coordinates Combined with the depth information of the object to be moved Generate a sequence of motion commands to control the movement of the robotic arm's end effector. , ,in, This indicates that the robotic arm has moved to the set gripping position. This indicates that the gripper is closed to complete the object grasping process; It can be noted that the camera mounted on the robotic arm, specifically a depth camera, can also collect depth information. Thus, in this embodiment, the relative position coordinates obtained from the previous processing can be combined to generate a sequence of motion commands. .
[0073] 9.2) Based on the sequence of action instructions Call the robotic arm's motion execution module Execute the sequence of action instructions sequentially. To complete the process of grasping the target object, and has ,in, This indicates whether the grabbing operation was successful, confirmed by the terminal feedback signal; After generating the action instruction sequence In this case, it can be used and invoked to complete the grabbing of the current object, i.e., the object to be moved.
[0074] 9.3) Based on the initial robotic arm position parameters from the previously initialized parameters. Construct a sequence of placement action instructions , ,in, This indicates that the robotic arm has returned to its starting position. This indicates that the grippers are open to complete the object placement; In layman's terms, the entire collaborative transportation process of an object, as mentioned here, also involves the action of placing the object.
[0075] Correspondingly, a sequence of placement action instructions can be created. The generation process also utilizes the initial robotic arm position parameters initialized in step S104. This is the specific position parameter of the robotic arm.
[0076] 9.4) According to the sequence of placement action instructions Call the robotic arm's motion execution module Place the object to be moved at the target location and release it, and have ,in, This indicates whether the placement operation was successful.
[0077] With the preceding sequence of action instructions Similarly, after generating the placement action instruction sequence In this case, it can be used and invoked to complete the placement of the current object, i.e., the object to be moved.
[0078] Step S1010: Complete the process of having the automated guided vehicle autonomously navigate back to the starting point according to the operating route.
[0079] After completing the transport of the object to be moved, the automated guided vehicle can continue to navigate back to the starting point according to the previously planned route, thus completing a collaborative object transport task.
[0080] Understandably, similar to the previous journey to the destination, during the navigation process, the automated guided vehicle will also combine environmental perception and path tracking algorithms to avoid obstacles and maintain driving stability, ensuring that the transportation task is completed safely and efficiently.
[0081] Furthermore, it is easy to understand that the system can also perform real-time monitoring, dynamic evidence storage, and post-event review of the processing dynamics involved in the above collaborative handling process, in order to meet the diverse application needs of various solutions.
[0082] Specifically, for some situations involving the application and processing of solutions for staff, the method of this application may also involve the display of real-time monitoring results, dynamic evidence storage results, and time-tracing results, and may also involve human-computer interaction for further control, such as the adjustment of specific display content / objects.
[0083] In conclusion, regarding the above solutions, this application addresses the issue of low automation levels in existing automated guided vehicle (AGV) handling strategies, leading to high task execution time costs and low system response efficiency. It also effectively solves the problems of insufficient environmental perception and visual positioning accuracy in existing AGV handling strategies, resulting in difficulties in positioning and low target acquisition success rates in complex or dynamic scenarios. Furthermore, it effectively addresses the lack of efficient human-computer interaction mechanisms, limited dynamic interaction capabilities, and inability to flexibly adapt to changing working conditions in existing AGV handling strategies. Finally, it effectively addresses the poor understanding of user natural language commands in existing AGV handling strategies. Therefore, in practical applications, this solution can meet the needs for intelligent, efficient, and precise handling.
[0084] Specifically, the above-mentioned scheme settings include the following more detailed explanations of its effects: 1) As the above steps S101 to S103 build a map in the operating environment, accept user-defined paths and combine real-time positioning technology, they not only realize high-precision perception and mapping of the working environment, but also support flexible setting of transport paths according to task requirements. At the same time, through dynamic obstacle avoidance control, they improve the autonomous navigation capability of the automated guided vehicle in complex environments, thereby effectively solving the technical problem of poor autonomous navigation capability in transport tasks. 2) Since step S106 extracts the feature information of objects in the image through the cross-attention mechanism algorithm, it can effectively improve the visual positioning accuracy of the multimodal large model, thereby more accurately locating the pixel position of the object and effectively solving the technical problem that existing handling robots have difficulty in accurately locating environmental objects. 3) Since steps S107 and S108 enable the multimodal large model to understand the visual information of the target object in the image through context instruction learning, the robotic arm can obtain the relative position of the target object in the physical world with relatively high accuracy, thereby effectively solving the technical problem that existing handling robots are difficult to achieve high-precision object grasping in the physical world. 4) Since step S109 parses the mechanical instructions of the robotic arm for grasping and placing, enabling the robotic arm to complete the specified object grasping task in an orderly manner, it effectively solves the technical problem that existing handling robots are unable to achieve automated handling tasks after reaching the destination.
[0085] The above is an introduction to the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV provided in this application. To facilitate better implementation of the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV provided in this application, this application also provides a collaborative handling device based on the integration of a multimodal large-scale robotic arm and an AGV from the perspective of functional modules.
[0086] Specifically, in this application, the collaborative handling device based on the integration of a multimodal large-scale robotic arm and an AGV may include the following structure: The mapping unit is used to build an environmental map of the preset operating environment in real time during the driving of the automated guided vehicle in the preset operating environment; The selection unit is used to select a route that passes through the starting point and the destination point based on the environment map. The navigation unit is used to enable the automated guided vehicle to navigate autonomously from the starting point to the target point based on the structured analysis of the running route; The initialization unit is used to initialize the parameters of the robotic arm driven by a multimodal large model in response to the environment of the target point. The imaging unit is used to capture images of the environment of the object to be moved at the target point using a camera mounted on the robotic arm. The enhancement unit is used to combine the cross-attention mechanism algorithm and the name of the object to be moved to perform data enhancement on the environmental image, resulting in an enhanced environmental image. The localization unit is used to input the enhanced environment image and predefined system prompts into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The conversion unit is used to convert pixel position coordinates into relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; The gripping and placing unit is used to call the motion execution module of the robotic arm to the designated position to grip the object to be transported based on the relative position coordinates, and then place the object to be transported. The return unit is used to allow the automated guided vehicle to autonomously navigate back to the starting point according to the operating route.
[0087] As an exemplary embodiment, the mapping unit is specifically used for: Control the automated guided vehicle to follow the initial path within the preset operating environment. During the lap, lidar data was continuously collected. With camera image data To form a set of environmental characteristic information , ; Set of environmental feature information Time synchronization and spatial registration are performed to generate a continuous sequence of sensor observations. , ,in, For sensor data synchronization and registration functions; Sensor observation sequence Input SLAM mapping module Output the current pose estimate Environmental Map , ,in, For automated guided vehicles in Pose information at any given moment SLAM mapping module Output a local or global environment map; Environmental Map Post-processing is performed to obtain the runtime environment map. The post-processing includes map boundary clipping, noise filtering, and coordinate standardization for use in subsequent path navigation.
[0088] As yet another exemplary embodiment, the enhancement unit is specifically used for: Environmental images and the name text of the object to be moved selected by the user. Both are input into the multimodal model. Extracting environmental images and name text The corresponding cross-attention map representation , Among them, multimodal model Cross-attention map representation is used to align spatial and semantic feature information of the input. The size of the rows and columns is the environmental image. The number of rows and columns after dividing into fixed-size blocks; Representing the cross-attention map As weights, the environmental image is generated through interpolation. Masks for data augmentation , ,in, For environmental images The number of columns, For environmental images The number of rows, This indicates that the cross-attention map is represented using an interpolation method. Convert the matrix size to and ; cover face With environmental images Fusion to generate enhanced environment images , ,in, Set the minimum visibility for the selected image.
[0089] As another exemplary embodiment, the positioning unit is specifically used for: Based on enhanced environmental images Combined with predefined system prompts Construct image-text pairs for input , Among them, predefined system prompt words Prompt words designed for locating pixel coordinates of objects in an image, based on contextual instructions; Input image-text pairs Input up to a multimodal large model To obtain the image of the object to be moved in the enhanced environment. pixel position coordinates , Among them, the multimodal large model with joint reasoning capabilities of image and language It contains a visual encoder, a language encoder, and a fusion module. pixel position coordinates Perform a range validity check; if it does not conform to the image pixel range, then call the multimodal large model again. Regenerate the pixel position coordinates.
[0090] As yet another exemplary embodiment, in the scope reasonableness check, for environmental images Let the pixel size be Then the pixel position coordinates The following inequalities must be satisfied: , .
[0091] As another exemplary embodiment, the conversion unit is specifically used for: Using the robotic arm parameters from the previous initialization parameters, a mapping model between pixels and physical coordinates is established, and an affine transformation model is obtained through an offline hand-eye calibration process. Affine transformation model Used to determine the pixel position in the image coordinate system Mapped to position in the physical coordinate system below the working plane of the robotic arm , ; pixel position coordinates Input to affine transformation model Outputs the relative position coordinates in the physical world coordinate system. , ,in, It is a corresponding affine transformation model The conversion function.
[0092] As another exemplary embodiment, the gripping and releasing unit is specifically used for: Based on relative position coordinates Combined with the depth information of the object to be moved Generate a sequence of motion commands to control the movement of the robotic arm's end effector. , ,in, This indicates that the robotic arm has moved to the set gripping position. This indicates that the gripper is closed to complete the object grasping process; According to the sequence of action instructions Call the robotic arm's motion execution module Execute the sequence of action instructions sequentially. To complete the process of grasping the target object, and has ,in, This indicates whether the grabbing operation was successful, confirmed by the terminal feedback signal; Based on the initial robotic arm position parameters from the previously initialized parameters. Construct a sequence of placement action instructions , ,in, This indicates that the robotic arm has returned to its starting position. This indicates that the grippers are open to complete the object placement; According to the sequence of placement action instructions Call the robotic arm's motion execution module Place the object to be moved at the target location and release it, and have ,in, This indicates whether the placement operation was successful.
[0093] This application also provides a processing device from a hardware architecture perspective. Specifically, the processing device of this application may include a processor, a memory, and input / output devices (if a device cluster is involved, the devices in the device cluster are also configured as a single processing device). The processor is used to execute computer programs stored in the memory to implement, for example... Figure 1 The corresponding embodiments describe the steps of the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV; or, when the processor executes the computer program stored in the memory, it implements the functions of each unit as described in the above embodiments, and the memory is used to store the processor's execution of the above... Figure 1The computer program required for the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV in the corresponding embodiment.
[0094] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory and executed by a processor to complete this application. One or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a computer device.
[0095] The processing device may include, but is not limited to, processors, memory, and input / output devices. Those skilled in the art will understand that the illustrations are merely examples of processing devices and do not constitute a limitation on the processing device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the processing device may also include network access devices, buses, etc., with processors, memory, input / output devices, etc., connected via buses.
[0096] A processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the processing device, connecting all parts of the device through various interfaces and lines.
[0097] Memory can be used to store computer programs and / or modules. The processor performs various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can primarily include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the processing device, etc. Furthermore, memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0098] When a processor executes a computer program stored in memory, it can specifically perform the following functions: The automated guided vehicle (AGV) constructs an environmental map of the preset operating environment in real time as it travels through the preset operating environment. Select the route that passes through the starting point and the destination point based on the environmental map; Based on the structured analysis of the route, the automated guided vehicle (AGV) can autonomously navigate from the starting point to the target point. The parameters of the robotic arm, driven by a multimodal large model, are initialized for the environment of the target point. These parameters include the robotic arm parameters and its position. Using a camera mounted on a robotic arm, the environment of the target point where the object needs to be moved is photographed to obtain an environmental image; By combining the cross-attention mechanism algorithm and the name of the object to be moved, data augmentation is performed on the environmental image to obtain an enhanced environmental image; The enhanced environmental image and predefined system prompts are input into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The pixel position coordinates are converted into the relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; Based on the relative position coordinates, the robot arm's motion execution module is invoked to move to the designated position to grab the object to be transported, and then the object to be transported is placed. The automated guided vehicle (AGV) will autonomously navigate back to its starting point based on the route it takes.
[0099] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described collaborative handling device, processing equipment, and its corresponding units based on the integration of a multimodal large-scale robotic arm and an AGV can be found in, for example... Figure 1 The description of the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV in the corresponding embodiment will not be repeated here.
[0100] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0101] Therefore, this application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the present application. Figure 1 The steps of the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV in the corresponding embodiment can be referred to as follows for specific operations. Figure 1The description of the collaborative handling method based on the integration of multimodal large-scale robotic arm and AGV in the corresponding embodiment will not be repeated here.
[0102] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0103] Because of the instructions stored in the computer-readable storage medium, the present application can be executed as described above. Figure 1 The steps of the collaborative handling method based on the integration of a multimodal large-scale robotic arm and an AGV in the corresponding embodiment can therefore achieve the results of this application. Figure 1 The beneficial effects that can be achieved by the collaborative handling method based on the integration of multimodal large-scale robotic arm and AGV in the corresponding embodiment are detailed in the preceding description and will not be repeated here.
[0104] The above provides a detailed description of the collaborative handling method, apparatus, processing equipment, and computer-readable storage medium based on the integration of a multimodal large-model robotic arm and an AGV. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the core ideas of this application. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A collaborative handling method based on the integration of a multimodal large-model robotic arm and an AGV, characterized in that, The method includes: During the operation of the automated guided vehicle in the preset operating environment, an environmental map of the preset operating environment is constructed in real time. Select a route that passes through the starting point and the target point based on the environmental map; Based on the structured analysis of the operating route, the automated guided vehicle (AGV) can autonomously navigate from the starting point to the target point. Initialize the parameters of the robotic arm driven by a multimodal large model for the environment of the target point; Using a camera mounted on the robotic arm, the environment of the object to be moved at the target point is photographed to obtain an environmental image; By combining the cross-attention mechanism algorithm and the name of the object to be moved, the environmental image is augmented to obtain an enhanced environmental image; The enhanced environment image and predefined system prompts are input into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The pixel position coordinates are converted into the relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; Based on the relative position coordinates, the motion execution module of the robotic arm is invoked to move to the designated position to grab the object to be transported, and then the object to be transported is placed. The automated guided vehicle (AGV) will autonomously navigate back to the starting point according to the route.
2. The method according to claim 1, characterized in that, The process of constructing an environmental map of the preset operating environment in real time during the operation of the automated guided vehicle includes: The automated guided vehicle is controlled to follow the initial path within the preset operating environment. During the lap, lidar data was continuously collected. With camera image data To form a set of environmental characteristic information , ; The environmental feature information set Time synchronization and spatial registration are performed to generate a continuous sequence of sensor observations. , ,in, For sensor data synchronization and registration functions; The sensor observation sequence Input SLAM mapping module Output the current pose estimate Environmental Map , ,in, For the automated guided vehicle in Pose information at any given moment For the SLAM mapping module Output a local or global environment map; For the environmental map Post-processing is performed to obtain the runtime environment map. The post-processing includes map boundary clipping, noise filtering, and coordinate standardization for use in subsequent path navigation.
3. The method according to claim 1, characterized in that, The method of combining the cross-attention mechanism algorithm and the name of the object to be moved to perform data augmentation on the image to obtain an enhanced environment image includes: Environmental images and the name text of the object to be moved selected by the user. Both are input into the multimodal model. Extract the environmental image. and the name text The corresponding cross-attention map representation , The multimodal model The cross-attention map is used to align the spatial and semantic features of the input. The size of the rows and columns is that of the environmental image. The number of rows and columns after dividing into fixed-size blocks; Represent the cross attention map As weights, the environmental image is generated through interpolation. Masks for data augmentation , ,in, For the environmental image The number of columns, For the environmental image The number of rows, This indicates that the cross-attention map is represented using an interpolation method. Convert the matrix size to and ; The mask With the environmental image Fusion to generate enhanced environment images , ,in, Set the minimum visibility for the selected image.
4. The method according to claim 1, characterized in that, The step of inputting the enhanced environment image and predefined system prompts into the multimodal large model to obtain the pixel position coordinates of the object to be transported includes: Based on enhanced environmental images Combined with predefined system prompts Construct image-text pairs for input , The predefined system prompt words Prompt words designed for locating pixel coordinates of objects in an image, based on contextual instructions; Input the image-text pair Input up to a multimodal large model The desired transported object is obtained in the enhanced environment image. pixel position coordinates , Among them, the multimodal large model with joint reasoning capabilities of image and language It contains a visual encoder, a language encoder, and a fusion module. pixel position coordinates A range validity check is performed. If the range does not conform to the image pixel range, the multimodal large model is invoked again. Regenerate the pixel position coordinates.
5. The method according to claim 4, characterized in that, In the range reasonableness check, for the environmental image Let the pixel size be Then the pixel position coordinates The following inequalities must be satisfied: , 。 6. The method according to claim 1, characterized in that, The step of converting the pixel position coordinates into the relative position coordinates of the robotic arm in the physical world using a hand-eye calibration algorithm includes: Using the robotic arm parameters from the previously initialized parameters, a mapping model between pixels and physical coordinates is established, and an affine transformation model is obtained through an offline hand-eye calibration process. The affine transformation model Used to determine the pixel position in the image coordinate system Mapped to position in the physical coordinate system below the working plane of the robotic arm , ; pixel position coordinates Input to the affine transformation model Output the relative position coordinates corresponding to the physical world coordinate system. , ,in, It corresponds to the affine transformation model. The conversion function.
7. The method according to claim 1, characterized in that, The step of calling the robotic arm's motion execution module to a designated position to grasp the object to be transported based on the relative position coordinates, and then placing the object to be transported, includes: Based on relative position coordinates Combined with the depth information of the object to be transported Generate a sequence of motion commands to control the movement of the robotic arm's end effector. , ,in, This indicates that the robotic arm has moved to the set grasping position. This indicates that the gripper is closed to complete the object grasping process; According to the action instruction sequence Call the motion execution module of the robotic arm The sequence of action instructions is executed sequentially. To complete the process of grasping the target object, and has ,in, This indicates whether the grabbing operation was successful, confirmed by the terminal feedback signal; Based on the initial robotic arm position parameters from the previously initialized parameters. Construct a sequence of placement action instructions , ,in, This indicates that the robotic arm has returned to its starting position. This indicates that the grippers are opened to complete the object placement; According to the placement action instruction sequence Call the motion execution module of the robotic arm Place the object to be transported at the target location and release it, and have ,in, This indicates whether the placement operation was successful.
8. A collaborative handling device based on the integration of a multimodal large-model robotic arm and an AGV, characterized in that, The device includes: The mapping unit is used to construct an environmental map of the preset operating environment in real time during the driving of the automated guided vehicle in the preset operating environment; The selection unit is used to select a route that passes through the starting point and the target point based on the environmental map. The navigation unit is used to enable the automated guided vehicle to autonomously navigate from the starting point to the target point based on the structured analysis of the running route; An initialization unit is used to initialize the parameters of a robotic arm driven by a multimodal large model for the environment of the target point, wherein the parameters of the robotic arm include robotic arm parameters and robotic arm position. The imaging unit is used to capture images of the environment of the object to be moved at the target point using a camera mounted on the robotic arm, thereby obtaining an environmental image. An enhancement unit is used to combine a cross-attention mechanism algorithm and the name of the object to be moved to perform data enhancement on the environmental image, thereby obtaining an enhanced environmental image. The positioning unit is used to input the enhanced environment image and predefined system prompts into the multimodal large model to obtain the pixel position coordinates of the object to be transported; The conversion unit is used to convert the pixel position coordinates into the relative position coordinates of the robotic arm in the physical world coordinate system using a hand-eye calibration algorithm; The gripping and placing unit is used to call the motion execution module of the robotic arm to a designated position to grip the object to be transported according to the relative position coordinates, and then place the object to be transported. The return unit is used to enable the automated guided vehicle to autonomously navigate back to the starting point according to the operating route.
9. A processing device, characterized in that, The method includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the method as described in any one of claims 1 to 7 when it invokes the computer program in the memory.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Clouded AGV application system of 5G smart factory
CN112731914A
Using method of omnidirectional AGV (Automatic Guided Vehicle) for cooperatively grabbing small devices based on double mechanical arms
CN117773936A
Intelligent mechanical arm control system scheme based on multi-model cooperation
CN119897866A
Path planning method and system based on four-AGV cooperative carrying
CN120778132A