Robot control method, apparatus, and electronic device
By combining large language models and multimodal large models, high precision and low cost of robot grasping of goods are achieved, solving the problems of low grasping accuracy and high cost in existing technologies, and realizing automated updating of the item database.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING PHOENIX TECHNOLOGY CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-06-23
Smart Images

Figure CN122253207A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics, and in particular to a robot control method, device, and electronic device. Background Technology
[0002] With the rapid iteration of smart retail technologies, intelligent goods positioning and task execution have become core components for improving the intelligence level of scenarios. Related technologies mainly revolve around three core aspects: goods location, downstream execution adaptation, and goods inventory maintenance, resulting in various technical solutions. However, regarding goods location, all related technologies suffer from low goods positioning accuracy. Furthermore, the lack of a unified adaptation interface for downstream execution routes means that traditional modular routes and Vision Language Action (VLA) routes cannot be seamlessly integrated, requiring modifications to existing systems for collaborative operation, which is costly. This leads to lower accuracy and higher costs for robots in grasping goods.
[0003] Therefore, improving the accuracy of robots in grasping goods and reducing the cost of grasping goods has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a robot control method, device, and electronic device, which can not only improve the accuracy of the robot in grasping goods, but also reduce the cost of grasping goods.
[0005] In a first aspect, embodiments of this application provide a robot control method, the method comprising:
[0006] Based on the initial demand instruction input by the target object, obtain the image corresponding to the shelf area where the target goods required to store the target object are stored, and extract at least one candidate goods area from the image;
[0007] The large language model is invoked to determine the target object's cargo demand information based on the initial demand instructions and the item library to which the target cargo belongs. Based on the cargo demand information, the initial demand instructions are normalized to obtain the target demand instructions.
[0008] Based on cargo demand information, a target candidate region is determined from at least one cargo candidate region;
[0009] The target requirement command and the target candidate region are input into a pre-trained multimodal large model to obtain control commands for controlling the robot;
[0010] Send control commands to the robot so that it can grasp the target goods based on the control commands.
[0011] In this embodiment, on the one hand, at least one candidate product area is determined from the image corresponding to the shelf area where the target goods required for the target object are stored. This allows for rapid determination of candidate product areas through coarse positioning. Then, based on the target object's goods demand information, a target candidate area is determined from the at least one candidate product area. This allows for accurate determination of the target candidate area corresponding to the goods that meet the user's needs through fine positioning. Therefore, dual positioning technology can quickly and accurately determine the target candidate area corresponding to the goods that meet the user's needs, thereby improving the accuracy of goods positioning and thus improving the accuracy of the robot's goods grasping. On the other hand, control commands for controlling the robot are generated based on the target candidate areas and target demand instructions. This allows for seamless conversion of goods positioning results into control commands, eliminating the need to modify the existing system. The command interface can seamlessly connect to traditional modular routes and VLA routes, thereby reducing the cost of the robot grasping goods. Therefore, this embodiment not only improves the accuracy of the robot's goods grasping but also reduces the cost of goods grasping.
[0012] In one optional implementation of the first aspect, determining a target candidate region from multiple candidate regions based on cargo demand information includes: determining the target cargo required by the target object based on cargo demand information; obtaining cargo information of the target cargo from an item database; and determining the target candidate region from multiple candidate regions based on the cargo information.
[0013] Using this implementation method, the target candidate region to which the target goods belong can be determined simply and quickly based on the goods demand information.
[0014] In an optional implementation of the first aspect, determining a target candidate region from multiple candidate regions based on cargo information includes: mapping cargo information into a vector space to obtain a first vector corresponding to the cargo information; mapping each candidate region in the multiple candidate regions into a vector space to obtain a second vector corresponding to each candidate region; determining the similarity between the first vector and each second vector; and selecting at least one candidate region corresponding to at least one second vector with a similarity greater than a preset similarity threshold as the target candidate region.
[0015] By using this implementation method, multiple candidate regions of goods are filtered by vector similarity based on the goods information required by the target object, thus quickly determining the target candidate region to which the target goods belong.
[0016] In an optional implementation of the first aspect, extracting at least one cargo candidate region from an image includes: invoking a pre-trained visual detection model to determine the object category and predicted bounding box corresponding to each object included in the image; and using the predicted bounding boxes corresponding to at least one object whose object category is cargo as at least one cargo candidate region.
[0017] Using this implementation method, multiple candidate areas for goods can be quickly identified by utilizing a visual inspection model.
[0018] In an optional implementation of the first aspect, the control command includes identification information and location information of the target goods required by the target object, and control parameters for controlling the robot; sending the control command to the robot so that the robot can grasp the target goods based on the control command includes: sending the control command to the robot so that the robot can locate the target goods from the shelf area based on the identification information and location information of the target goods, and grasp the target goods based on the control parameters.
[0019] In this implementation method, since the control command includes the identification information of the target goods, after the control command is transmitted to the traditional modular robot, the robot can obtain the goods information of the target goods from the item warehouse based on the identification information, and accurately locate the target goods from the shelf area based on the goods information and location information, and complete the grasping of the target goods.
[0020] In an optional implementation of the first aspect, the control command includes the location information and appearance attributes of the target goods required by the target object, as well as control parameters for controlling the robot; sending the control command to the robot so that the robot can grasp the target goods based on the control command includes: sending the control command to the robot so that the robot can locate the target goods from the shelf area based on the location information and appearance attributes of the target goods, and grasp the target goods based on the control parameters.
[0021] In this implementation method, since the control command includes the appearance attributes of the target goods, and the appearance attributes are visual attributes, the VLA execution route can be seamlessly connected. Thus, when the robot grasps the target goods, it can match the visual features in the collected image with the appearance attributes in the control command to locate the target goods from the shelf area and complete the precise grasping of the target goods.
[0022] In an optional implementation of the first aspect, the method further includes: after the robot grasps the target goods, obtaining the identification information of the target goods; removing the identification information from the item library and deducting the inventory quantity corresponding to the target goods to obtain an updated item library; the updated item library is used to determine the goods demand information of the target object in the next step.
[0023] By adopting this implementation method, after the robot grabs the target goods, the item database can be automatically updated based on the identification information of the target goods. This can minimize the need for manual inventory checks and manual entry of goods information into the item database, thereby reducing labor costs.
[0024] In an optional implementation of the first aspect, the method further includes: when there is new goods to be added to the item library, acquiring first image data and second image data of the new goods; the first image data is acquired when the new goods are not placed in the shelf area, and the second image data is acquired when the new goods are already placed in the shelf area; determining the goods characteristics of the new goods based on the first image data, and determining the location information of the new goods based on the second image data; determining new identification information corresponding to the new goods based on the goods characteristics and location information; performing structured processing on the new identification information, and adding the processed new identification information to the item library to obtain an updated item library; the updated item library is used to determine the goods demand information of the target object in the next step.
[0025] By adopting this implementation method, when there are new goods to be added to the item database, the new goods can be automatically and accurately entered into the database, and the item database can be automatically updated. This can minimize the need for manual entry of goods information into the item database, thereby reducing labor costs.
[0026] Secondly, embodiments of this application provide a robot control device, the device comprising:
[0027] The acquisition and processing module is used to acquire an image of the shelf area where the target goods are stored, based on the initial demand command input by the target object, and extract at least one candidate goods area from the image.
[0028] The determination module is used to call the large language model, determine the goods demand information of the target object based on the initial demand instruction and the item library to which the target goods belong, and perform normalization processing on the initial demand instruction based on the goods demand information to obtain the target demand instruction.
[0029] The determination module is also used to determine a target candidate region from at least one candidate region of goods based on the goods demand information;
[0030] The processing module is used to input the target requirement command and the target candidate region into the pre-trained multimodal large model to obtain the control command for controlling the robot;
[0031] The sending module is used to send control commands to the robot so that the robot can grasp the target goods based on the control commands.
[0032] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method provided in the first aspect above.
[0033] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the first aspect above.
[0034] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method provided in the first aspect above.
[0035] Regarding the beneficial effects of any of the technical solutions in the second to fifth aspects mentioned above, refer to the beneficial effects of the corresponding technical solutions in the first aspect; repeated examples will not be listed here. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is an optional flowchart illustrating a robot control method provided in an embodiment of this application;
[0038] Figure 2 This is a schematic diagram of an optional process for automatic updating of an item database provided in an embodiment of this application;
[0039] Figure 3 This is another optional flowchart illustrating a robot control method provided in an embodiment of this application;
[0040] Figure 4 This is a schematic diagram of an optional structure of a robot control device provided in an embodiment of this application;
[0041] Figure 5 This is a schematic diagram of another optional structure of a robot control device provided in the embodiments of this application;
[0042] Figure 6 This is a schematic diagram of an optional structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0044] The robot control method provided in the embodiments of this application will be described below.
[0045] Please see Figure 1 , Figure 1 This is an optional flowchart illustrating a robot control method provided in an embodiment of this application. The method can be executed by an electronic device. The electronic device can be a terminal device or a server. Terminal devices mentioned herein can include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc., and portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. The server mentioned herein can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, etc., without limitation. Figure 1 As shown, the robot control method may include, but is not limited to, the following steps:
[0046] S101. Based on the initial demand instruction input by the target object, obtain the image corresponding to the shelf area where the target goods required to store the target object are stored, and extract at least one candidate goods area from the image.
[0047] In one optional implementation, the electronic device acquires an image of the shelf area corresponding to the target goods required by the target object based on the initial demand command input by the target object. This can be achieved by: determining the initial goods requirement of the target object based on the initial demand command input by the target object, wherein the initial goods requirement includes initial information of the target goods required by the target object; and acquiring an image of the shelf area corresponding to the target goods required by the target object based on the initial information.
[0048] In some embodiments, the electronic device acquires an image of the shelf area corresponding to the goods required for storing the target object based on initial information. This can be achieved by: the electronic device acquiring the location information of the shelf area where the target goods are stored from the item warehouse to which the target goods belong; the electronic device sending an image acquisition command to the robot, and the robot receiving the image acquisition command from the electronic device, the image acquisition command including the location information of the shelf area; the robot moving to the location of the shelf area based on the location information of the shelf area included in the image acquisition command, and acquiring an image corresponding to the shelf area; the robot sending the image corresponding to the shelf area where the target goods are stored to the electronic device, and the electronic device acquiring the image corresponding to the shelf area where the target goods are stored.
[0049] The robots are pre-deployed in a space that stores various types of goods. For example, this space could be a supermarket, and the various types of goods may include, but are not limited to, snacks, beverages, daily necessities, and fresh fruits and vegetables.
[0050] For example, suppose the initial demand instruction input by the target object is "Give me a bag of potato chips". In this case, the electronic device can determine the initial information of the target object's required goods as potato chips based on the initial demand instruction. At this time, the electronic device can obtain the location information of the shelf area where the potato chips are stored from the item library, such as the third shelf of shelf A. Then, the electronic device can send an image acquisition instruction to the robot, which includes the above location information, namely the third shelf of shelf A. After that, the robot can move to the side of shelf A based on the location information included in the image acquisition instruction and acquire the image corresponding to the third shelf area of shelf A. Finally, the robot can send the acquired image to the electronic device.
[0051] In other embodiments, the electronic device acquires an image of the shelf area corresponding to the goods required for storing the target object based on initial information. Alternatively, the electronic device may send an image acquisition command to a robot, and the robot receives the image acquisition command from the electronic device, the image acquisition command including initial information about the target goods; in response to the image acquisition command, the robot acquires the location information of the shelf area where the target goods are stored from the item warehouse to which the target goods belong, based on the initial information; the robot moves to the location of the shelf area based on the location information of the shelf area and acquires an image corresponding to the shelf area; the robot sends the image corresponding to the shelf area where the target goods are stored to the electronic device, and the electronic device acquires images corresponding to shelf areas storing multiple types of goods.
[0052] The robots are pre-deployed in a space that stores various types of goods. For example, this space could be a supermarket, and the various types of goods may include, but are not limited to, snacks, beverages, daily necessities, and fresh fruits and vegetables.
[0053] For example, suppose the initial demand instruction input by the target object is "Give me a bag of potato chips". In this case, the electronic device can determine the initial information of the target goods required by the target object based on the initial demand instruction, which is potato chips. At this time, the electronic device can send an image acquisition instruction to the robot. The image acquisition instruction includes the initial information of the target goods, namely potato chips. Then, the robot can obtain the location information of the shelf area where potato chips are stored from the item library based on the image acquisition instruction, such as the third shelf of shelf A. After that, the robot can move to the side of shelf A based on the above location information and acquire the image corresponding to the third shelf area of shelf A. Finally, the robot can send the acquired image to the electronic device.
[0054] In one alternative implementation, the electronic device extracts at least one cargo candidate region from the image by: calling a pre-trained visual detection model to determine at least one cargo candidate region based on the image.
[0055] Optionally, the visual detection model can be a lightweight network model, for example, a lightweight U-Net network.
[0056] S102. Call the large language model, determine the target object's cargo demand information based on the initial demand instruction and the item library to which the target cargo belongs, and standardize the initial demand instruction based on the cargo demand information to obtain the target demand instruction.
[0057] Among them, the Large Language Model (LLM) is a deep learning model in the field of artificial intelligence, specifically designed for understanding and generating human language. Optionally, the Large Language Model can be an open-source model obtained by an electronic device from a comprehensive model platform.
[0058] In some embodiments, the initial demand instruction may be obtained by the electronic device by: displaying an information input box in the user interface; monitoring the information entered into the information input box; and, if the information is detected, treating the information as an initial demand instruction for the target object.
[0059] For example, an initial demand instruction could be something like, “Give me a bag of potato chips.”
[0060] In some embodiments, the electronic device invokes a large language model to determine the target object's goods demand information based on the initial demand instructions input by the target object and a multi-category goods database. This can be achieved by: integrating the multi-category goods database as a knowledge base into the large language model to obtain a target large language model; inputting the initial demand instructions input by the target object into the target large language model, and determining the target object's goods demand information through multi-turn dialogue.
[0061] For example, suppose user 1's initial demand instruction is "Give me a bag of potato chips". The electronic device can use the target large language model, based on the corresponding item database, to output the information "There are potato chips of brand A and brand B. Which brand do you want?" In this case, user 1 can input the instruction "I want potato chips of brand A". Then, the electronic device can further use the target large language model to output the information "Brand A potato chips have spicy, cucumber, lime and tomato flavors. Which flavor do you want?" User 1 can input the instruction "I want spicy". In this case, the electronic device can determine that user 1's goods demand information is: spicy potato chips of brand A.
[0062] In some embodiments, the electronic device normalizes the initial demand instruction based on the goods demand information to obtain the target demand instruction. This can be achieved by: determining the information of the target goods from the item library based on the goods demand information, including location information, image information, etc.; and normalizing the initial demand instruction based on the target goods information to obtain the target demand instruction.
[0063] For example, continuing from the above example, assuming user 1's goods demand information is spicy potato chips of brand A, the electronic device can determine the information of spicy potato chips of brand A from the item database based on the goods demand information, such as location information (on the second shelf of shelf A), etc. Then, the electronic device can normalize the initial demand instruction "Give me a bag of potato chips" based on the information of spicy potato chips of brand A, and get "Grab spicy potato chips of brand A on the second shelf of shelf A and give them to user 1".
[0064] S103. Based on cargo demand information, determine the target candidate region from at least one cargo candidate region.
[0065] In one alternative implementation, the electronic device determines a target candidate region from multiple candidate regions based on cargo demand information. This can be achieved by: determining cargo information of the target cargo required by the target object based on the cargo demand information; and determining the target candidate region from multiple candidate regions based on the cargo information.
[0066] S104. Input the target requirement command and the target candidate region into the pre-trained multimodal large model to obtain the control command for controlling the robot.
[0067] The control commands may include, but are not limited to, the location information of the target cargo required by the target object and the control parameters used to control the robot.
[0068] In some embodiments, the multimodal large model may be an open-source model obtained by the electronic device from a comprehensive modeling platform.
[0069] S105. Send control commands to the robot so that the robot can grasp the target goods based on the control commands.
[0070] In this process, after the robot grabs the target goods based on control commands, it can also deliver the target goods to the target object.
[0071] In one optional implementation, after step S105, the electronic device can also update the item database corresponding to multiple types of goods after the robot grabs the target goods, to obtain an updated item database; wherein, the updated item database is used to determine the goods demand information of the target object in the next step.
[0072] In this embodiment, on the one hand, at least one candidate product area is determined from the image corresponding to the shelf area where the target goods required for the target object are stored. This allows for rapid determination of candidate product areas through coarse positioning. Then, based on the target object's goods demand information, a target candidate area is determined from the at least one candidate product area. This allows for accurate determination of the target candidate area corresponding to the goods that meet the user's needs through fine positioning. Therefore, dual positioning technology can quickly and accurately determine the target candidate area corresponding to the goods that meet the user's needs, thereby improving the accuracy of goods positioning and thus improving the accuracy of the robot's goods grasping. On the other hand, control commands for controlling the robot are generated based on the target candidate areas and target demand instructions. This allows for seamless conversion of goods positioning results into control commands, eliminating the need to modify the existing system. The command interface can seamlessly connect to traditional modular routes and VLA routes, thereby reducing the cost of the robot grasping goods. Therefore, this embodiment not only improves the accuracy of the robot's goods grasping but also reduces the cost of goods grasping.
[0073] In one alternative implementation, Figure 1 Step S103 of the robot control method shown, namely, the way in which the electronic device determines the target candidate region from at least one cargo candidate region based on cargo demand information, may be: determining the target cargo required by the target object based on cargo demand information; obtaining cargo information of the target cargo from the item library; and determining the target candidate region from at least one cargo candidate region based on the cargo information.
[0074] For example, assuming the goods demand information is spicy potato chips of brand A, the electronic device can determine that the target goods required by the target object are spicy potato chips of brand A.
[0075] The information about the target goods may include, but is not limited to, images, location information, warehousing time, and calorie content. For example, assuming the target goods are spicy potato chips of brand A, the electronic device can retrieve images of spicy potato chips of brand A from the item database, such as image 1 of the outer packaging; location information, such as the second shelf of shelf A; warehousing time, such as 11:00:00 on April 20, 2026; and calorie content, such as 400 kcal.
[0076] In some embodiments, the electronic device obtains the goods information of the target goods from the item database, which may be: retrieving the goods information of the target goods from the item database based on the name of the target goods.
[0077] In some embodiments, the electronic device determines the target candidate region from at least one cargo candidate region based on cargo information by: mapping the cargo information into a vector space to obtain a first vector corresponding to the cargo information; mapping each cargo candidate region in the at least one cargo candidate region into a vector space to obtain a second vector corresponding to each cargo candidate region; determining the similarity between the first vector and each second vector; and selecting at least one cargo candidate region corresponding to at least one second vector with a similarity greater than a preset similarity threshold as the target candidate region. In this way, by performing vector similarity filtering on at least one cargo candidate region based on the cargo information of the target cargo required by the target object, the target candidate region to which the target cargo belongs can be quickly determined.
[0078] Optionally, the preset similarity threshold can be determined based on expert experience, multiple trials, or manually defined, etc., without limitation here.
[0079] Optionally, the electronic device determines the similarity between the first vector and each of the second vectors, which may be by determining the cosine similarity between the first vector and each of the second vectors, or by determining the Euclidean distance between the first vector and each of the second vectors, etc., without limitation here.
[0080] Using this implementation method, the target candidate region to which the target goods belong can be determined simply and quickly based on the goods demand information.
[0081] In one alternative implementation, Figure 1 In step S101 of the robot control method shown, the electronic device may extract at least one cargo candidate region from the image by: calling a pre-trained visual detection model to determine the object category and predicted bounding box corresponding to each object included in the image; and taking the predicted bounding boxes corresponding to at least one object whose object category is cargo as at least one cargo candidate region.
[0082] The visual detection model can be a lightweight neural network model. For example, the visual detection model can be a lightweight U-Net network.
[0083] In some embodiments, the visual detection model can be trained by an electronic device in the following manner: acquiring training images and corresponding annotation information, the annotation information including the actual category and actual bounding box of each object in the training image; inputting the training images into a pre-built initial visual detection model to obtain the predicted category and predicted bounding box of each object in the training image; inputting the predicted category, predicted bounding box, actual category, and actual bounding box of each object into a preset loss function to obtain a loss value; wherein the preset loss function may include a classification loss term and a regression loss term, the classification loss term being used to determine the difference between the predicted category and the actual category of each object in the training image, and the regression loss term being used to determine the difference between the predicted bounding box and the actual bounding box of each object in the training image; updating the model parameters of the initial visual detection model with the goal of minimizing the loss value to obtain the visual detection model.
[0084] In some embodiments, the electronic device invokes a pre-trained visual detection model to determine the object category and predicted bounding box corresponding to each object in the image. This can be achieved by: denoising the image to obtain a processed image; and inputting the processed image into the pre-trained visual detection model to obtain the predicted category and predicted bounding box corresponding to each object in the image. By denoising the image separately, image quality can be improved, thereby improving the visual detection results.
[0085] For example, suppose that for image 1, the electronic device can perform denoising processing on image 1 to obtain a processed image 1; the processed image 1 is then input into a pre-trained visual detection model to obtain the predicted category and predicted bounding box corresponding to each object in the processed image 1, wherein the predicted category corresponding to each object is, for example, any of the following: aluminum can, beverage bottle, box, snack bag, etc. In this case, the electronic device can determine that aluminum can, beverage bottle, and snack bag are goods. At this time, the electronic device can use the predicted bounding boxes corresponding to aluminum can, beverage bottle, and box, respectively, as three candidate goods regions corresponding to image 1.
[0086] Using this implementation method, multiple candidate areas for goods can be quickly identified by utilizing a visual inspection model.
[0087] In one alternative implementation, Figure 1In the robot control method shown, the control command includes the identification information and location information of the target goods required by the target object, and the control parameters used to control the robot; step S105, that is, the electronic device sends the control command to the robot so that the robot can grasp the target goods based on the control command, can be: sending the control command to the robot so that the robot can locate the target goods from the shelf area based on the identification information and location information of the target goods, and grasp the target goods based on the control parameters.
[0088] In some embodiments, the control parameters for controlling the robot may include a first control parameter for controlling the robot's manipulator and a second control parameter for controlling the robot to move to the location corresponding to the location information. In this embodiment, after receiving a control command from an electronic device, the robot, in response to the control command, retrieves the target goods information from the item warehouse based on the identification information of the target goods; locates the target goods from the shelf area based on the goods information and location information; moves to the location corresponding to the location information based on the second control parameter; and at the location corresponding to the location information, grasps the target goods based on the first control parameter.
[0089] Optionally, the cargo information of the target cargo may include, but is not limited to, images of the target cargo, the time of entry into the warehouse, etc.
[0090] Optionally, the control parameters for controlling the robot may also include a third control parameter for controlling the robot to move to the location of the target object; after the robot grabs the target goods based on the first control parameter, it may also move to the location of the target object based on the third control parameter and pass the target goods to the target object.
[0091] In this implementation method, since the control command includes the identification information of the target goods, after the control command is transmitted to the traditional modular robot, the robot can obtain the goods information corresponding to the target goods from the item warehouse based on the identification information, and accurately locate the target goods from the shelf area based on the goods information and location information, and complete the grasping of the target goods.
[0092] In one alternative implementation, Figure 1 In the robot control method shown, the control command includes the location information and appearance attributes of the target goods required by the target object, as well as the control parameters used to control the robot; the electronic device sends the control command to the robot so that the robot can grasp the target goods based on the control command. This can be done by sending the control command to the robot so that the robot can locate the target goods from the shelf area based on the location information and appearance attributes of the target goods, and grasp the target goods based on the control parameters.
[0093] In some embodiments, the control parameters for controlling the robot may include a first control parameter for controlling the robot's manipulator and a second control parameter for controlling the robot to move to the location corresponding to the location information. In this embodiment, after receiving a control command from an electronic device, the robot, in response to the control command, moves to the location corresponding to the location information based on the second control parameter and acquires an image of the goods on the shelf corresponding to that location; extracts multiple goods features from the goods image; determines a target goods feature from the multiple goods features based on appearance attributes; locates the target goods from the shelf area based on the target goods feature; and grasps the target goods based on the first control parameter.
[0094] Optionally, the robot determines the target cargo feature from multiple cargo features based on appearance attributes. This can be achieved by: mapping the appearance attributes to a feature space to obtain the appearance features corresponding to the appearance attributes; determining the similarity between the appearance features and multiple cargo features respectively, and determining at least one candidate cargo feature from multiple cargo features based on the multiple similarities; performing geometric space verification on at least one candidate cargo feature based on the geometric structure of the appearance features, and using the verified candidate cargo feature as the target cargo feature.
[0095] Optionally, the robot determines at least one candidate cargo feature from multiple cargo features based on multiple similarities. This can be achieved by: determining at least one similarity from multiple similarities that has a similarity greater than a preset similarity threshold; and using the cargo feature corresponding to at least one similarity as at least one candidate cargo feature.
[0096] The preset similarity threshold can be determined based on expert experience, multiple trials, or is defined manually, etc., and no limitation is made here.
[0097] In this implementation method, since the control command includes the appearance attributes of the target goods, and the appearance attributes are visual attributes, the VLA execution route can be seamlessly connected. Thus, when the robot grasps the target goods, it can match the visual features in the collected image with the appearance attributes in the control command to locate the target goods from the shelf area and complete the precise grasping of the target goods.
[0098] In one alternative implementation, Figure 1 In the robot control method shown, the electronic device can also obtain the identification information of the target goods after the robot grabs the target goods; remove the identification information from the item library and deduct the inventory quantity corresponding to the target goods to obtain an updated item library; the updated item library is used to determine the goods demand information of the target object in the next time.
[0099] In some embodiments, after the robot grabs the target goods, the electronic device obtains the identification information of the target goods, which may be done by: obtaining a transaction log or task operation record after the robot grabs the target goods; and obtaining the identification information of the target goods from the transaction log or task operation record.
[0100] In some embodiments, after obtaining the updated item library, the electronic device can also verify the updated item library; if the verification result shows that the inventory quantity is an integer and the identification information of the target goods is correct, the updated item library is fed back to the large language model so that the large language model can determine the goods demand information of the target object based on the updated item library in the next task.
[0101] By adopting this implementation method, after the robot grabs the target goods, the item database can be automatically updated based on the identification information of the target goods. This can minimize the need for manual inventory checks and manual entry of goods information into the item database, thereby reducing labor costs.
[0102] In one alternative implementation, Figure 1 In the robot control method shown, the electronic device can also acquire first image data and second image data of the new goods to be added to the item library when there are new goods to be added. The first image data is acquired when the new goods are not placed in the shelf area, and the second image data is acquired when the new goods are already placed in the shelf area. Based on the first image data, the characteristics of the new goods are determined, and based on the second image data, the location information of the new goods is determined. Based on the characteristics and location information, the corresponding new identification information of the new goods is determined. The new identification information is structured and added to the item library to obtain an updated item library. The updated item library is used to determine the goods demand information of the target object in the next step.
[0103] In some embodiments, both the first image data and the second image data can be collected by the robot through its own image acquisition device; wherein, the robot may collect image data corresponding to the new goods before placing the new goods in the shelf area for storing multiple types of goods, and use the image data as the first image data; the robot may also collect image data corresponding to the new goods and the shelf position after placing the new goods in the shelf area for storing multiple types of goods, and use the image data as the second image data.
[0104] The first image data may include the appearance attributes and specifications of the newly added goods, or the characteristics of the goods may include the appearance attributes and specifications of the newly added goods; the second image data may include the appearance attributes and specifications of the newly added goods, as well as the shelf location information corresponding to the newly added goods.
[0105] In some embodiments, the electronic device determines the cargo characteristics of the newly added cargo based on the first image data, and determines the location information of the newly added cargo based on the second image data. This can be achieved by: inputting the first image data into a multimodal large model to obtain the cargo characteristics of the newly added cargo, and inputting the second image data into the multimodal large model to obtain the location information of the newly added cargo.
[0106] In some embodiments, the electronic device determines the new identification information corresponding to the newly added goods based on the characteristics and location information of the goods. This can be achieved by inputting the characteristics and location information of the goods into a multimodal large model to obtain the new identification information corresponding to the newly added goods.
[0107] In other embodiments, the electronic device determines the new identification information corresponding to the new cargo based on cargo characteristics and location information. This can be achieved by using hash calculation to obtain the new identification information corresponding to the new cargo based on cargo characteristics and location information.
[0108] In some embodiments, the electronic device performs structured processing on the newly added identification information, which may include: verifying the newly added identification information to obtain a verification result; and performing structured processing on the newly added identification information if the verification result indicates that the verification has passed.
[0109] Optionally, the electronic device can verify the newly added identification information by comparing it with multiple identification information in the item database. If the comparison result indicates that the newly added identification information does not exist in the item database, the verification result indicates that the verification has passed.
[0110] In some embodiments, after obtaining the updated item library, the electronic device can also feed the updated item library back to the large language model, so that the large language model can determine the target object's goods demand information based on the updated item library in the next task.
[0111] By adopting this implementation method, when there are new goods to be added to the item database, the new goods can be automatically and accurately entered into the database, and the item database can be automatically updated. This can minimize the need for manual entry of goods information into the item database, thereby reducing labor costs.
[0112] Please see Figure 2 , Figure 2 This is a schematic diagram of an optional process for automatically updating an item database according to an embodiment of this application. This process corresponds to the aforementioned electronic device's process of determining the updated item database. For example... Figure 2As shown, after the robot grasps the target goods, the electronic device can first determine whether there are any new goods to be added to the item library. If so, it acquires the first image data of the new goods before they are placed on the shelf area and inputs the first image data into the multimodal large model to obtain the characteristics of the new goods. It then acquires the second image data of the new goods after they are placed on the shelf and inputs the second image data into the multimodal large model to obtain the location information of the new goods. The goods characteristics and location information are input into the multimodal large model to obtain the new identification information corresponding to the new goods. The new identification information is structured and entered into the item library to obtain the updated item library, thus completing the entry of the new goods into the inventory. If not, it acquires the transaction log or task operation record. It then obtains the identification information of the target goods from the transaction log or task operation record. From the item library, it removes the identification information and deducts the corresponding inventory quantity of the target goods to obtain the updated item library. The updated item library is then verified. If the verification result shows that the inventory quantity is an integer and the identification information of the target goods matches correctly, the update of the item library is confirmed to be complete. This method can accurately complete the automatic entry of new goods into the warehouse (item warehouse) and the automatic update of the item warehouse after the robot grabs the target goods. In this way, we can get rid of the dependence on manual inventory and manual data entry as much as possible.
[0113] The following is combined with Figure 3 This paper provides an overall description of the robot control method provided in the embodiments of this application. Please refer to [link / reference]. Figure 3 , Figure 3 This is another optional flowchart illustrating a robot control method provided in an embodiment of this application. Figure 3 As shown, the robot control method may include, but is not limited to, the following steps:
[0114] S301. Based on the initial demand instruction input by the target object, obtain the image corresponding to the shelf area where the target goods required to store the target object are stored, and extract at least one candidate goods area from the image.
[0115] In one alternative implementation, the electronic device extracts at least one cargo candidate region from the image by: calling a pre-trained visual detection model to determine the object category and predicted bounding box corresponding to each object included in the image; and taking the predicted bounding boxes corresponding to at least one object whose object category is cargo as at least one cargo candidate region.
[0116] S302. Call the large language model, determine the target object's cargo demand information based on the initial demand instruction and the item library to which the target cargo belongs, and standardize the initial demand instruction based on the cargo demand information to obtain the target demand instruction.
[0117] In an optional implementation, the relevant description of step S302 can be found in the description of step S102 above, and will not be repeated here.
[0118] In this application, the execution order of steps S301 and S302 is not limited. For example, the electronic device may execute step S301 first and then step S302; or it may execute step S302 first and then step S301; or it may execute steps S301 and S302 simultaneously, etc., without limitation.
[0119] S303. Based on cargo demand information, determine the target candidate region from at least one cargo candidate region.
[0120] In some embodiments, the electronic device determines a target candidate region from at least one candidate region based on cargo demand information. This can be achieved by: determining the target cargo required by the target object based on the cargo demand information; obtaining cargo information of the target cargo from an item library; mapping the cargo information to a vector space to obtain a first vector corresponding to the cargo information; and mapping each candidate region of the at least one candidate region to a vector space to obtain a second vector corresponding to each candidate region; determining the similarity between the first vector and each second vector; and selecting the candidate region corresponding to at least one second vector with a similarity greater than a preset similarity threshold as the target candidate region.
[0121] S304. Input the target requirement command and the target candidate region into the pre-trained multimodal large model to obtain the control command for controlling the robot.
[0122] S305. Send control commands to the robot so that the robot can grasp the target goods based on the control commands.
[0123] In one optional implementation, the control command includes the identification information and location information of the target goods required by the target object, and control parameters for controlling the robot; the electronic device sends the control command to the robot so that the robot can grasp the target goods based on the control command, which can be: sending the control command to the robot so that the robot can locate the target goods from the shelf area based on the identification information and location information of the target goods, and grasp the target goods based on the control parameters.
[0124] In another optional implementation, the control command includes the location information and appearance attributes of the target goods required by the target object, as well as control parameters for controlling the robot; the electronic device sends control commands to the robot so that the robot can grasp the target goods based on the control commands, which can be: sending control commands to the robot so that the robot can locate the target goods from the shelf area based on the location information and appearance attributes of the target goods, and grasp the target goods based on the control parameters.
[0125] S306. After the robot grabs the target goods, determine whether there are any new goods to be added to the item library. If yes, proceed to steps S307 to S310; otherwise, proceed to steps S311 to S312.
[0126] S307. Acquire first image data and second image data of the newly added goods; the first image data is acquired when the newly added goods are not placed in the shelf area, and the second image data is acquired when the newly added goods are placed in the shelf area.
[0127] S308. Based on the first image data, determine the characteristics of the newly added goods, and based on the second image data, determine the location information of the newly added goods.
[0128] S309. Based on cargo characteristics and location information, determine the new identification information corresponding to the newly added cargo.
[0129] S310. The newly added identification information is structured and added to the item library to obtain the updated item library; the updated item library is used to determine the goods demand information of the target object in the next step.
[0130] S311. Obtain transaction logs or task operation records, and obtain the identification information of the target goods from the transaction logs or task operation records.
[0131] S312. Remove the identification information from the item database and deduct the inventory quantity corresponding to the target goods to obtain the updated item database; the updated item database is used to determine the goods demand information of the target object in the next step.
[0132] In this embodiment, on the one hand, at least one candidate product area is determined from the image corresponding to the shelf area where the target goods required by the target object are stored. This allows for rapid determination of candidate product areas through coarse positioning. Then, based on the target object's goods demand information, a target candidate area is determined from the at least one candidate product area. This allows for accurate determination of the target candidate area corresponding to the goods that meet the user's needs through fine positioning. Therefore, dual positioning technology can quickly and accurately determine the target candidate area corresponding to the goods that meet the user's needs, thereby improving the accuracy of goods positioning and thus improving the accuracy of the robot's goods grasping. On the other hand, control commands for controlling the robot are generated based on the target candidate areas and target demand instructions. This allows for seamless conversion of goods positioning results into control commands, eliminating the need to modify the existing system. The command interface can seamlessly connect to traditional modular routes and VLA routes, enabling efficient coordination between target goods positioning and target goods grasping. This not only reduces the difficulty of implementation but also lowers the cost of robot goods grasping. Therefore, this embodiment not only improves the accuracy of robot goods grasping but also reduces the cost of goods grasping.
[0133] Furthermore, in this embodiment, after the robot grabs the target goods, the electronic device can accurately complete two major scenarios: automatic entry of newly added goods into the warehouse (item warehouse) and automatic updating of the item warehouse after the robot grabs the target goods. In this way, on the one hand, through the automated maintenance technology of the item warehouse, the reliance on manual inventory and manual data entry can be reduced as much as possible, thereby reducing labor costs; on the other hand, in the next execution of robot control tasks, based on the updated item warehouse, the user's initial demand command can be parsed to obtain more accurate information on the user's goods demand.
[0134] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0135] Based on the same inventive concept, this application also provides a robot control device for implementing the robot control method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more robot control device embodiments provided below can be found in the limitations of the robot control method described above, and will not be repeated here.
[0136] Please see Figure 4 , Figure 4 This is a schematic diagram of an optional structure of a robot control device provided in an embodiment of this application. For example... Figure 4 As shown, the robot control device may include, but is not limited to:
[0137] The acquisition and processing module 401 is used to acquire an image of the shelf area corresponding to the target goods required to store the target object based on the initial demand instruction input by the target object, and extract at least one candidate goods area from the image.
[0138] The determination module 402 is used to call the large language model, determine the goods demand information of the target object based on the initial demand instruction and the item library to which the target goods belong, and perform normalization processing on the initial demand instruction based on the goods demand information to obtain the target demand instruction.
[0139] The determination module 402 is also used to determine a target candidate region from at least one cargo candidate region based on cargo demand information;
[0140] Processing module 403 is used to input the target demand command and the target candidate region into a pre-trained multimodal large model to obtain control commands for controlling the robot;
[0141] The sending module 404 is used to send control commands to the robot so that the robot can grasp the target goods based on the control commands.
[0142] In some embodiments, when determining a target candidate region from multiple candidate regions based on cargo demand information, the determining module 402 is specifically used to: determine the target cargo required by the target object based on cargo demand information; obtain cargo information of the target cargo from the item library; and determine the target candidate region from multiple candidate regions based on the cargo information.
[0143] In some embodiments, when determining a target candidate region from multiple candidate regions based on cargo information, the determining module 402 is specifically configured to: map cargo information into a vector space to obtain a first vector corresponding to the cargo information; and map each candidate region in the multiple candidate regions into a vector space to obtain a second vector corresponding to each candidate region; determine the similarity between the first vector and each second vector; and take at least one candidate region corresponding to at least one second vector with a similarity greater than a preset similarity threshold as the target candidate region.
[0144] In some embodiments, when the acquisition and processing module 401 extracts at least one cargo candidate region from an image, it is specifically used to: call a pre-trained visual detection model to determine the object category and predicted bounding box corresponding to each object included in the image; and take the predicted bounding boxes corresponding to at least one object whose object category is cargo as at least one cargo candidate region.
[0145] In some embodiments, the control command includes the identification information and location information of the target goods required by the target object, and control parameters for controlling the robot; when the sending module 404 sends the control command to the robot so that the robot can grasp the target goods based on the control command, it is specifically used to: send the control command to the robot so that the robot can locate the target goods from the shelf area based on the identification information and location information of the target goods, and grasp the target goods based on the control parameters.
[0146] In some embodiments, the control command includes the location information and appearance attributes of the target goods required by the target object, as well as control parameters for controlling the robot; when the sending module 404 sends the control command to the robot so that the robot can grasp the target goods based on the control command, it is specifically used to: send the control command to the robot so that the robot can locate the target goods from the shelf area based on the location information and appearance attributes of the target goods, and grasp the target goods based on the control parameters.
[0147] In some embodiments, the acquisition and processing module 401 is further configured to acquire the identification information of the target goods after the robot grabs the target goods; remove the identification information from the item library and deduct the inventory quantity corresponding to the target goods to obtain an updated item library; the updated item library is used to determine the goods demand information of the target object in the next step.
[0148] In some embodiments, the acquisition and processing module 401 is further configured to acquire first image data and second image data of the new goods to be added to the item library when there are new goods to be added to the item library; the first image data is acquired when the new goods are not placed in the shelf area, and the second image data is acquired when the new goods are placed in the shelf area; based on the first image data, determine the characteristics of the new goods, and based on the second image data, determine the location information of the new goods; based on the characteristics and location information, determine the new identification information corresponding to the new goods; perform structured processing on the new identification information, and add the processed new identification information to the item library to obtain an updated item library; the updated item library is used to determine the goods demand information of the target object in the next step.
[0149] It is understood that the specific implementation of each module in the robot control device provided in this application embodiment and the beneficial effects that can be achieved can be referred to the description of the robot control method embodiment above, and will not be repeated here.
[0150] Please see Figure 5 , Figure 5 This is a schematic diagram of another optional structure of a robot control device provided in an embodiment of this application. For example... Figure 5 As shown, the robot control device may include, but is not limited to: a visual detection module 501, an intent clarification and standardization module 502 based on a large language model and an item database, a standardized task instruction generation module 503 based on a multimodal large model, and an item database update module 504; wherein, the above modules are seamlessly connected and work together through standardized data interfaces to form a closed loop of the entire process from the initial requirement instruction input of the target object to task execution and item database update.
[0151] The visual detection module 501 is used to detect the image corresponding to the shelf area where the target goods are stored, and to obtain at least one candidate goods area.
[0152] The intent definition and standardization module 502, based on a large language model and an item database, is used to call the large language model, determine the target object's goods demand information based on the initial demand instruction input by the target object and the item database to which the target goods belong, and standardize the initial demand instruction based on the goods demand information to obtain the target demand instruction.
[0153] The standardized task instruction generation module 503 based on a multimodal large model is used to determine a target candidate region from at least one candidate region of goods based on goods demand information, and input the target demand instruction and the target candidate region into a pre-trained multimodal large model. Based on the downstream adaptation route, control instructions for controlling the robot are obtained; the control instructions are sent to the robot so that the robot can grasp the target goods based on the control instructions; wherein, when the downstream adaptation route is a traditional modular operation route, the control instructions include the identification information of the target goods; when the downstream adaptation route is a VLA route, the control instructions include the appearance attributes of the target goods.
[0154] The item database update module 504 is used to update the item database by adding new goods to the database when it is determined that there are new goods to be added to the database, or to update the item database based on the identification information of the captured target goods when it is determined that there are no new goods to be added to the database.
[0155] Each module in the aforementioned robot control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the onboard terminal device in hardware form or independently of it, or they can be stored in the memory of the robot control device in software form, so that the processor can call and execute the operations corresponding to each module.
[0156] In one exemplary embodiment, an electronic device is provided, the internal structure of which can be shown as follows: Figure 6 As shown, this electronic device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a robot control method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the electronic device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads installed inside the electronic device.
[0157] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0158] In one exemplary embodiment, this application provides an electronic device, including a memory and a motor domain electronic device, wherein the memory stores a computer program; when the motor domain electronic device executes the computer program, it implements the steps in the robot control methods described above.
[0159] In one exemplary embodiment, this application provides a computer-readable storage medium having a computer program stored thereon. When executed by an electromechanical electronic device, the computer program implements the steps in the robot control methods described above.
[0160] In one exemplary embodiment, this application provides a computer program product, including a computer program. When executed by an electromechanical electronic device, the computer program implements the steps in the robot control methods described above.
[0161] It should be noted that the data involved in this application (including but not limited to data used for analysis, data stored, data displayed, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0162] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0163] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0164] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A robot control method, characterized in that, The method includes: Based on the initial demand command input by the target object, obtain an image corresponding to the shelf area where the target goods required by the target object are stored, and extract at least one candidate goods area from the image; The large language model is invoked to determine the goods demand information of the target object based on the initial demand instruction and the item library to which the target goods belong. Based on the goods demand information, the initial demand instruction is normalized to obtain the target demand instruction. Based on the cargo demand information, a target candidate region is determined from at least one of the cargo candidate regions; The target requirement command and the target candidate region are input into a pre-trained multimodal large model to obtain control commands for controlling the robot; The control command is sent to the robot so that the robot can grasp the target goods based on the control command.
2. The method according to claim 1, characterized in that, The step of determining a target candidate region from multiple candidate regions based on the cargo demand information includes: Based on the cargo demand information, the target cargo required by the target object is determined; Obtain the cargo information of the target cargo from the item database; Based on the cargo information, a target candidate region is determined from multiple cargo candidate regions.
3. The method according to claim 2, characterized in that, The step of determining a target candidate region from multiple candidate cargo regions based on the cargo information includes: The cargo information is mapped into a vector space to obtain a first vector corresponding to the cargo information; and each of the multiple cargo candidate regions is mapped into the vector space to obtain a second vector corresponding to each cargo candidate region. Determine the similarity between the first vector and each of the second vectors; At least one cargo candidate region corresponding to at least one of the second vectors whose similarity is greater than a preset similarity threshold is taken as the target candidate region.
4. The method according to claim 1, characterized in that, Extracting at least one cargo candidate region from the image includes: A pre-trained visual detection model is invoked to determine the object category and predicted bounding box of each object included in the image; The predicted bounding boxes corresponding to at least one object whose object category is cargo are used as at least one cargo candidate region.
5. The method according to claim 1, characterized in that, The control instructions include the identification information and location information of the target cargo required by the target object, and the control parameters for controlling the robot; Sending the control command to the robot so that the robot can grasp the target goods based on the control command includes: The control command is sent to the robot so that the robot can locate the target goods from the shelf area based on the identification information and the location information of the target goods, and grab the target goods based on the control parameters.
6. The method according to claim 1, characterized in that, The control instructions include the location information and appearance attributes of the target cargo required by the target object, as well as the control parameters for controlling the robot. Sending the control command to the robot so that the robot can grasp the target goods based on the control command includes: The control command is sent to the robot so that the robot can locate the target goods from the shelf area based on the location information and appearance attributes of the target goods, and grasp the target goods based on the control parameters.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: After the robot grabs the target goods, it acquires the identification information of the target goods; The identification information is removed from the item database, and the inventory quantity corresponding to the target goods is deducted to obtain an updated item database; the updated item database is used to determine the goods demand information of the target object in the next step.
8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: If there are new goods to be added to the item library, acquire first image data and second image data of the new goods; the first image data is acquired when the new goods are not placed in the shelf area, and the second image data is acquired when the new goods are placed in the shelf area. Based on the first image data, the characteristics of the newly added goods are determined, and based on the second image data, the location information of the newly added goods is determined; Based on the cargo characteristics and the location information, determine the new identification information corresponding to the newly added cargo; The newly added identification information is structured and then added to the item database to obtain an updated item database. The updated item database is used to determine the goods demand information of the target object in the next step.
9. A robot control device, characterized in that, The device includes: The acquisition and processing module is used to acquire an image of the shelf area where the target goods required by the target object are stored, based on the initial demand command input by the target object, and to extract at least one candidate goods area from the image. The determination module is used to call the large language model, determine the goods demand information of the target object based on the initial demand instruction and the item library to which the target goods belong, and perform normalization processing on the initial demand instruction based on the goods demand information to obtain the target demand instruction; The determining module is further configured to determine a target candidate region from at least one of the cargo candidate regions based on the cargo demand information; The processing module is used to input the target demand command and the target candidate region into a pre-trained multimodal large model to obtain control commands for controlling the robot; The sending module is used to send the control command to the robot so that the robot can grasp the target goods based on the control command.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program; when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.