Article distribution method and device, readable storage medium and robot
By acquiring images after the robot reaches the delivery task destination and using the object recognition model to identify the object categories in the storage area, the problem of robot invalid stay is solved, and more efficient delivery task execution is achieved.
Patent Information
- Application Number
- CN202410125016.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-07-29
AI Technical Summary
After arriving at the delivery destination, the robot stays due to identifying an object, resulting in an increase in invalid residence time, affecting the efficiency of delivering objects.
By obtaining the image of the target storage area after arriving at the delivery task destination, the trained object recognition model is used to identify the object categories in the storage area. If all are preset categories, the collection is determined to be completed, and the next delivery task is performed without manual confirmation from the user.
It reduces the robot's invalid residence time and improves the efficiency of delivering goods. Especially when preset objects such as small tickets are not taken away in restaurant scenes, the robot can directly perform the next task.
Smart Images

Figure CN120382476A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of robots, and particularly relates to an article delivery method, device, readable storage medium, and robot. Background Art
[0002] With the rapid development of robot technology, intelligent robots have been widely used. In order to save costs and provide more efficient services, robots are beginning to be used to replace humans in delivering articles.
[0003] Currently, during the process of a robot delivering an article, after the robot reaches the destination, when the robot recognizes an object, it stays at the destination and requires manual handling by the user to trigger the execution of the next task. However, this situation increases the idle time of the robot at the destination and affects the efficiency of the robot in delivering articles. Summary of the Invention
[0004] Embodiments of this application provide an article delivery method, device, readable storage medium, and robot, which can solve the problem of the robot staying when detecting an object, resulting in idle stay.
[0005] In a first aspect, embodiments of this application provide an article delivery method, which is applied to a robot provided with at least one layer of article placement area;
[0006] The method includes:
[0007] After reaching the destination of the delivery task, obtain an image of the target article placement area, where the target article placement area is the article placement area for placing the articles of the delivery task;
[0008] If it is determined according to the image that there are objects in the target article placement area and each of the objects is an object of a preset category, then determine that the article collection is completed;
[0009] If it is determined that the robot has a next delivery task, then execute the next delivery task.
[0010] In an embodiment, if it is determined according to the image that there are objects in the target article placement area and each of the objects is an object of a preset category, then determining that the article collection is completed includes:
[0011] Input the image into a trained object recognition model to obtain a detection result output by the trained object recognition model;
[0012] If the detection result indicates that there are objects in the target article placement area and each of the objects is an object of the preset category, then determine that the article collection is completed.
[0013] In one embodiment, the trained object recognition model includes a backbone network, a neck network, and a detection network, and the neck network includes a first detection head, a second detection head, and a third detection head;
[0014] Inputting the image into the trained object recognition model to obtain the detection result output by the trained object recognition model includes:
[0015] Inputting the image into the backbone network; extracting features from the image through the backbone network to obtain feature maps of different sizes; obtaining a first feature map from the backbone network through the first detection head, obtaining a second feature map from the backbone network through the second detection head, and obtaining a third feature map from the backbone network through the third detection head, and then performing upsampling and merging on the first feature map, the second feature map, and the third feature map through the neck network to obtain a fused feature map; detecting the fused feature map through the detection network to obtain and output the detection result, so as to obtain the detection result output by the trained object recognition model;
[0016] Wherein, the size of the first feature map is larger than the size of the second feature map, and the size of the second feature map is larger than the size of the third feature map.
[0017] In one embodiment, before inputting the image into the trained object recognition model, it further includes:
[0018] Obtaining training data, where the training data includes real images and confused images, the real images include objects of the preset category, and the confused images include objects that the object recognition model misjudges as belonging to the preset category;
[0019] Using the training data to train the object recognition model until the loss value of the preset loss function is less than the preset loss value to obtain the trained object recognition model.
[0020] In one embodiment, before inputting the image into the trained object recognition model, it further includes:
[0021] Converting the image into a grayscale image;
[0022] If the grayscale value of the grayscale image is less than the preset grayscale value, performing image enhancement processing on the grayscale image to obtain an enhanced image;
[0023] Correspondingly, inputting the image into the trained object recognition model includes:
[0024] Inputting the enhanced image into the trained object recognition model.
[0025] In one embodiment, when there is an object in the target storage area, the detection result includes the detection frame of the object, and / or, category information;
[0026] If the detection result indicates that there is an object in the target storage area, and each of the objects is an object of the preset category, it is determined that the object fetching is completed, including:
[0027] If the detection result indicates that there is an object in the target storage area, and it is determined that each of the objects is an object of the preset category according to the category information of each of the objects, it is determined that the object fetching is completed;
[0028] Or, if the detection result indicates that there is an object in the target storage area, and it is determined that each of the objects is an object of the preset category according to the texture features of the imaging regions of each of the objects, it is determined that the object fetching is completed, where the imaging region of the object is obtained by extracting from the image through the detection frame of the object.
[0029] In one embodiment, the method further includes:
[0030] If it is determined according to the image that there is no object in the target storage area, and it is determined that the robot has the next delivery task, then execute the next delivery task;
[0031] If a command indicating that the object fetching is completed is received, and it is determined that the robot has the next delivery task, then execute the next delivery task.
[0032] In a second aspect, an embodiment of the present application provides an article delivery device, including:
[0033] An image acquisition module, configured to acquire an image of a target storage area after arriving at the destination of the delivery task, where the target storage area is a storage area for placing the articles of the delivery task;
[0034] An object recognition module, configured to determine that the object fetching is completed if it is determined according to the image that there is an object in the target storage area, and each of the objects is an object of a preset category;
[0035] A task processing module, configured to execute the next delivery task if it is determined that the robot has the next delivery task.
[0036] In a third aspect, an embodiment of the present application provides a robot, including
[0037] A storage area for placing objects;
[0038] An image collector for collecting images;
[0039] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described in any one of the above first aspects is implemented.
[0040] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method described in any one of the above first aspects is implemented.
[0041] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a robot, the robot is enabled to execute the method described in any one of the above first aspects.
[0042] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:
[0043] After arriving at the destination of the delivery task, the embodiment of the present application acquires an image of the target storage area, where the target storage area is the storage area for placing the items of the delivery task; if it is determined according to the image that there are objects in the target storage area and all the objects are objects of a preset category, it is determined that the item collection is completed; if it is determined that the robot has a next delivery task, the next delivery task is executed, so that the robot can also determine whether to execute the next delivery task by confirming the category of the objects in the storage area, and when all the items that need to be taken away by the user in the storage area have been taken away and only the objects that do not affect the completion of the delivery task remain, the robot directly executes the next task, such as the receipt in a restaurant scenario, without the user having to manually click to confirm completion or take away all the items to execute the next task, reducing the ineffective stay time of the robot, shortening the delivery time, and improving the delivery efficiency.
[0044] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the above first aspect and will not be elaborated here. Description of the Drawings
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0046] Figure 1 It is the first flow diagram of the item delivery method provided by an embodiment of the present application;
[0047] Figure 2 It is an example diagram of the receipt provided by an embodiment of the present application;
[0048] Figure 3 It is the second flow schematic diagram of the article delivery method provided by an embodiment of the present application;
[0049] Figure 4 It is the third flow schematic diagram of the article delivery method provided by an embodiment of the present application;
[0050] Figure 5 It is the fourth flow schematic diagram of the article delivery method provided by an embodiment of the present application;
[0051] Figure 6 It is the fifth flow schematic diagram of the article delivery method provided by an embodiment of the present application;
[0052] Figure 7 It is the structural schematic diagram of the article delivery device provided by an embodiment of the present application;
[0053] Figure 8 It is the structural schematic diagram of the robot provided by an embodiment of the present application. Detailed implementation manners
[0054] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are put forward to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0055] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0056] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.
[0057] As used in the specification and the appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0058] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0059] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear at different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0060] In one embodiment, Figure 1 is the first flow schematic diagram of the article delivery method provided by an embodiment of the present application. As Figure 1 shown, the method is applied to a robot provided with at least one layer of storage areas.
[0061] In application, this method can be specifically executed by an article delivery device, which can be implemented by software and / or hardware, and can be configured in the above-mentioned robot or in a background server.
[0062] Specifically, it includes the following steps:
[0063] S11: After arriving at the destination of the delivery task, obtain an image of the target storage area.
[0064] Wherein, the target storage area is a storage area for placing the articles of the delivery task.
[0065] Exemplarily, when the robot is a food delivery robot, the storage area can be a tray, the target storage area can be a tray for placing the articles of the delivery task, and the articles of the delivery task are dishes. When the robot is a goods delivery robot, the storage area can be a compartment, the target storage area can be a compartment for placing the articles of the delivery task, and the articles of the delivery task are goods. Exemplarily, the image can be obtained by an RGB camera set above the storage area. Exemplarily, after the robot arrives at the destination of the delivery task, it can guide the user to pick up the meal through the flashing of the lights corresponding to the storage area.
[0066] In application, after triggering the delivery task, the robot transports the articles to the destination of the delivery task. Then, it can take a picture of the target storage area through a camera set above the storage area to obtain an image of the target storage area.
[0067] Exemplarily, in a restaurant scenario, dishes can be placed in the tray area of the robot. The image includes the imaging area of the tray and the dishes, that is, the shooting range of the camera can cover the entire tray area and a certain range around the tray, such as the area within 5 cm of the edge.
[0068] It can be understood that after arriving at the destination of the delivery task, an image of the target storage area is obtained. That is, during the process of the robot departing for the destination of the delivery task, images do not need to be obtained, which can reduce memory consumption and also reduce the impact of detection during the movement of the robot on motion navigation, and reduce the occurrence probability of other anomalies. At the same time, it is not necessary to perform visual detection and obtain images during the movement of the robot. This enables the robot to consume less resources when arriving at the destination and maintain overall performance.
[0069] S12: If it is determined according to the image that there are objects in the target storage area and all the objects are objects of a preset category, it is determined that the object taking is completed.
[0070] In the application, when the robot delivers items, there may be objects of a preset category attached, but the objects of the preset category do not belong to the items of the delivery task. After the user takes away the items delivered by the robot and only the objects of the preset category remain, based on these objects of the preset category, no user operation is required, and at this time, the robot can directly execute the next task without waiting for the user to operate.
[0071] After obtaining the image of the target storage area, object detection is performed on the image to obtain the storage situation of the target storage area. The storage situation can be the presence situation of the objects in the target storage area. When it is determined according to the image that there are objects in the target storage area, the categories of the objects are determined; when it is determined that all the objects are objects of a preset category, it is determined that the user has completed the object taking.
[0072] When it is determined that not all the objects are objects of a preset category, it is determined that the user has not completed the object taking and needs to continue waiting for the user to take the object until it is determined that the object taking is completed.
[0073] Among them, the objects of the preset category can include receipts. Figure 2 It is an example diagram of a receipt provided by an embodiment of the present application. As Figure 2 shown, the receipt can be a small rectangular white paper, and information such as the order number, item name, and destination of the delivery task can be printed on the paper. Generally, the receipt is placed in the storage area of the robot along with the delivered items to facilitate the service staff to check the dish information and serve the dishes. Since the receipt reflects the order situation and does not affect the completion of the delivery task itself, it can be taken away together with the delivered items at the destination or not taken away together with the delivered items. When not taken away, it will be left in the target storage area.
[0074] In a restaurant scenario, when using a robot for delivery, it is easy for the receipt to be left in the storage area. For example, when the robot delivers the dishes for Table 1 on the first layer, at the food pickup area, the restaurant staff near the food pickup area often place the dishes for Table 1 and the corresponding receipt on the first layer of the robot at the same time, so that when the robot arrives at Table 1, the restaurant staff in the dining area can pick up and serve the food according to the receipt information. That is, for robots without the function of identifying the types of dishes or those that do not display guidance after identification, or when the restaurant wants to improve the serving efficiency of the staff, the receipt is often used frequently. And identifying the types of dishes requires high computing power for the robot, and there are obvious differences among different restaurants, and the accuracy is not high. In order to reduce the hardware requirements for the robot, the robot may not identify the specific types of dishes during use, but only identify the presence or absence of objects. At this time, when the receipt is not taken away and is identified as having an object, since the robot cannot distinguish whether it is a dish or a receipt, the robot will stop. And at this time, the robot should execute the next task. Therefore, in the embodiments of the present application, the receipt is regarded as one of the objects of the preset category, which can identify and distinguish the special situation of only having a receipt, facilitating the robot to quickly and accurately complete the meal pickup confirmation, execute the next task, and improve the efficiency.
[0075] In a possible implementation manner, the objects of the preset category can be objects that do not affect the completion of the delivery task when only these objects are left in the storage area. The objects of the preset category can also include tissues, tableware, etc.
[0076] It can be understood that the objects of the preset category can be determined according to the operation habits of the restaurant. For example, when the restaurant is used to delivering the tableware corresponding to each table by the robot, then at this time the tableware does not belong to the objects of the preset category, and taking away the corresponding tableware is considered to complete the meal pickup and execute the next task, otherwise the robot will wait. For another example, when the restaurant is used to place spare tableware on the robot for guests to take and use, at this time the tableware is not bound to a certain table, then the tableware belongs to the objects of the preset category. When the tableware is not taken away, it does not affect the robot to complete the meal pickup confirmation and continue with the next task.
[0077] Exemplarily, in a restaurant scenario, after the robot arrives at the destination of the delivery task, it acquires an image. When the restaurant is used to placing tissues on the robot for all guests to take, it can be determined according to the image that there are objects in the target storage area, and it is determined that these objects are only receipts and tissues, and it is determined that the object pickup is completed. That is, in this embodiment, first, it is judged whether there are objects in the target storage area according to the image. When there are objects, the objects are further distinguished to judge whether they are objects of the preset category, and each object is an object of the preset category, accurately identifying the situation of "there are objects, but the objects are all objects that do not affect the meal pickup confirmation", and improving the delivery efficiency of the robot.
[0078] For example, if the user actively sends a confirmation instruction to the robot, the robot determines that the meal is taken and enters the subsequent task. For example, the user can complete the meal take confirmation by pressing the confirmation button on the robot screen, and the robot receives the confirmation instruction.
[0079] For example, the robot uses a display screen and voice announcements to guide customers to pick up their food and receipts. However, due to the restaurant's acoustic environment, customers may turn off the voice announcement function or miss the announcement. In other cases, because the tray and the screen are facing away from each other, customers, for efficiency, may choose to pick up their food directly based on the receipt information, rather than looking at the screen. Even if the customer removes the tray, the receipt remains on the storage area.
[0080] S13: If it is determined that the robot has the next delivery task, the next delivery task is executed.
[0081] For example, if it is determined that the robot has a next delivery task, the robot goes to the destination of the next delivery task and performs the next delivery task.
[0082] In addition, if it is determined that the robot does not have the next delivery task, it can perform the return mission and wait for the triggering of the delivery task.
[0083] This embodiment obtains an image of the target storage area after arriving at the destination of the delivery task. The target storage area is the storage area where the items for the delivery task are placed. If it is determined based on the image that there are objects in the target storage area, and all objects are objects of a preset category, it is determined that the object retrieval is completed. If it is determined that the robot has the next delivery task, the next delivery task is executed, so that the robot can also determine whether to execute the next delivery task by confirming the category of the objects in the storage area. When all the items that the user needs to take away in the storage area are taken away and only objects that do not affect the completion of the delivery task are left, the robot directly executes the next task. For example, in a restaurant scenario, there is no need for the user to manually click to confirm completion or take away all items before executing the next task. This reduces the time the robot stays ineffectively, shortens the delivery time, and improves the delivery efficiency.
[0084] In one embodiment, Figure 3 This is a second flow chart of the article delivery method provided by an embodiment of the present application. Figure 3 As shown, the method further includes:
[0085] S14: If it is determined based on the image that there is no object in the target storage area and it is determined that the robot has the next delivery task, the next delivery task is executed.
[0086] In an application, when it is determined that there are no objects in the target storage area based on an image, it indicates that the user has taken away all the objects, and the next step can be carried out. Then, it is determined whether the robot has a next delivery task. When it is determined that the robot has a next delivery task, it goes to the destination of the next delivery task.
[0087] Exemplarily, in a restaurant scenario, if the dishes for table 1 are placed on the first layer and the dishes for table 2 are placed on the second layer, when the robot reaches the delivery stop for table 1, the target storage area to be judged is the first layer. If there are no objects on the first layer and there is a next delivery task, the next delivery task is executed. If the dishes for table 1 are placed on both the first layer and the second layer, the target storage area includes the first layer and the second layer. When there are no objects on both the first layer and the second layer and there is a next delivery task, the next delivery task is executed.
[0088] S15: If a pick-up completion instruction is received and it is determined that the robot has a next delivery task, the next delivery task is executed.
[0089] In an application, when a pick-up completion instruction is received, it indicates that the user has manually confirmed the completion of the delivery task, and the next step can be carried out. Then, it is determined whether the robot has a next delivery task. When it is determined that the robot has a next delivery task, it goes to the destination of the next delivery task.
[0090] In this embodiment, by setting multiple judgment conditions for executing the next delivery task, it can be determined in a timely manner that the pick-up is completed, reducing the ineffective stay time of the robot, shortening the delivery time, and improving the delivery efficiency.
[0091] In one embodiment, Figure 4 is the third flowchart of the article delivery method provided by an embodiment of the present application. As Figure 4 shown, if it is determined based on an image that there are objects in the target storage area and all the objects are objects of a preset category, it is determined that the pick-up is completed, including:
[0092] S121: Input the image into a trained object recognition model to obtain the detection result output by the trained object recognition model.
[0093] In an application, object detection is performed on the image through a deep learning model: the trained object recognition model to obtain the detection result regarding the object.
[0094] Exemplarily, the trained object recognition model can be implemented through the Yolo model.
[0095] Due to the small volume or surface area of objects of a preset category, for example, in a restaurant scenario, receipts, tissues, tableware, etc. have a small volume or surface area, and these objects such as receipts, tissues, tableware, etc. are placed at different positions in the storage area. Generally, object recognition models have a high recognition rate for objects with a large volume and surface area, but a low recognition rate for small objects. To improve the accuracy and recall rate of the model for recognizing small objects, the structure of the object recognition model is improved. In a possible implementation, the detection heads of the model are increased to obtain feature maps related to the characteristics of small objects, so as to improve the accuracy and recall rate of recognizing small objects and reduce the missed detection of small objects located at different positions. The trained object recognition model includes a backbone network, a neck network, and a detection network. The neck network includes a first detection head, a second detection head, and a third detection head.
[0096] Input the image into the trained object recognition model to obtain the detection results output by the trained object recognition model, including:
[0097] S21: Input the image into the backbone network; perform feature extraction on the image through the backbone network to obtain feature maps of different sizes.
[0098] In applications, the backbone network performs multiple downsamplings on the image, and each downsampling realizes feature extraction to obtain feature maps of different sizes. That is, the size of the feature map after downsampling will be reduced.
[0099] It can be understood that multiple downsamplings can reduce the computational amount, prevent overfitting, and increase the receptive field, so that the subsequent network can learn global feature information. For example, in a restaurant scenario, the global features of a tray can be learned.
[0100] Exemplarily, the size of the image input into the backbone network is 640*480. After 4 times of downsampling, the size of the obtained feature map is 160*120. Then after 8 times of downsampling, the size of the obtained feature map is 80*60. Then after 16 times of downsampling, the size of the obtained feature map is 40*30.
[0101] S22: After obtaining the first feature map from the backbone network through the first detection head, the second feature map from the backbone network through the second detection head, and the third feature map from the backbone network through the third detection head, perform upsampling and merging on the first feature map, the second feature map, and the third feature map through the neck network to obtain a fused feature map; wherein, the size of the first feature map is larger than the size of the second feature map, and the size of the second feature map is larger than the size of the third feature map.
[0102] In an application, in the neck network, the added first detection head obtains a first feature map from the backbone network. The size of this first feature map is the largest, and the corresponding receptive field is also large, enabling the subsequent network to detect small objects. The third detection head obtains a third feature map from the backbone network, and this third feature map enables the subsequent network to detect large objects. The second detection head obtains a second feature map from the backbone network, and this second feature map enables the subsequent network to detect objects whose size is between that of small objects and large objects.
[0103] Exemplarily, the added first detection head obtains a first feature map of 160*120, and the corresponding receptive field is 4*4. The second detection head obtains a second feature map of 80*60. The third detection head obtains a third feature map of 40*30.
[0104] Then, the first feature map, the second feature map, and the third feature map are upsampled, and then merged in the channel dimension to fuse the features of each feature map.
[0105] It can be understood that after upsampling and merging each feature map to achieve feature fusion, the accuracy and generalization ability of the model to recognize objects can be improved.
[0106] S23: The fused feature map is detected by the detection network to obtain and output a detection result, so as to obtain the detection result output by the trained object recognition model.
[0107] In an application, in the detection network, after performing multiple depthwise separable convolution processes on the fused feature map, convolution processing is performed to obtain a convolved fused feature map. The convolved fused feature map is detected to obtain a detection result, and the detection result is output.
[0108] It can be understood that performing depthwise separable convolution processing can reduce the computational amount of the detection network and speed up the inference speed of the model. Performing convolution processing can reduce the dimension and the computational amount of the network.
[0109] S122: If the detection result indicates that there are objects in the target storage area and all objects are objects of a preset category, it is determined that the object taking is completed.
[0110] In an application, when there are objects in the target storage area, the detection result will include information about the objects. At this time, the detection result indicates that there are objects in the target storage area. Then, the category of the objects is determined according to the information about the objects. When only objects of the preset category remain in the target storage area, it is determined that the object taking is completed. It can be understood that since items such as receipts, tableware, and tissues are significantly different from the dishes, the computing power required to identify the category of the objects is much less than the computing power required to identify the category of the dishes.
[0111] In a possible implementation manner, when there are objects in the target storage area, the detection result includes the detection frame of the object and / or category information.
[0112] Among them, the detection box of the object involves the position of the center point of the detection box and the width and height of the detection box.
[0113] If the detection result indicates that there is an object in the target storage area and each object is an object of a preset category, it is determined that the object taking is completed, including:
[0114] If the detection result indicates that there is an object in the target storage area and it is determined that each object is an object of a preset category according to the category information of each object, it is determined that the object taking is completed.
[0115] In the application, since the category result of the object in the detection result can represent the category of the object, the category of each object can be directly determined according to the category information of the object. When each object is an object of a preset category, it is determined that the object taking is completed.
[0116] Or, if the detection result indicates that there is an object in the target storage area and it is determined that each object is an object of a preset category according to the texture feature of the imaging area of each object, it is determined that the object taking is completed, and the imaging area of the object is obtained by extracting from the image through the detection box of the object.
[0117] In the application, there may be objects misclassified by the trained object recognition model. For example, in a restaurant scenario, the receipt is a small white paper. When there are white and small dishes, tissues, etc. that are easily misclassified. At this time, the imaging area of the object can be extracted through the center point position and width and height of the detection box, and the texture feature extraction algorithm is used to extract the texture feature of the imaging area of the object. The classifier determines the category of the object according to the texture feature of the object. When each object is an object of a preset category, it is determined that the object taking is completed.
[0118] It can be understood that by extracting the imaging area of the object through the detection box, it is ensured that the extracted texture feature comes from the object itself, reducing the influence of other texture features on the determination of the category of the object. The category of the object can be accurately determined through the texture feature of the object.
[0119] In this embodiment, the image is input into the trained object recognition model to obtain the detection result output by the trained object recognition model; if the detection result indicates that there is an object in the target storage area and each object is an object of a preset category, it is determined that the object taking is completed, and the object can be accurately recognized.
[0120] In one embodiment, Figure 5 is the fourth flow schematic diagram of the article distribution method provided by an embodiment of the present application. As Figure 5 shown, before the image is input into the trained object recognition model, it further includes:
[0121] S31: Obtain training data, where the training data includes real images and confused images.
[0122] Among them, the real image includes objects of a preset category, and the confused image includes objects misjudged by the object recognition model as belonging to the preset category.
[0123] In practice, some objects similar to the objects of the preset category, such objects include objects similar in color and / or shape. For example, tissues and white tableware similar to receipts in a restaurant scene will cause the object recognition model to misjudge such objects as objects of the preset category. Increase the number of confused images in the training data so that the object recognition model can learn the features of such objects, so that such objects can be distinguished from the objects of the preset category. When the objects of the preset category only include receipts, the model can classify the detection results into three categories: objects to be retrieved, receipts, and empty trays. At this time, when training the model, tissues and white tableware are trained as objects to be retrieved. The method of data augmentation can be used to increase the number of picture samples of receipts, tissues, and white tableware, so that the model can have more opportunities to learn the features of these categories during training and distinguish the pictures of tissues and white tableware.
[0124] S32: Use the training data to train the object recognition model until the loss value of the preset loss function is less than the preset loss value, and obtain the trained object recognition model.
[0125] In the application, the preset loss function can be set as the cross-entropy loss function. During the process of training the object recognition model, the stochastic gradient descent method is used to adjust the weight update of the model. When the model converges, the trained object recognition model is obtained.
[0126] In this embodiment, the training data includes real images and confused images. Using the training data, the object recognition model is trained until the loss value of the preset loss function is less than the preset loss value, and the trained object recognition model is obtained, so that the object recognition model learns the features of the objects in the confused images and reduces the probability of misjudgment.
[0127] In one embodiment, Figure 6 is the fifth process schematic diagram of the article distribution method provided by an embodiment of the present application. As Figure 6 shown, before inputting the image into the trained object recognition model, it further includes:
[0128] S41: Convert the image into a grayscale image.
[0129] S42: If the grayscale value of the grayscale image is less than the preset grayscale value, perform image enhancement processing on the grayscale image to obtain the enhanced image.
[0130] In an application, after converting an image into a grayscale image, the grayscale values of the pixels are calculated. The distribution of the grayscale values of the pixels is statistically analyzed to determine the brightness and darkness of the grayscale image. If the grayscale values are concentrated within a preset interval, for example, the grayscale values of more than 80% of the pixels are within [0, 80], it indicates that the brightness of the grayscale image is relatively dark, and image enhancement processing is required to increase the contrast of the grayscale image and make the grayscale image brighter. The brightness of the enhanced image is greater than the preset grayscale value, and the shape of the object in the enhanced image is clearer.
[0131] Correspondingly, inputting the image into a trained object recognition model includes:
[0132] S121`: Inputting the enhanced image into the trained object recognition model.
[0133] In this embodiment, by converting the image into a grayscale image, if the grayscale value of the grayscale image is less than the preset grayscale value, image enhancement processing is performed on the grayscale image to obtain an enhanced image. For example, in a restaurant scene, the light changes greatly, and it is possible that the light at the destination is insufficient, resulting in a relatively dark brightness of the obtained image. Image enhancement processing is performed on the grayscale image to make the shape of the object in the enhanced image clearer, and the trained object recognition model can more easily detect the object, reducing the situation of model missing detection.
[0134] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0135] Corresponding to the method described in the above embodiments, for the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0136] In one embodiment, Figure 7 is a schematic structural diagram of an article delivery device provided by an embodiment of the present application. As Figure 7 shown, the device includes:
[0137] An image acquisition module 10, configured to acquire an image of a target storage area after arriving at the destination of the delivery task, where the target storage area is a storage area for placing the items of the delivery task;
[0138] An object recognition module 11, configured to determine that the item picking is completed if it is determined that there are objects in the target storage area according to the image and all the objects are objects of a preset category;
[0139] A task processing module 12, configured to execute the next delivery task if it is determined that the robot has a next delivery task.
[0140] In a possible implementation, the object recognition module is specifically configured to input an image into a trained object recognition model to obtain a detection result output by the trained object recognition model; if the detection result indicates that there are objects in the target storage area and all the objects are objects of a preset category, it is determined that the object taking is completed.
[0141] In a possible implementation, the trained object recognition model includes a backbone network, a neck network, and a detection network, and the neck network includes a first detection head, a second detection head, and a third detection head.
[0142] The object recognition module is specifically configured to input an image into the backbone network; extract features from the image through the backbone network to obtain feature maps of different sizes; after obtaining a first feature map from the backbone network through the first detection head, a second feature map from the backbone network through the second detection head, and a third feature map from the backbone network through the third detection head, perform upsampling and merging on the first feature map, the second feature map, and the third feature map through the neck network to obtain a fused feature map; detect the fused feature map through the detection network to obtain and output a detection result, so as to obtain the detection result output by the trained object recognition model; wherein, the size of the first feature map is larger than the size of the second feature map, and the size of the second feature map is larger than the size of the third feature map.
[0143] In a possible implementation, the device further includes a training module;
[0144] The training module is configured to obtain training data, where the training data includes real images and confused images, the real images include objects of a preset category, and the confused images include objects that the object recognition model misjudges as belonging to the preset category; use the training data to train the object recognition model until the loss value of a preset loss function is less than a preset loss value to obtain a trained object recognition model.
[0145] In a possible implementation, the device further includes an image processing module:
[0146] The image processing module is configured to convert the image into a grayscale image; if the grayscale value of the grayscale image is less than a preset grayscale value, perform image enhancement processing on the grayscale image to obtain an enhanced image.
[0147] Correspondingly, the object recognition module is configured to input the enhanced image into the trained object recognition model.
[0148] In a possible implementation, when there are objects in the target storage area, the detection result includes a detection frame of the object and / or category information.
[0149] The object recognition module is specifically configured to determine that the object fetching is completed if the detection result indicates that there is an object in the target object placement area and it is determined that each object is an object of a preset category according to the category information of each object; or, if the detection result indicates that there is an object in the target object placement area and it is determined that each object is an object of a preset category according to the texture features of the imaging areas of each object, then determine that the object fetching is completed. The imaging area of the object is obtained by extracting from the image through the detection frame of the object.
[0150] Figure 8 FIG. is a schematic structural diagram of a robot provided by an embodiment of the present application. As Figure 8 shown, the robot 2 of this embodiment includes: at least one processor 20 ( Figure 8 only one is shown in the figure), a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20. When the processor 20 executes the computer program 22, the steps in any of the above method embodiments are implemented. The robot further includes. The robot further includes an object placement area and an image collector (not shown in the figure). Among them, the object placement area is used to place objects, and the image collector is used to collect images.
[0151] The robot 2 may include, but is not limited to, the processor 20 and the memory 21. Those skilled in the art can understand that Figure 8 this is only an example of the robot 2 and does not constitute a limitation on the robot 2. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0152] The so-called processor 20 may be a central processing unit (CPU). The processor 20 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0153] In some embodiments, the memory 21 may be an internal storage unit of the robot 2, such as the hard disk or memory of the robot 2. In other embodiments, the memory 21 may also be an external storage device of the robot 2, such as a plug-in hard disk equipped on the robot 2, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 21 may also include both the internal storage unit of the robot 2 and external storage devices. The memory 21 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 21 may also be used to temporarily store data that has been output or will be output.
[0154] It should be noted that for the content such as information interaction and execution process between the above-mentioned device / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0155] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details are not described herein again.
[0156] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0157] The embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can execute the steps in the above-mentioned method embodiments.
[0158] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some cases, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0159] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0160] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0161] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0162] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0163] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An article distribution method, characterized in that, A robot applied with at least one layer of storage areas; The method includes: After arriving at the destination of the delivery task, acquiring an image of a target storage area, where the target storage area is a storage area for placing the items of the delivery task; If it is determined according to the image that there are objects in the target storage area and all the objects are objects of a preset category, then determine that the item picking is completed; If it is determined that the robot has a next delivery task, then execute the next delivery task.
2. The method according to claim 1, wherein If it is determined according to the image that there are objects in the target storage area and all the objects are objects of a preset category, then determining that the item picking is completed includes: Inputting the image into a trained object recognition model to obtain a detection result output by the trained object recognition model; If the detection result indicates that there are objects in the target storage area and all the objects are objects of the preset category, then determine that the item picking is completed.
3. The method according to claim 2, wherein The trained object recognition model includes a backbone network, a neck network, and a detection network, and the neck network includes a first detection head, a second detection head, and a third detection head; Inputting the image into a trained object recognition model to obtain a detection result output by the trained object recognition model includes: Inputting the image into the backbone network; extracting features of the image through the backbone network to obtain feature maps of different sizes; obtaining a first feature map from the backbone network through the first detection head, obtaining a second feature map from the backbone network through the second detection head, and obtaining a third feature map from the backbone network through the third detection head, then performing upsampling and merging on the first feature map, the second feature map, and the third feature map through the neck network to obtain a fused feature map; detecting the fused feature map through the detection network to obtain and output the detection result, so as to obtain the detection result output by the trained object recognition model; Wherein, the size of the first feature map is larger than the size of the second feature map, and the size of the second feature map is larger than the size of the third feature map.
4. The method according to claim 2 or 3, characterized in that, Before inputting the image into the trained object recognition model, it further includes: Acquiring training data, where the training data includes real images and confused images, the real images include objects of the preset category, and the confused images include objects that the object recognition model wrongly determines to belong to the preset category; Using the training data to train the object recognition model until the loss value of a preset loss function is less than a preset loss value to obtain a trained object recognition model.
5. The method according to claim 4, wherein Before inputting the image into the trained object recognition model, it further includes: Converting the image into a grayscale image; If the grayscale value of the grayscale image is less than a preset grayscale value, then performing image enhancement processing on the grayscale image to obtain an enhanced image; Correspondingly, inputting the image into the trained object recognition model includes: Inputting the enhanced image into the trained object recognition model.
6. The method according to claim 5, wherein When there are objects in the target storage area, the detection result includes detection frames of the objects, and / or, category information; If the detection result indicates that there are objects in the target storage area and each of the objects is an object of the preset category, it is determined that the item picking is completed, including: If the detection result indicates that there are objects in the target storage area and it is determined that each of the objects is an object of the preset category according to the category information of each of the objects, it is determined that the item picking is completed; Or, if the detection result indicates that there are objects in the target storage area and it is determined that each of the objects is an object of the preset category according to the texture features of the imaging areas of each of the objects, it is determined that the item picking is completed, and the imaging area of the object is obtained by extracting from the image through the detection frame of the object.
7. The method according to claim 1, characterized in that It further includes: If it is determined according to the image that there are no objects in the target storage area and it is determined that the robot has the next delivery task, the next delivery task is executed; If a pick-up completion instruction is received and it is determined that the robot has the next delivery task, the next delivery task is executed.
8. An article delivery device, characterized in that, It includes: An image acquisition module, configured to acquire an image of a target storage area after arriving at the destination of the delivery task, where the target storage area is a storage area for placing the items of the delivery task; An object recognition module, configured to determine that the item picking is completed if it is determined according to the image that there are objects in the target storage area and each of the objects is an object of a preset category; A task processing module, configured to execute the next delivery task if it is determined that the robot has the next delivery task.
9. A robot, including: A storage area for placing objects; An image collector for collecting images; A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.