An object crawling method, apparatus, computer device, and storage medium

By determining the confidence level of target candidates in the robot and performing optimization decisions, the problem of task execution difficulties for robots in untrained tasks is solved, thereby improving task success rate and efficiency.

CN115713514BActive Publication Date: 2026-04-03BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Robots struggle to complete tasks or instructions they haven't been trained for, and in scenarios with similar object properties, they are prone to having to identify multiple objects one by one, resulting in low task execution success rate and efficiency.

Method used

By identifying target candidate objects based on the state description information in the target object capture command and the confidence level of candidate objects in the current scene, the system executes preset actions for capturing and asking questions, and optimizes decision-making using confidence level and reward information.

Benefits of technology

It improves the accuracy and efficiency of task execution, reduces the number of interactions with users, and quickly captures target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713514B_ABST
    Figure CN115713514B_ABST
Patent Text Reader

Abstract

This disclosure provides an object grasping method, apparatus, computer device, and storage medium. The method includes: in response to receiving a target object grasping instruction, determining a second confidence level for each candidate object as a target object based on state description information of the target object and a first confidence level that each candidate object in the current scene has each preset state; determining target candidate objects and reward information after performing preset actions on the target candidate objects based on the first and second confidence levels; and determining and executing target preset actions based on the reward information. This disclosure eliminates the need for pre-training the robot, accurately identifies target candidate objects that meet user requirements, and reduces the number of interactions with the user by determining the optimal target preset action, allowing for faster target object grasping and improving task execution success rate and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to an object grasping method, apparatus, computer device, and storage medium. Background Technology

[0002] With the development of technology, service robots are gradually appearing in people's daily lives. Some robots typically have the ability to disambiguate during conversations. When a user gives a command to a robot, if the robot has any doubts about the command, it can proactively ask the user questions, clarify the meaning of the command by combining the user's answers, and successfully complete the user's task.

[0003] Typically, robots only execute pre-trained tasks or process pre-trained instructions. When faced with untrained tasks or instructions, the robot struggles to complete the task. Furthermore, in scenarios where objects have similar properties, the robot may end up pointing to multiple objects one by one. Therefore, these factors contribute to the robot's low task execution success rate and efficiency. Summary of the Invention

[0004] This disclosure provides at least one object grabbing method, apparatus, computer device, and storage medium.

[0005] In a first aspect, embodiments of this disclosure provide an object crawling method, including:

[0006] In response to receiving a target object grabbing instruction, based on the target object's state description information contained in the target object grabbing instruction and the first confidence level of each candidate object in the current scene having each preset state, the second confidence level of each candidate object is determined to be the target object;

[0007] Based on the first confidence level and the second confidence level corresponding to each of the candidate objects, target candidate objects are determined, along with reward information after performing each preset action on the target candidate object; the preset actions include capturing the target candidate object and asking a question;

[0008] Based on the reward information corresponding to each preset action, the target preset action to be executed is determined, and the target preset action is executed for the target candidate object.

[0009] In one optional implementation, the first confidence level of the candidate object having each preset state includes: a third confidence level that the candidate object has each preset feature attribute, a fourth confidence level that the candidate object is located in different preset location regions, and a fifth confidence level that the candidate object has each preset positional relationship with each of the other candidate objects.

[0010] In one optional implementation, the third confidence level is obtained through the following steps:

[0011] The two-dimensional scene image corresponding to the current scene is segmented to obtain the image region corresponding to each of the candidate objects;

[0012] Based on the first similarity between the image region corresponding to the candidate object and the description information of each preset feature attribute, the third confidence level of the candidate object with each preset feature attribute is determined.

[0013] In one optional implementation, the fourth confidence level is obtained through the following steps:

[0014] Obtain the two-dimensional scene image corresponding to the current scene;

[0015] Based on the position range of the candidate object in the two-dimensional scene image and the overlap area between the candidate object and each preset position region in the two-dimensional scene image, a fourth confidence level is determined for the candidate object to be located in each preset position region.

[0016] In one optional implementation, the fifth confidence level is obtained through the following steps:

[0017] Based on the three-dimensional position information of each candidate object in the three-dimensional scene image corresponding to the current scene, the shortest distance vector between each candidate object and the other candidate objects is determined.

[0018] Based on the projection length of the shortest distance vector in each preset direction, a fifth confidence level is determined to indicate that there are preset positional relationships between the candidate object and other candidate objects.

[0019] In one optional implementation, the step of determining the second confidence level of each candidate object as the target object based on the state description information of the target object contained in the target object grabbing instruction and the first confidence level of each candidate object having each preset state in the current scene includes:

[0020] Based on the target feature attributes contained in the state description information of the target object, and the third confidence level of each candidate object having the target feature attributes, each candidate object is determined as the sixth confidence level of the target object;

[0021] Based on the second similarity between the state description information of the target object and the standard location area description information corresponding to each preset location area, a seventh confidence level is determined that the state description information of the target object contains the target preset location area; based on the seventh confidence level and the fourth confidence level that each candidate object is located in the target preset location area, an eighth confidence level is determined that each candidate object is the target object;

[0022] Based on the third similarity between the state description information of the target object and the standard candidate object description information corresponding to each candidate object, a ninth confidence level is determined that the state description information of the target object contains the target candidate object; based on the fourth similarity between the state description information of the target object and the standard positional relationship description information corresponding to each preset positional relationship, a tenth confidence level is determined that the state description information of the target object contains the target preset positional relationship; based on the ninth confidence level, the tenth confidence level, and the fifth confidence level that each candidate object has the target preset positional relationship with each of the other candidate objects, an eleventh confidence level is determined that each candidate object is the target object.

[0023] Based on the sixth confidence level, the eighth confidence level, and the eleventh confidence level, the second confidence level of each candidate object is determined as that of the target object.

[0024] In one optional implementation, the question-posing includes posing a question for the target candidate object; wherein the posed question includes feature attribute description information referencing feature attributes; the question is determined through the following steps:

[0025] Based on the twelfth confidence level of each preset feature attribute pointing to the target candidate object, and the third confidence level of the target candidate object having each preset feature attribute, the reference feature attribute is selected from each preset feature attribute;

[0026] The problem is to generate feature attribute description information that includes the reference feature attributes.

[0027] In one optional implementation, the target preset action includes asking the question; asking the question includes asking a question to the target object;

[0028] After executing the target preset action, the process further includes:

[0029] In response to receiving a response to the question, a fifth similarity is determined between the received response and each preset response.

[0030] Based on the fifth similarity, update the first confidence level of each candidate object in the current scene for each preset state;

[0031] Based on the updated first confidence level, the second confidence level of each candidate object is updated to that of the target object.

[0032] In one optional implementation, the target preset action includes asking the question; asking the question includes asking a question to the target object;

[0033] After executing the target preset action, the process further includes:

[0034] In response to receiving supplementary description information for the problem, based on the supplementary description information and the first confidence level, the second confidence level of each candidate object is updated to be the target object, and the step of determining the reward information is returned.

[0035] Secondly, embodiments of this disclosure also provide an object grasping device, comprising:

[0036] The first determining module is used to respond to receiving a target object grabbing instruction, and determine the second confidence level of each candidate object as the target object based on the target object's state description information contained in the target object grabbing instruction and the first confidence level of each candidate object in the current scene having each preset state.

[0037] The second determining module is used to determine target candidate objects and reward information after performing preset actions on the target candidate objects based on the first confidence level and the second confidence level corresponding to each of the candidate objects respectively; the preset actions include capturing the target candidate object and asking questions;

[0038] The third determining module is used to determine the target preset action to be executed based on the reward information corresponding to each preset action, and to execute the target preset action for the target candidate object.

[0039] Thirdly, embodiments of this disclosure also provide a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.

[0040] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect or any possible implementation of the first aspect.

[0041] The object grasping method provided in this disclosure allows the robot to analyze the probability that each candidate object in the current scene is the target object. This enables the robot to accurately determine the target candidate object that meets the user's requirements without prior training, thereby improving the accuracy of task execution. Furthermore, the optimal target preset action can be determined based on the feedback information after executing each preset action. This not only reduces the number of interactions with the user but also allows the robot to grasp the target object specified by the user as quickly as possible, improving the success rate and efficiency of task execution.

[0042] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0043] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0044] Figure 1 A flowchart of an object grabbing method provided by an embodiment of this disclosure is shown;

[0045] Figure 2 This diagram illustrates a scene view of a current scene provided by an embodiment of the present disclosure;

[0046] Figure 3 This diagram illustrates another scene view of the current scenario provided by an embodiment of the present disclosure;

[0047] Figure 4 A flowchart of another object grabbing method provided by an embodiment of this disclosure is shown;

[0048] Figure 5 This diagram illustrates the architecture of an object grasping device provided in an embodiment of the present disclosure.

[0049] Figure 6A schematic diagram of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0051] During the dialogue and disambiguation process between the robot and the user, when the user gives the robot an instruction, the robot can proactively ask the user questions if it has any doubts about the user's instructions. By combining the user's answers, the robot clarifies the meaning of the instructions and successfully completes the user's task.

[0052] However, robots typically only execute pre-trained tasks or process pre-trained instructions. When faced with untrained tasks or instructions, they struggle to complete them. Furthermore, in scenarios where objects have similar properties, the robot's disambiguation ability deteriorates, easily leading to situations where the robot identifies multiple objects one by one. Therefore, these factors result in low task execution success rates and efficiency for the robot.

[0053] Based on this, this disclosure provides an object crawling method, comprising: responding to receiving a target object crawling instruction, determining a second confidence level for each candidate object as the target object based on the state description information of the target object contained in the target object crawling instruction and a first confidence level for each candidate object in the current scene having each preset state; determining a target candidate object and reward information after performing each preset action on the target candidate object based on the first confidence level and the second confidence level corresponding to each candidate object; the preset action includes crawling the target candidate object and asking a question; determining a target preset action to be executed based on the reward information corresponding to each preset action, and executing the target preset action on the target candidate object.

[0054] In the above process, the robot can analyze the probability that each candidate object in the current scene is the target object. This enables it to accurately identify the target candidate object that meets the user's requirements without prior training, thus improving the accuracy of task execution. Furthermore, it can determine the optimal target preset action based on the feedback information after executing each preset action. This not only reduces the number of conversations with the user but also allows the robot to quickly capture the target object specified by the user, thereby improving the success rate and efficiency of task execution.

[0055] The deficiencies of the above solutions and the proposed solutions are the result of the inventor's practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.

[0056] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0057] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0058] To facilitate understanding of this embodiment, a detailed description of the object grasping method disclosed in this disclosure is provided first. The object grasping method provided in this disclosure is generally executed by a computer device with certain computing capabilities. The object grasping method provided in this disclosure can be applied to scenarios where a user instructs a robot to grasp a target object. The robot used to grasp the target object may have a computer device internally configured to execute the object grasping method. The robot may also have a grasping device for grasping the target object. After the internal computer device determines the target candidate object to be grasped from the candidate objects, it can instruct the grasping device to grasp the target candidate object.

[0059] See Figure 1 The diagram shows a flowchart of an object crawling method provided in an embodiment of this disclosure. The method includes steps S101 to S103, wherein:

[0060] S101: In response to receiving a target object grabbing instruction, based on the target object's state description information contained in the target object grabbing instruction and the first confidence level of each candidate object in the current scene having each preset state, determine that each candidate object is the second confidence level of the target object.

[0061] In this embodiment of the disclosure, the target object grasping instruction can be a voice command issued by the user to the robot to grasp a target object based on the current scene. The target object grasping instruction may include state description information of the target object. Specifically, the state description information may include at least one of the following: its own characteristic attributes (e.g., name, color, shape, size, category, etc.), its location area in the current scene, and its positional relationship with other candidate objects in the current scene.

[0062] The robot can capture images of the current scene. The captured images of the current scene can include at least one candidate object. For example, if the current scene includes a table surface, the robot's captured images of the table surface can include items such as pliers, wrenches, screwdrivers, and hammers placed on the table surface.

[0063] In one implementation, the robot can capture the scene of the current scene through an image acquisition device before the user issues the instruction to grab the target object, and determine the first confidence level of each candidate object in the current scene for each preset state.

[0064] Here, the preset state may include a preset position area in the scene view with preset characteristic attributes, located in the current scene, and a preset positional relationship between the candidate object and each other candidate object.

[0065] The preset feature attributes may include at least one of name, color, shape, size, and category. In one implementation, the scene image in the current scene can be divided into multiple location regions according to preset rules. For example, the scene image in the current scene can be divided into, for instance, as shown below. Figure 2 The diagram shows multiple positional regions: front, back, left, right, center, left-front, right-front, left-back, and right-back. A preset positional region may include at least one of these regions. The positional relationship between a candidate object and each other candidate object can be represented by the orientation information of the candidate object relative to each of the other candidate objects, with each other serving as a reference. Generally, for a candidate object placed statically on a plane, the orientation information of the candidate object relative to each of the other reference objects can refer to the orientation information in the two-dimensional coordinate system of the plane. For example... Figure 3 In the illustration of another current scene, the hammer is located to the left of the double-ended wrench, behind the pliers, and to the left rear of the screwdriver. In other cases, such as when the candidate object is suspended in the target space, the orientation information of the candidate object relative to each other reference object can also refer to the orientation information in the three-dimensional spatial coordinate system.

[0066] The first confidence level of a candidate object having each preset state may include: the third confidence level that the candidate object has each preset feature attribute; the fourth confidence level that the candidate object is located in different preset location regions; and the fifth confidence level that the candidate object has each preset location relationship with each other candidate object.

[0067] In one implementation, the third confidence level of each candidate object for each preset feature attribute can be obtained through the following steps:

[0068] Step a1: Segment the two-dimensional scene image corresponding to the current scene to obtain the image regions corresponding to each candidate object.

[0069] Here, an image acquisition device can be used to acquire an image of the current scene, obtaining a two-dimensional scene image corresponding to the current scene. This two-dimensional scene image can be a red-green-blue (RGB) image. In one implementation, a detector (Detic) in a pre-trained vision-language model can be used to detect candidate objects in the two-dimensional scene image, determining the probability of each candidate object belonging to each category. Then, the two-dimensional scene image is segmented to obtain the image region corresponding to each candidate object.

[0070] Step a2: Based on the first similarity between the image region corresponding to the candidate object and the description information of each preset feature attribute, determine the third confidence level of the candidate object having each preset feature attribute.

[0071] Here, the preset feature attribute description information can be text information that describes the image region using preset feature attributes. For example, the preset feature attribute description information can be "a picture of red pliers", "a picture of yellow pliers", ..., "a picture of blue pliers".

[0072] First, the initial similarity between the image region corresponding to the candidate object and the pre-defined feature attribute description information can be determined. In one implementation, the image region corresponding to the candidate object and the pre-defined feature attribute description information can be input into a Contrastive Language-Image Pre-Training (CLIP) model, and the initial similarity between the image region corresponding to the candidate object and the pre-defined feature attribute description information can be output.

[0073] Then, for each preset feature attribute, the first similarity between the image region corresponding to each candidate object and the description information of the preset feature attribute corresponding to that preset feature attribute is normalized to obtain the third confidence level that the candidate object possesses that preset feature attribute. Here, the higher the first similarity between the image region corresponding to the candidate object and the description information of the preset feature attribute corresponding to a certain preset feature attribute, the higher the third confidence level that the candidate object possesses that preset feature attribute.

[0074] Typically, to more clearly describe the third confidence level of each preset feature attribute possessed by a candidate object, in one implementation, a preset feature attribute (e.g., color) can be selected. Then, the first similarity between the image region corresponding to the candidate object and the preset feature attribute description information corresponding to different preset feature attribute values ​​(e.g., red, yellow, blue, etc.) under that preset feature attribute can be determined. Then, the third confidence level of the candidate object possessing each preset feature attribute value under that preset feature attribute can be determined. For example, the candidate object has a third confidence level of red, the candidate object has a third confidence level of yellow, and the candidate object has a third confidence level of blue.

[0075] In one implementation, the fourth confidence level of the candidate object located in different preset location regions can be determined through the following steps:

[0076] Step b1: Obtain the 2D scene image corresponding to the current scene.

[0077] Alternatively, an image acquisition device can be used to capture images of the current scene, obtaining a two-dimensional scene image corresponding to the current scene. This two-dimensional scene image can also be an RGB image.

[0078] Step b2: Based on the position range of the candidate object in the two-dimensional scene image and the overlap area between it and each preset position region in the two-dimensional scene image, determine the fourth confidence level of the candidate object in each preset position region.

[0079] Here, the position range of the candidate object in the 2D scene image can refer to the position range of the candidate object's bounding box in the 2D scene image. A planar coordinate system can be established in the 2D scene image, and then the position range of the candidate object's bounding box in that planar coordinate system can be determined. Based on the position range of the candidate object's bounding box in that planar coordinate system and various preset position regions, the overlap area between the two can be determined.

[0080] Then, for each preset location region, the position range of each candidate object in the 2D scene image and the overlap area between the candidate object and the preset location region are normalized to obtain the fourth confidence level of the candidate object being located in the preset location region. Here, the larger the position range of the candidate object in the 2D scene image and the larger the overlap area between the candidate object and a certain preset location region, the higher the fourth confidence level of the candidate object being located in the preset location region.

[0081] In one implementation, the fifth confidence level regarding the pre-defined positional relationship between a candidate object and each other candidate object can be determined through the following steps:

[0082] Step c1: Based on the 3D position information of each candidate object in the 3D scene image corresponding to the current scene, determine the shortest distance vector between each candidate object and all other candidate objects.

[0083] Here, an image acquisition device can be used to acquire a 2D scene image and a depth scene image corresponding to the current scene. Then, the 2D scene image and the depth scene image can be used to obtain a 3D scene image corresponding to the current scene.

[0084] Here, the 3D position information of the candidate object can refer to the point cloud information formed by the 3D coordinates of each point on the candidate object. The shortest distance vector can refer to the vector between the two points with the shortest distance from one candidate object to another.

[0085] For each candidate object, the shortest distance vector between the candidate object and other candidate objects can be obtained based on the point cloud information of the candidate object and the point cloud information of other candidate objects.

[0086] Step c2: Based on the projection length of the shortest distance vector in each preset direction, determine the fifth confidence level of the preset positional relationship between the candidate object and other candidate objects.

[0087] Here, the preset direction can refer to a direction on a two-dimensional plane, such as front, back, left, right, etc. For each candidate object, the shortest distance vector between the candidate object and other candidate objects is projected onto each preset direction to obtain the projection length of the shortest distance vector in each preset direction.

[0088] For each preset direction, the projection length of the shortest distance vector between each candidate object and all other candidate objects in that preset direction is normalized to obtain the fifth confidence level that the preset positional relationship exists between the candidate object and the other candidate objects. Here, the longer the projection length of the shortest distance vector between the candidate object and the other candidate objects in that preset direction, the higher the fifth confidence level that the preset positional relationship exists between the candidate object and the other candidate objects.

[0089] The first confidence level of each candidate object in the current scene, obtained by the above process, for each preset state, can be used to determine the second confidence level of each candidate object as the target object after receiving the target object grabbing instruction from the user.

[0090] In one implementation, after receiving a target object capture instruction from a user, the state description information of the target object contained in the instruction can be split according to a syntax tree to obtain three parts of state description information: the target feature attributes of the target object, the target location region it is located in, and the target location relationship between the target object and each other candidate object. Then, for the above three parts of state description information, the second confidence level of each candidate object can be determined as the target object, which can be performed according to steps d1-d4.

[0091] Step d1: Based on the target feature attributes of the target object, the sixth confidence level of each candidate object can be determined by the target feature attributes contained in the state description information of the target object and the third confidence level of the target feature attributes of each candidate object.

[0092] In one approach, the CLIP model can be used to determine the sixth confidence level of each candidate object as the target object. State description information containing the target object's characteristic attributes can be used as e. self Represented by x. Each candidate can be represented by x. i Indicated by i, which is a positive integer. By inputting state description information containing the target feature attributes of the target object and each candidate object into the CLIP model, the CLIP output result CLIP(e) can be obtained. self ,x i Therefore, the sixth confidence level for each candidate object to be the target object can be determined as: P(e self |x i )∝CLIP(e self ,x i ).

[0093] Step d2: For the target location area where the target object is located, the seventh confidence level is determined based on the second similarity between the state description information of the target object and the standard location area description information corresponding to each preset location area. Based on the seventh confidence level and the fourth confidence level that each candidate object is located in the target preset location area, the eighth confidence level is determined for each candidate object to be the target object.

[0094] Here, the accuracy of the target object's state description information containing the target's preset location region can first be determined. Specifically, based on the second similarity between the target object's state description information and the standard location region description information corresponding to each preset location region, the seventh confidence level can be determined as to whether the target object's state description information contains the target's preset location region. The seventh confidence level characterizes the probability that the target location region described by the user in the state description information is a certain preset location region.

[0095] Among them, the seventh confidence level of the target object's state description information, which includes the target's preset location region, can be represented by L(p). i e loc ) indicates; p i Indicates each preset location area; e loc This represents the state description information of the target object contained in the user-issued target object retrieval command. The fourth confidence score of each candidate object located in the preset target location region can be represented by P(p i |x i The eighth confidence level of each candidate object as the target object can be expressed as:

[0096] Step d3: For the target positional relationship between the target object and each other candidate object, the ninth confidence level of the target object's state description information containing the target preset positional relationship can be determined based on the third similarity between the target object's state description information and the standard positional relationship description information corresponding to each preset positional relationship; based on the ninth confidence level and the fifth confidence level of each candidate object having a target preset positional relationship with each other candidate object, the tenth confidence level of each candidate object being the target object can be determined.

[0097] Here, the accuracy of the target object's state description information containing candidate objects can be determined. Specifically, based on the third similarity between the target object's state description information and the standard positional relationship description information corresponding to each candidate object, the ninth confidence level can be determined. The ninth confidence level represents the probability that the target object described by the user in the state description information is a certain candidate object. Furthermore, the accuracy of the target object's state description information containing a target preset positional relationship can also be determined. Specifically, based on the fourth similarity between the target object's state description information and the standard positional relationship description information corresponding to each preset positional relationship, the tenth confidence level can be determined. The tenth confidence level represents the probability that the target positional relationship described by the user in the state description information is a certain preset positional relationship.

[0098] Among them, the ninth confidence level of the candidate object included in the state description information of the target object can be represented by P(x).j |e rel ) indicates; x j This represents the target object described by the user, where j is a positive integer; e rel This represents the state description information of the target object included in the user-issued target object retrieval command. The tenth confidence level of the target object's preset positional relationship within the state description information can be represented by L(r). i,j ,e rel ) indicates; r i,j Let i represent the target preset positional relationship, where i is a positive integer. The fifth confidence level for the existence of the target preset positional relationship between each candidate object and each of the other candidate objects is P(r). i,j |x i ,x j The eleventh confidence level of each candidate object as the target object can be expressed as:

[0099] Step d4: Finally, based on the sixth, eighth, and eleventh confidence levels, determine the second confidence level of each candidate object as the target object.

[0100] Here, the second confidence level of each candidate object can be expressed as: P(x i |e)∝P(x i )P(e self |x i )P(e loc |x i )P(e rel |x i Where e represents the state description information of the target object contained in the target object grabbing instruction; P(x i ) represents the probability that each candidate object belongs to its respective category.

[0101] S102: Based on the first confidence level and the second confidence level corresponding to each of the candidate objects, determine the target candidate object and the reward information after performing each preset action on the target candidate object; the preset action includes capturing the target candidate object and asking a question.

[0102] Here, in the decision-making module, by combining the first confidence level corresponding to each candidate object and the obtained second confidence level, the target candidate object that may meet the user's requirements and the reward information after performing each preset action on the target candidate object can be determined.

[0103] Here, a partially observable Markov Decision Process (POMDP) ​​framework can be used for decision-making. Preset actions may include identifying target candidates and posing questions. Posing questions can include asking questions about the target candidates and asking questions about the target object itself.

[0104] The reward information can be used to guide the robot to make the best decision in order to grasp the target object indicated by the user in the shortest possible time. Different reward information can be set for different preset actions. For example, if the robot grasps the target object, the reward information can be +10 points; if the robot grasps the wrong object, the reward information can be -10 points; if the robot asks a question, the reward information can be -1 point.

[0105] In practice, different questions can be set in advance, and preset response information can be set for different questions.

[0106] In one implementation, when the question is posed to a target candidate, the question is determined by the following steps: based on the twelfth confidence level of each preset feature attribute pointing to the target candidate and the third confidence level of the target candidate having each preset feature attribute, a reference feature attribute is selected from each preset feature attribute; and a question is generated that includes feature attribute description information of the reference feature attribute.

[0107] Here, in each iteration of question generation, reference feature attributes can be selected from each preset feature attribute. In each iteration, reference feature attributes can be selected by the maximum product of the twelfth confidence level of each preset feature attribute pointing to the target candidate object and the third confidence level of the target candidate object for each preset feature attribute. In the next iteration, the reference feature attribute with the largest product of the twelfth confidence level of each preset feature attribute pointing to the target candidate object and the third confidence level of the target candidate object for each preset feature attribute is selected from the remaining preset feature attributes.

[0108] Here, the reference feature attribute for each filtering can be represented by a. t express, Wherein, P(x i |a t P(a, ..., a1) represents the twelfth confidence level of each preset feature attribute pointing to the target candidate object, where P(a...a1) represents the twelfth confidence level of each preset feature attribute pointing to the target candidate object. t |x i ) indicates that the target candidate object has the third confidence level for each preset feature attribute.

[0109] S103: Based on the reward information corresponding to each preset action, determine the target preset action to be executed, and execute the target preset action for the target candidate object.

[0110] When the preset target and preset action is to grab a target candidate object, if the target candidate object is the target object, the grabbing task is successful and the task ends; if the target candidate object is not the target object, the grabbing task fails, and the user can reissue the target object grabbing command or end the task.

[0111] When the predetermined target and predetermined action is to raise a question, raising a question can include raising a question for the target candidate or raising a question for the target object.

[0112] In this context, when asking questions, if the user's response information is used to determine whether the target candidate is the target object, the target candidate object can be retrieved if it is determined to be the target object.

[0113] In this context, posing questions includes asking questions specifically to the target audience. Users can then respond to these questions, providing either a direct answer or supplementary descriptive information. For example, if the question is "Do you want the blue box?", the user could answer "No, I want a red box." Here, "No" represents the direct answer, and "I want a red box" represents the supplementary descriptive information. The robot can process the direct answer and supplementary descriptive information separately.

[0114] In one implementation, after executing the target preset action, in response to receiving the response information to the question raised, a fifth similarity between the response information and each preset response information can be determined; based on the fifth similarity, a first confidence that each candidate object in the current scene has each preset state is updated; based on the updated first confidence, a second confidence that each candidate object is the target object is updated.

[0115] Here, the fifth similarity between the determined response information and each preset response information can be normalized and used as the first confidence score for each candidate object in the current scene to have each preset state. After the first confidence score is updated, the second confidence score of each candidate object as the target object can be updated based on the updated first confidence score, and subsequent steps can continue to be executed until the target object is captured, or the task ends after a certain number of dialogues.

[0116] In one implementation, after performing the target preset action, in response to receiving supplementary description information for the question raised, the second confidence level of each candidate object can be updated based on the supplementary description information and the first confidence level, and the step of determining the reward information can be returned.

[0117] Receiving supplementary description information for the question raised is equivalent to receiving new status description information. Here, the second confidence level of each candidate object can be re-determined as the target object, and the step of determining the reward information can be returned until the target object is captured or the task ends after the dialogue reaches a preset number of times.

[0118] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0119] Based on the same inventive concept, this disclosure also provides an object grasping device corresponding to the object grasping method. Since the principle of the device in this disclosure for solving the problem is similar to the object grasping method described above in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0120] like Figure 4 The diagram illustrates another object grasping method. First, the robot's camera captures scene images of the current environment; specifically, it can capture both a 2D scene image and a 3D scene image. Object detection is then performed on the 3D scene image to identify candidate objects. Finally, the 2D scene image is segmented for each candidate object to obtain its corresponding image region.

[0121] Next, based on the first similarity between the image region corresponding to the candidate object and each preset color, the first confidence level of the candidate object having each preset color is determined.

[0122] The second confidence level of the candidate object is determined based on the position range of the bounding box of the candidate object and the overlap area between the bounding box and each preset grid region in the two-dimensional scene image.

[0123] Based on the point cloud information of each candidate object in the 3D scene image, the shortest distance vector between each candidate object and all other candidate objects is determined. Then, based on the projection length of the shortest distance vector in each preset direction, a third confidence level is determined to determine the existence of preset positional relationships between the candidate object and other candidate objects. These preset directions can include the front, back, left, and right directions in a two-dimensional plane.

[0124] An observation model can be constructed by using a first confidence level that candidate objects possess each preset color, a second confidence level that candidate objects are located in each preset grid, and a third confidence level that each candidate object has each preset positional relationship with other candidate objects. This observation model is used to predict the reward information after the robot executes each preset action, based on a fourth confidence level that each candidate object is the target object, after the user issues a target object grasping command.

[0125] After a user issues a target object retrieval command, the fourth confidence level can be determined based on the target object's state description information contained in the retrieval command, as well as the aforementioned first, second, and third confidence levels. For example, if the user's target object retrieval command is "Please give me a pair of pliers," then the fourth confidence level for each candidate object can be determined as "pliers" based on the above process.

[0126] Next, based on the observation model and the fourth confidence level that each candidate object is a "pincer," the reward information after the robot performs various preset actions is determined. These preset actions may include asking questions about attribute characteristics (e.g., "What color is it?", "Is it color XX?"), asking questions about location range (e.g., "Where is it?", "Is it in direction XX?"), asking questions about location relationships (e.g., "In which direction of XX is it?", "Is it in direction XX?"), confirming the target candidate object after asking questions, and grasping the target candidate object.

[0127] After the robot performs different preset actions, it can receive different reward information. For example, it can get -1 point after asking a question; it can get +10 points after the target candidate object is the target object; and it can get -10 points after the target candidate object is not the target object.

[0128] Based on the predicted feedback information after performing each preset action, the robot can determine the target preset action, i.e., the optimal preset action. Here, determining the target preset action to ask the user a question about the positional relationship can be done by asking the user, "In which direction is it facing these red and black pincers?"

[0129] Once the user provides a response to the question, the robot can continue to determine the target preset action based on the observation model and the response information, until the target object is determined to be grasped, or the number of questions asked reaches the preset number, at which point the task ends.

[0130] Reference Figure 5 The diagram shown is an architectural schematic of an object grasping device provided in an embodiment of this disclosure. The device includes:

[0131] The first determining module 501 is used to respond to receiving a target object grabbing instruction, and determine the second confidence level of each candidate object as the target object based on the target object state description information contained in the target object grabbing instruction and the first confidence level of each candidate object in the current scene having each preset state.

[0132] The second determining module 502 is used to determine target candidate objects and reward information after performing preset actions on the target candidate objects based on the first confidence level and the second confidence level corresponding to each of the candidate objects respectively; the preset actions include capturing the target candidate object and asking questions;

[0133] The third determining module 503 is used to determine the target preset action to be executed based on the reward information corresponding to each preset action, and to execute the target preset action for the target candidate object.

[0134] In one feasible implementation, the first confidence level of the candidate object having each preset state includes: a third confidence level of the candidate object having each preset feature attribute, a fourth confidence level of the candidate object being located in different preset location regions, and a fifth confidence level of the candidate object having each preset location relationship with each other candidate object.

[0135] In one optional implementation, the apparatus further includes a fourth determining module for determining the third confidence level;

[0136] The fourth determining module is specifically used for:

[0137] The two-dimensional scene image corresponding to the current scene is segmented to obtain the image region corresponding to each of the candidate objects;

[0138] Based on the first similarity between the image region corresponding to the candidate object and the description information of each preset feature attribute, the third confidence level of the candidate object with each preset feature attribute is determined.

[0139] In one optional implementation, the apparatus further includes a fifth determining module for determining the fourth confidence level;

[0140] The fifth determining module is specifically used for:

[0141] Obtain the two-dimensional scene image corresponding to the current scene;

[0142] Based on the position range of the candidate object in the two-dimensional scene image and the overlap area between the candidate object and each preset position region in the two-dimensional scene image, a fourth confidence level is determined for the candidate object to be located in each preset position region.

[0143] In one optional implementation, the apparatus further includes a sixth determining module for determining the fifth confidence level;

[0144] The sixth determining module is specifically used for:

[0145] Based on the three-dimensional position information of each candidate object in the three-dimensional scene image corresponding to the current scene, the shortest distance vector between each candidate object and the other candidate objects is determined.

[0146] Based on the projection length of the shortest distance vector in each preset direction, a fifth confidence level is determined to indicate that there are preset positional relationships between the candidate object and other candidate objects.

[0147] In one optional implementation, the first determining module 501 is specifically used for:

[0148] Based on the target feature attributes contained in the state description information of the target object, and the third confidence level of each candidate object having the target feature attributes, each candidate object is determined as the sixth confidence level of the target object;

[0149] Based on the second similarity between the state description information of the target object and the standard location area description information corresponding to each preset location area, a seventh confidence level is determined that the state description information of the target object contains the target preset location area; based on the seventh confidence level and the fourth confidence level that each candidate object is located in the target preset location area, an eighth confidence level is determined that each candidate object is the target object;

[0150] Based on the third similarity between the state description information of the target object and the standard candidate object description information corresponding to each candidate object, a ninth confidence level is determined that the state description information of the target object contains the target candidate object; based on the fourth similarity between the state description information of the target object and the standard positional relationship description information corresponding to each preset positional relationship, a tenth confidence level is determined that the state description information of the target object contains the target preset positional relationship; based on the ninth confidence level, the tenth confidence level, and the fifth confidence level that each candidate object has the target preset positional relationship with each of the other candidate objects, an eleventh confidence level is determined that each candidate object is the target object.

[0151] Based on the sixth confidence level, the eighth confidence level, and the eleventh confidence level, the second confidence level of each candidate object is determined as that of the target object.

[0152] In one optional implementation, the question-posing includes posing a question for the target candidate object; wherein the posed question includes feature attribute description information referencing feature attributes; the device further includes a seventh determining module for determining the question;

[0153] The seventh determining module is specifically used for:

[0154] Based on the twelfth confidence level of each preset feature attribute pointing to the target candidate object, and the third confidence level of the target candidate object having each preset feature attribute, the reference feature attribute is selected from each preset feature attribute;

[0155] The problem is to generate feature attribute description information that includes the reference feature attributes.

[0156] In one optional implementation, the target preset action includes asking the question; asking the question includes asking a question to the target object;

[0157] After the step of performing the target preset action, the device includes:

[0158] The eighth determining module is used to determine the fifth similarity between the received response information and each preset response information in response to receiving the response information for the question.

[0159] The first update module is used to update the first confidence level of each candidate object in the current scene with each preset state based on the fifth similarity;

[0160] The second update module is used to update the second confidence of each candidate object as the target object based on the updated first confidence.

[0161] In one optional implementation, the target preset action includes asking the question; asking the question includes asking a question to the target object;

[0162] After the step of performing the target preset action, the device includes:

[0163] The third update module, in response to receiving supplementary description information for the problem, updates the second confidence level of each candidate object to the target object based on the supplementary description information and the first confidence level, and returns the step of determining the reward information.

[0164] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0165] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 6 The diagram shows the structure of a computer device 600 provided in this embodiment of the present disclosure, including a processor 601, a memory 602, and a bus 603. The memory 602 stores execution instructions and includes main memory 6021 and external memory 6022. The main memory 6021, also called internal memory, is used to temporarily store computational data in the processor 601 and data exchanged with external memory 6022 such as a hard disk. The processor 601 exchanges data with the external memory 6022 through the main memory 6021. When the computer device 600 is running, the processor 601 and the memory 602 communicate through the bus 603, causing the processor 601 to execute the following instructions:

[0166] In response to receiving a target object grabbing instruction, based on the target object's state description information contained in the target object grabbing instruction and the first confidence level of each candidate object in the current scene having each preset state, the second confidence level of each candidate object is determined to be the target object;

[0167] Based on the first confidence level and the second confidence level corresponding to each of the candidate objects, target candidate objects are determined, along with reward information after performing each preset action on the target candidate object; the preset actions include capturing the target candidate object and asking a question;

[0168] Based on the reward information corresponding to each preset action, the target preset action to be executed is determined, and the target preset action is executed for the target candidate object.

[0169] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the object-grabbing method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0170] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the object grabbing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0171] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0172] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0173] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0174] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0175] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0176] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. An object crawling method, characterized in that, include: In response to receiving a target object grabbing instruction, based on the target object's state description information contained in the target object grabbing instruction and the first confidence level of each candidate object in the current scene having each preset state, the second confidence level of each candidate object is determined to be the target object; Based on the first confidence level and the second confidence level corresponding to each of the candidate objects, a target candidate object is determined, along with the reward information after performing each preset action on the target candidate object. The preset actions include capturing the target candidate object and asking a question; Based on the reward information corresponding to each preset action, the target preset action to be executed is determined, and the target preset action is executed for the target candidate object.

2. The method according to claim 1, characterized in that, The candidate object has a first confidence level for each preset state, including: The candidate objects each have a third confidence level for each preset feature attribute, a fourth confidence level for each candidate object located in different preset location regions, and a fifth confidence level for each preset location relationship between each candidate object and each other.

3. The method according to claim 2, characterized in that, The third confidence level is obtained through the following steps: The two-dimensional scene image corresponding to the current scene is segmented to obtain the image region corresponding to each of the candidate objects; Based on the first similarity between the image region corresponding to the candidate object and the description information of each preset feature attribute, the third confidence level of the candidate object with each preset feature attribute is determined.

4. The method according to claim 2, characterized in that, The fourth confidence level is obtained through the following steps: Obtain the two-dimensional scene image corresponding to the current scene; Based on the position range of the candidate object in the two-dimensional scene image and the overlap area between the candidate object and each preset position region in the two-dimensional scene image, a fourth confidence level is determined for the candidate object to be located in each preset position region.

5. The method according to claim 2, characterized in that, The fifth confidence level is obtained through the following steps: Based on the three-dimensional position information of each candidate object in the three-dimensional scene image corresponding to the current scene, the shortest distance vector between each candidate object and the other candidate objects is determined. Based on the projection length of the shortest distance vector in each preset direction, a fifth confidence level is determined to indicate the existence of preset positional relationships between the candidate object and other candidate objects.

6. The method according to claim 2, characterized in that, The step of determining the second confidence level of each candidate object as the target object based on the state description information of the target object contained in the target object capture instruction and the first confidence level of each candidate object having each preset state in the current scene includes: Based on the target feature attributes contained in the state description information of the target object, and the third confidence level of each candidate object having the target feature attributes, each candidate object is determined as the sixth confidence level of the target object; Based on the second similarity between the state description information of the target object and the standard location area description information corresponding to each preset location area, a seventh confidence level is determined that the state description information of the target object contains the target preset location area; based on the seventh confidence level and the fourth confidence level that each candidate object is located in the target preset location area, an eighth confidence level is determined that each candidate object is the target object; Based on the third similarity between the state description information of the target object and the standard candidate object description information corresponding to each candidate object, a ninth confidence level is determined that the state description information of the target object contains the target candidate object; based on the fourth similarity between the state description information of the target object and the standard positional relationship description information corresponding to each preset positional relationship, a tenth confidence level is determined that the state description information of the target object contains the target preset positional relationship; based on the ninth confidence level, the tenth confidence level, and the fifth confidence level that each candidate object has the target preset positional relationship with each of the other candidate objects, an eleventh confidence level is determined that each candidate object is the target object. Based on the sixth confidence level, the eighth confidence level, and the eleventh confidence level, the second confidence level of each candidate object is determined as that of the target object.

7. The method according to claim 2, characterized in that, The process of raising questions includes raising questions for the target candidate object; wherein, the raised questions include feature attribute description information referencing feature attributes; the questions are determined through the following steps: Based on the twelfth confidence level of each preset feature attribute pointing to the target candidate object, and the third confidence level of the target candidate object having each preset feature attribute, the reference feature attribute is selected from each preset feature attribute; The problem is to generate feature attribute description information that includes the reference feature attributes.

8. The method according to claim 1, characterized in that, The target preset action includes asking the question; asking the question includes asking a question to the target object; After executing the target preset action, the method further includes: In response to receiving a response to the question, a fifth similarity is determined between the received response and each preset response. Based on the fifth similarity, update the first confidence level of each candidate object in the current scene for each preset state; Based on the updated first confidence level, the second confidence level of each candidate object is updated to that of the target object.

9. The method according to claim 1, characterized in that, The target preset action includes asking the question; asking the question includes asking a question to the target object; After executing the target preset action, the method further includes: In response to receiving supplementary description information for the problem, based on the supplementary description information and the first confidence level, the second confidence level of each candidate object is updated to be the target object, and the step of determining the reward information is returned.

10. An object grasping device, characterized in that, include: The first determining module is used to respond to receiving a target object grabbing instruction, and determine the second confidence level of each candidate object as the target object based on the target object's state description information contained in the target object grabbing instruction and the first confidence level of each candidate object in the current scene having each preset state. The second determining module is used to determine the target candidate object and the reward information after performing each preset action on the target candidate object based on the first confidence level and the second confidence level corresponding to each of the candidate objects respectively; The preset actions include capturing the target candidate object and asking a question; The third determining module is used to determine the target preset action to be executed based on the reward information corresponding to each preset action, and to execute the target preset action for the target candidate object.

11. A computer device, characterized in that, include: The computer device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and the processor communicates with the memory via the bus when the computer device is running, and the machine-readable instructions, when executed by the processor, perform the steps of the object grasping method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the object grasping method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Object recognition method and device based on image recognition model and electronic equipment

    CN112633384A

  • Network training method and device, robot control method and device, equipment and storage medium

    CN114397817A