Object information recognition method, path planning method, device, terminal and robot

Through the segmentation processing of two-dimensional image and depth information and neural network model recognition, the problem of low object information recognition efficiency is solved, and fast and accurate object information acquisition is achieved, which is suitable for robot obstacle avoidance and path planning.

CN115167409BActive Publication Date: 2025-08-22ANKER INNOVATIONS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210760127.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-08-22
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

In the prior art, object information recognition efficiency is low, and the recognition time is mainly caused by measuring the depth information of an object circle.

Method used

By obtaining the two-dimensional image and depth information of the preset area, performing segmentation processing, obtaining the target depth information and target pixel information of the target object, and using a neural network model to identify object information, including identification of volume size, object posture and position information.

Benefits of technology

It realizes that object information can be quickly obtained without measuring the depth information of an object without measuring the object in one circle, which improves recognition efficiency and is suitable for robot obstacle avoidance and path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115167409B_ABST
    Figure CN115167409B_ABST
Patent Text Reader

Abstract

The present application relates to an object information recognition method, a path planning method, a device, a terminal, and a robot. The method comprises: obtaining a two-dimensional image and depth information obtained by image acquisition of a preset area; performing segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to a target object located in the preset area and target pixel information of the target object in the two-dimensional image; and performing information recognition on the target object based on the target pixel information and the target depth information to obtain object information of the target object, wherein the object information is used to indicate the space occupied by the target object in the preset area. The use of this method can improve the efficiency of object information recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information recognition technology, and in particular to an object information recognition method, path planning method, device, terminal and robot. Background Art

[0002] With the development of information recognition technology, there are many scenarios where object information needs to be recognized. For example, cleaning products such as sweepers need to identify the object information of obstacles during the cleaning process in order to perform obstacle avoidance.

[0003] In the related art, a depth sensor is mainly used to measure the depth information of a circle of the object to be identified, and the object information of the object is determined based on the depth information of the circle of the object.

[0004] However, it takes a long time to measure the depth information of an object to be identified, resulting in low efficiency in object information identification. Summary of the Invention

[0005] Based on this, it is necessary to provide an object information recognition method, path planning method, device, terminal and robot that can improve the efficiency of object information recognition in order to address the above technical problems.

[0006] In a first aspect, the present application provides a method for identifying object information, comprising:

[0007] Acquire a two-dimensional image and depth information obtained by image acquisition of a preset area;

[0008] performing segmentation processing according to the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image;

[0009] Information recognition is performed on the target object according to the target pixel information and the target depth information to obtain object information of the target object, where the object information is used to indicate a space occupied by the target object in the preset area.

[0010] In one embodiment, the object information includes volume size, object posture, and position information of the target object in the preset area, and the performing information recognition on the target object based on the target pixel information and the target depth information to obtain the object information of the target object includes:

[0011] Obtaining a target object category of the target object;

[0012] The target object is identified according to the target object category, the target pixel information and the target depth information to obtain the volume size, object posture and position information of the target object.

[0013] In one embodiment, obtaining the target object category of the target object includes:

[0014] Inputting the target pixel information into a trained object category recognition model to obtain the object category and confidence level output by the object category recognition model, wherein the confidence level is used to indicate the credibility of the object category output by the object category recognition model;

[0015] If the confidence level is higher than a preset confidence level, the object category output by the object category recognition model is used as the target object category of the target object.

[0016] In one embodiment, the performing information recognition on the target object based on the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes:

[0017] Generate a first target three-dimensional feature descriptor according to the target object category, the target pixel information and the target depth information;

[0018] Inputting the first target three-dimensional feature descriptor into a trained object feature extraction model to obtain the centroid position of the target object and a first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position of the target object output by the object feature extraction model, wherein the object feature extraction model is obtained by training the first three-dimensional feature descriptor as input to a neural network model and using the centroid position of the object and the relative position relationship between the first three-dimensional feature descriptor and the centroid position of the object as outputs of the neural network model, wherein the first three-dimensional feature descriptor is generated based on object category, pixel information, and depth information;

[0019] determining a volume size of the target object according to the first target three-dimensional feature descriptor;

[0020] Determining the position information of the target object according to the centroid position of the target object;

[0021] The object posture of the target object is determined according to the relative position relationship of the first target.

[0022] In one embodiment, identifying the volume size, object posture, and position information of the target object based on the target three-dimensional feature descriptor and the target relative position relationship includes:

[0023] The volume size of the target object is determined according to a matching result of the first target three-dimensional feature descriptor in a global coordinate system, where the global coordinate system is a coordinate system of the preset area.

[0024] In one embodiment, the performing information recognition on the target object based on the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes:

[0025] According to the target object category, searching a database for a target volume size corresponding to the target object category, and using the target volume size as the volume size of the target object, wherein the database includes volume sizes corresponding to different object categories;

[0026] The object posture and position information of the target object are determined according to the target pixel information and the target depth information.

[0027] In one embodiment, determining the object posture and position information of the target object according to the target pixel information and the target depth information includes:

[0028] generating a second target three-dimensional feature descriptor according to the target pixel information and the target depth information;

[0029] Inputting the second target three-dimensional feature descriptor into a trained posture recognition model corresponding to the target object category to obtain the centroid position of the target object output by the posture recognition model and a second target relative position relationship between the second target three-dimensional feature descriptor and the centroid position of the target object, wherein the posture recognition model is trained for each object type by using the second three-dimensional feature descriptor as input to a neural network model and using the centroid position of the object and the relative position relationship between the second three-dimensional feature descriptor and the centroid position of the object as outputs of the neural network model, wherein the second three-dimensional feature descriptor is generated based on pixel information and depth information;

[0030] Determining the position information of the target object according to the centroid position of the target object;

[0031] The object posture of the target object is determined according to the relative position relationship of the second target.

[0032] In one embodiment, the performing segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image includes:

[0033] determining a segmentation position between a foreground and a background of the two-dimensional image according to the depth information;

[0034] Acquiring image position information of the target object in the two-dimensional image;

[0035] Segmentation processing is performed according to the segmentation position and the image position information to obtain the target depth information and the target pixel information.

[0036] In a second aspect, the present application provides a path planning method, comprising:

[0037] Acquire a two-dimensional image and depth information obtained by image acquisition of a preset area;

[0038] performing segmentation processing according to the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image;

[0039] performing information recognition on the target object according to the target pixel information and the target depth information to obtain object information of the target object, where the object information is used to indicate a space occupied by the target object in the preset area;

[0040] The robot's travel path is planned according to the object information.

[0041] In a third aspect, the present application provides an object information recognition device, comprising:

[0042] An acquisition module is used to acquire a two-dimensional image and depth information obtained by collecting images of a preset area;

[0043] a segmentation processing module, configured to perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to a target object located in the preset area and target pixel information of the target object in the two-dimensional image;

[0044] An information recognition module is used to perform information recognition on the target object according to the target pixel information and the target depth information to obtain object information of the target object, where the object information is used to indicate a space occupied by the target object in the preset area.

[0045] In a fourth aspect, the present application provides a path planning device, comprising:

[0046] An acquisition module is used to acquire a two-dimensional image and depth information obtained by collecting images of a preset area;

[0047] a segmentation processing module, configured to perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to a target object located in the preset area and target pixel information of the target object in the two-dimensional image;

[0048] an information recognition module, configured to perform information recognition on the target object based on the target pixel information and the target depth information to obtain object information of the target object, wherein the object information is used to indicate a space occupied by the target object in the preset area;

[0049] The path planning module is used to plan the robot's travel path according to the object information.

[0050] In a fifth aspect, the present application provides a terminal comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0051] In a sixth aspect, the present application provides a robot, comprising:

[0052] A first image acquisition device is used to acquire an image of a preset area to obtain a two-dimensional image;

[0053] The second image acquisition device is used to acquire images of a preset area to obtain depth information;

[0054] A processor is used to implement the steps of the above method.

[0055] In a seventh aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0056] In an eighth aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.

[0057] The above-mentioned object information recognition method, path planning method, device, terminal and robot obtain a two-dimensional image and depth information obtained by image acquisition of a preset area; perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image; perform information recognition on the target object based on the target pixel information and the target depth information to obtain object information of the target object, and the object information is used to indicate the space occupied by the target object in the preset area; since the object information of the target object located in the preset area can be obtained through the two-dimensional image and depth information of the preset area, the object information of the object can be obtained without measuring the depth information of a circle of the object, thereby achieving the technical effect of improving the efficiency of object information recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 A schematic diagram of an application environment of an object information recognition method provided in an embodiment of the present application;

[0059] Figure 2 A schematic diagram of an application environment of another object information recognition method provided in an embodiment of the present application;

[0060] Figure 3 A flowchart of an object information recognition method provided in an embodiment of the present application;

[0061] Figure 4 A detailed flow chart of identifying the target object based on the target pixel information and the target depth information provided in an embodiment of the present application;

[0062] Figure 5 A detailed flowchart of segmentation processing based on the two-dimensional image and the depth information provided in an embodiment of the present application;

[0063] Figure 6 A flowchart of an object information recognition method provided in an embodiment of the present application;

[0064] Figure 7 A schematic diagram of the structure of an object information recognition device provided in an embodiment of the present application;

[0065] Figure 8 A schematic diagram of the structure of a path planning device provided in an embodiment of the present application;

[0066] Figure 9 This is a diagram of the internal structure of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0068] refer to Figure 1 , Figure 1 Schematic diagram of the application environment of an object information recognition method provided in an embodiment of the present application. Figure 1 As shown, the terminal of this embodiment includes a first image acquisition device 110 and a second image acquisition device 120 .

[0069] In this embodiment, the first image acquisition device 110 is used to acquire images of a preset area to obtain a two-dimensional image. The second image acquisition device 120 is used to acquire images of the preset area to obtain depth information. The terminal processor recognizes object information based on the two-dimensional image and depth information.

[0070] Optionally, terminals may include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, portable wearable devices, and servers. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. The server may be implemented as a standalone server or a server cluster consisting of multiple servers.

[0071] Optionally, the first image acquisition device 110 may include but is not limited to a conventional camera that acquires two-dimensional images, etc., and this embodiment does not limit this. The second image acquisition device 120 may include but is not limited to a tof (time of flight) camera and a depth sensor that can acquire depth information through image acquisition, etc., and this embodiment does not limit this. Generally, depth sensors include active depth sensors and passive depth sensors. Active depth sensor: A sensor that detects distance depth information by actively emitting energy (including but not limited to electromagnetic waves, microwaves, etc.). Passive depth sensor: A sensor that detects distance depth information by receiving energy (electromagnetic signals, etc.) reflected or emitted from the outside world.

[0072] For example, in this embodiment, if a user wants to know the object information of an object in a preset area, the user can use the terminal to capture an image of the preset area to obtain a two-dimensional image and depth information. At this time, the object information of the object can be identified based on the two-dimensional image and depth information.

[0073] It should be noted that the terminal may not need to be provided with the first image acquisition device 110 and the second image acquisition device 120. In this case, the preset area image is acquired by other image acquisition devices to obtain a two-dimensional image and depth information, and the terminal can then perform object information recognition based on the two-dimensional image and depth information.

[0074] refer to Figure 2 , Figure 2 This is a schematic diagram of an application environment for another object information recognition method provided in an embodiment of the present application. Figure 2 As shown, the robot of this embodiment includes a first image acquisition device 210 and a second image acquisition device 220 .

[0075] In this embodiment, robots include but are not limited to sweeping robots, cargo handling robots, etc., and this embodiment does not impose any restrictions.

[0076] Exemplarily, the first image acquisition device 210 and the second image acquisition device 220 are used to capture images of a preset area in the direction of travel of the robot, thereby obtaining a two-dimensional image and depth information in front of the robot. The object information of obstacles in the direction of travel can be determined based on the two-dimensional image and depth information, thereby performing obstacle avoidance processing.

[0077] It is understandable that the object information recognition method provided in the embodiments of the present application can be applied to scenarios including but not limited to the above-mentioned scenarios.

[0078] refer to Figure 3 , Figure 3 This is a flow chart of an object information recognition method provided in an embodiment of the present application. Figure 1 Terminal or Figure 2 In one embodiment, the robot in FIG. Figure 3 As shown, the object information recognition method includes:

[0079] Step 310: Acquire a two-dimensional image and depth information obtained by capturing images of a preset area.

[0080] The preset area refers to the area where the first image acquisition device and the second image acquisition device can capture images. Generally, a two-dimensional image is a planar image that does not contain depth information. Two-dimensional refers to the four directions of left, right, up, and down, without front and back. The content on a piece of paper can be considered two-dimensional. Depth information is used to indicate the distance between the second image acquisition device and various positions in the preset area. Specifically, the depth information of the preset area can be obtained using a TOF camera.

[0081] Step 320 : Perform segmentation processing according to the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image.

[0082] In this embodiment, the target object can be an object selected by the user from the two-dimensional image, or any object in a preset area captured by the two-dimensional image can be used as the target object. This embodiment does not limit this. Optionally, the target object can be a person or an object. In this embodiment, people and objects are collectively referred to as objects. For example, in Figure 1 In the application environment, suppose there is a table and a ball in the preset area, and the captured two-dimensional image includes the table and the ball. If the user selects the table, the object information of the table is recognized; if the user selects the ball, the object information of the ball is recognized; or the table and the ball are respectively selected as the target objects, and the object information of the table and the object information of the ball are recognized respectively. For another example, in Figure 2 In an application environment, since the robot needs to move, obstacle avoidance is required. Any object in the preset area captured by the two-dimensional image can be used as the target object, and the target depth information refers to the depth information corresponding to the target object. Specifically, since the depth information of the preset area is obtained, and generally, the object does not occupy the entire preset area, it is necessary to obtain the target depth information corresponding to the target object. Target pixel information refers to the pixel information of the target object in the two-dimensional image. Specifically, the target pixel information may include but is not limited to one or more of the number of pixels occupied by the target object, the occupied pixel area, the position of the occupied pixel area in the image, and the RGB (color system) value of each occupied pixel point.

[0083] Segmentation processing refers to the processing performed to obtain the target depth information and target pixel information corresponding to the target object. Specifically, the segmentation processing can be to segment the regional position of the object in the preset area, so as to determine the target pixel information and target depth information corresponding to the target object based on the regional position. Optionally, the segmentation processing can be performed by classical algorithms, modern AI-related algorithms, and other algorithms. Classic algorithms mainly include threshold segmentation, region growing, watershed, edge detection, wavelet transform, and genetic algorithm; modern AI-related models mainly include MASK RCNN, SetNet, PSPNetd, FCN, Unet, SegNet, etc.

[0084] Step 330: Perform information recognition on the target object according to the target pixel information and the target depth information to obtain object information of the target object.

[0085] The object information indicates the space occupied by the target object within the preset area. It should be noted that if two-dimensional object information is to be determined, the occupancy information may be area. If three-dimensional object information is to be determined, the occupancy information may include, but is not limited to, at least one of the following: the target object's volume, its posture, and its position within the preset area. This embodiment does not impose any limitation thereto.

[0086] The technical solution of this embodiment can obtain the object information of the target object located in the preset area through the two-dimensional image and depth information of the preset area, and can obtain the object information of the object without measuring the depth information of a circle of the object, thereby achieving the technical effect of improving the efficiency of object information recognition.

[0087] It should be noted that, in this embodiment, the two-dimensional image and depth information may be obtained by capturing an image of a preset area at a single angle.

[0088] refer to Figure 4 , Figure 4 A detailed flow chart of identifying the target object according to the target pixel information and the target depth information is provided in an embodiment of the present application. In one embodiment, the object information includes volume size, object posture and position information, such as Figure 4 As shown, the performing information recognition on the target object according to the target pixel information and the target depth information to obtain object information of the target object includes:

[0089] Step 410: Obtain the target object category of the target object.

[0090] The target object category refers to the category to which the target object belongs; for example, it may be a table, a chair, a ball, etc., which is determined according to the category actually described by the target object.

[0091] Step 420 : performing information recognition on the target object according to the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object.

[0092] Volume dimensions refer to the shape parameters of the target object. Specifically, they can reflect the target object's outline, shape, and size. Object pose indicates how the target object is positioned, such as its orientation. Position information indicates the specific location of the target object within a preset area.

[0093] In this embodiment, information of the target object is identified based on the target object category, target pixel information and target depth information to obtain the volume size, object posture and position information of the target object. Since the volume size, object posture and position information of the target object are obtained through multiple parameters, the obtained volume size, object posture and position information are more accurate.

[0094] In one possible implementation, obtaining the target object category of the target object includes:

[0095] Inputting the target pixel information into a trained object category recognition model to obtain the object category and confidence level output by the object category recognition model, wherein the confidence level is used to indicate the credibility of the object category output by the object category recognition model;

[0096] If the confidence level is higher than a preset confidence level, the object category output by the object category recognition model is used as the target object category of the target object.

[0097] Among them, the object recognition model can be trained by using pixel information as the input of the object category recognition model and the object category as the output of the object category recognition model, thereby obtaining a trained object category recognition model. Optionally, the object category recognition model includes but is not limited to the yolo series model, the SSD series model, the ResNet, the MobildeNet series model, etc. The preset confidence level can be set as needed, for example, to 80%. In this embodiment, optionally, if the confidence level is lower than the preset confidence level, the target pixel information is reacquired and the new target pixel information is input into the object category recognition model until the confidence level is higher than the preset confidence level.

[0098] The technical solution of this embodiment uses an object category recognition model to identify the target object category, and only when the confidence level is higher than a preset confidence level is the object category output by the object category recognition model used as the target object category of the target object. This improves the ease of obtaining the target object category and also improves the accuracy of the obtained target object category.

[0099] In one possible implementation, performing information recognition on the target object based on the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes:

[0100] generating a first target three-dimensional feature descriptor according to the target object category, the target pixel information, and the target depth information;

[0101] Inputting the first target three-dimensional feature descriptor into a trained object feature extraction model to obtain the centroid position of the target object and a first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position of the target object output by the object feature extraction model, wherein the object feature extraction model is obtained by training the first three-dimensional feature descriptor as input to a neural network model and using the centroid position of the object and the relative position relationship between the first three-dimensional feature descriptor and the centroid position of the object as outputs of the neural network model, wherein the first three-dimensional feature descriptor is generated based on object category, pixel information, and depth information;

[0102] determining a volume size of the target object according to the first target three-dimensional feature descriptor;

[0103] Determining the position information of the target object according to the centroid position of the target object;

[0104] The object posture of the target object is determined according to the relative position relationship of the first target.

[0105] The first target 3D feature descriptor is generated based on the target object category, target pixel information, and target depth information. Specifically, the first target 3D feature descriptor includes the target object category, target pixel information, and target depth information. Object feature extraction models include, but are not limited to, one or more of a convolutional neural network, a residual neural network, a radial basis function neural network, and a Hopfield network. 3D feature descriptors are often used to represent key features of 3D perception data of an object. 3D feature descriptors include local feature descriptors and global feature descriptors. The information encoded by a 3D feature descriptor primarily falls into two categories: one relates to the shape information of a local surface, including the spatial distribution of points and geometric properties. The other includes both color and shape information of the local surface. The center of mass refers to physical objects, while the centroid refers to abstract geometric entities. For physical objects with uniform density, the center of mass and centroid coincide. In this embodiment, the centroid is the geometric center of the target object. The centroid position refers to the location of the centroid within a preset area. Optionally, the centroid position can be expressed as coordinates.

[0106] In this embodiment, the centroid position of the target object and the first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position of the target object are obtained through the object feature extraction model, and then the volume size of the target object is determined according to the first target three-dimensional feature descriptor, the position information of the target object is determined according to the centroid position of the target object, and the object posture of the target object is determined according to the first target relative position relationship.

[0107] Specifically, the volume of the target object can be determined based on the first target 3D feature descriptor, as a key feature used to represent the 3D perception data of the object. Alternatively, the position information of the target object can be determined based on the centroid position of the target object. Alternatively, the relative position relationship of the first target is the object pose of the target object.

[0108] It should be noted that a model that can directly output volume size, object posture and position information can also be trained, and this embodiment does not limit this.

[0109] In one possible implementation, determining the volume size of the target object according to the first target three-dimensional feature descriptor includes:

[0110] The volume size of the target object is determined according to a matching result of the first target three-dimensional feature descriptor in a global coordinate system, where the global coordinate system is a coordinate system of the preset area.

[0111] In this embodiment, the first target three-dimensional feature descriptor includes the spatial distribution information of the key points of the target object. Based on the matching results of the spatial distribution information of the key points in the global coordinate system, the coordinates of the key points can be determined, and the volume size of the target object can be determined based on the coordinates. Among them, the key points can be points that can reflect the size and volume of the target object. For example, for a table, the key points are the intersection points of the edges of the table. This embodiment does not limit this. In this embodiment, the relative position relationship of the target can be regarded as the object posture of the target object.

[0112] In one possible implementation, performing information recognition on the target object based on the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes:

[0113] According to the target object category, searching a database for a target volume size corresponding to the target object category, and using the target volume size as the volume size of the target object, wherein the database includes volume sizes corresponding to different object categories;

[0114] The object posture and position information of the target object are determined according to the target pixel information and the target depth information.

[0115] In this embodiment, the target volume size corresponding to the target object category found in the database is used as the target object's volume size. This allows the volume size to be determined without the need for a trained model, reducing the computing power required for volume size determination. After the target object's volume size is found in the database, the target object's pose is determined based on the target pixel information and target depth information.

[0116] In one possible implementation, determining the object posture and position information of the target object according to the target pixel information and the target depth information includes:

[0117] generating a second target three-dimensional feature descriptor according to the target pixel information and the target depth information;

[0118] Inputting the second target three-dimensional feature descriptor into a trained posture recognition model corresponding to the target object category to obtain the centroid position of the target object output by the posture recognition model and a second target relative position relationship between the second target three-dimensional feature descriptor and the centroid position of the target object, wherein the posture recognition model is trained for each object type by using the second three-dimensional feature descriptor as input to a neural network model and using the centroid position of the object and the relative position relationship between the second three-dimensional feature descriptor and the centroid position of the object as outputs of the neural network model, wherein the second three-dimensional feature descriptor is generated based on pixel information and depth information;

[0119] Determining the position information of the target object according to the centroid position of the target object;

[0120] The object posture of the target object is determined according to the relative position relationship of the second target.

[0121] In this embodiment, a corresponding posture recognition model is trained for each object type. After determining the target object category, the trained posture recognition model corresponding to the target object category is obtained to determine the target object's posture and position information. The second target 3D feature descriptor is generated based on target pixel information and target depth information. Specifically, the second target 3D feature descriptor may include target pixel information and target depth information.

[0122] Specifically, this embodiment obtains the centroid position of the target object and the second target relative position relationship between the second target three-dimensional feature descriptor and the centroid position of the target object through a posture recognition model corresponding to the target object category, thereby determining the position information of the target object based on the centroid position, and determining the object posture based on the second target relative position relationship.

[0123] The above embodiments all use target depth information and target pixel information to identify object information. The following embodiments describe how to obtain target depth information and target pixel information based on any of the above embodiments.

[0124] In some cases of the examples, segmentation processing can be performed using two-dimensional images or depth information to obtain target pixel information and target depth information. However, if segmentation processing is performed only using two-dimensional images or depth information, the segmentation may be inaccurate, resulting in inaccurate target depth information and target pixel information, which in turn leads to inaccurate object information recognition.

[0125] refer to Figure 5 , Figure 5 A detailed flow chart of segmentation processing based on the two-dimensional image and the depth information is provided in an embodiment of the present application. In one embodiment, Figure 5 As shown, segmentation processing is performed according to the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image, including:

[0126] Step 510: Determine a segmentation position between the foreground and background of the two-dimensional image according to the depth information.

[0127] The foreground refers to a person or object close to the second image acquisition device. The background refers to a person or object far from the second image acquisition device. In this embodiment, the segmentation position between the foreground and the background can be considered as the location of the target object.

[0128] Step 520: Obtain image position information of the target object in the two-dimensional image.

[0129] In this embodiment, pixel information of each pixel in a two-dimensional image can be obtained, and a connected region with matching pixel information can be used as the image of the target object. The position of this connected region can then be used as the image position information of the target object. For example, if a table is generally of the same color, a connected region with similar RGB values ​​can be used as the image of the table to determine the image position information of the table.

[0130] Step 530: Perform segmentation processing according to the segmentation position and the image position information to obtain the target depth information and the target pixel information.

[0131] In this embodiment, the position of the target object in the preset area can be accurately determined by combining the segmentation position and the image position information, thereby accurately obtaining the target depth information and the target pixel information.

[0132] The technical solution of this embodiment performs segmentation processing based on depth information and two-dimensional images, which can improve the accuracy of the obtained target depth information and target pixel information.

[0133] refer to Figure 6 , Figure 6 This is a flow chart of an object information recognition method provided in an embodiment of the present application. Figure 1 Terminal or Figure 2 In one embodiment, the robot in FIG. Figure 6 As shown, the path planning method includes:

[0134] Step 610: Acquire a two-dimensional image and depth information obtained by capturing an image of a preset area.

[0135] Step 620 : Perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image.

[0136] Step 630: Perform information recognition on the target object according to the target pixel information and the target depth information to obtain object information of the target object, where the object information is used to indicate a space occupied by the target object in the preset area.

[0137] Step 640: Plan the robot's travel path based on the object information.

[0138] In this embodiment, the robot's travel path is planned according to the object information, so that the robot avoids the target object during movement, thereby achieving obstacle avoidance.

[0139] The technical solution of this embodiment can obtain the object information of the target object located in the preset area through the two-dimensional image and depth information of the preset area, and can obtain the object information of the object without measuring the depth information of a circle of the object, thereby improving the efficiency of object information recognition. Therefore, the efficiency of corresponding path planning is also improved, and the cleaning efficiency is also improved.

[0140] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0141] Based on the same inventive concept, embodiments of the present application also provide an object information recognition device for implementing the object information recognition method described above. The solution provided by this device is similar to the solution described in the method described above. Therefore, the specific limitations of one or more of the following embodiments of the object information recognition device can be found in the above-mentioned limitations of the object information recognition method and will not be further elaborated here.

[0142] refer to Figure 7 , Figure 7 This is a schematic diagram of the structure of an object information recognition device provided in an embodiment of the present application. In one embodiment, Figure 7 As shown, the object information recognition device includes an acquisition module 710, a segmentation processing module 720 and an information recognition module 730, wherein:

[0143] An acquisition module 710 is configured to acquire a two-dimensional image and depth information obtained by capturing an image of a preset area;

[0144] a segmentation processing module 720 configured to perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image;

[0145] The information recognition module 730 is configured to perform information recognition on the target object based on the target pixel information and the target depth information to obtain object information of the target object, where the object information indicates a space occupied by the target object in the preset area.

[0146] In one embodiment, the object information includes volume size, object posture, and location information of the target object in the preset area. The information recognition module 730 includes:

[0147] an acquiring unit, configured to acquire a target object category of the target object;

[0148] An information recognition unit is used to perform information recognition on the target object according to the target object category, the target pixel information and the target depth information to obtain the volume size, object posture and position information of the target object.

[0149] In one embodiment, the acquisition unit is specifically configured to input the target pixel information into a trained object category recognition model to obtain the object category and confidence level output by the object category recognition model, wherein the confidence level is used to indicate the credibility of the object category output by the object category recognition model;

[0150] If the confidence level is higher than a preset confidence level, the object category output by the object category recognition model is used as the target object category of the target object.

[0151] In one embodiment, the information identification unit includes:

[0152] The first information recognition subunit is used to generate a first target three-dimensional feature descriptor based on the target object category, the target pixel information and the target depth information; input the first target three-dimensional feature descriptor into the trained object feature extraction model to obtain the centroid position of the target object output by the object feature extraction model and the first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position of the target object, wherein the object feature extraction model is obtained by training the first three-dimensional feature descriptor as the input of a neural network model and the centroid position of the object and the relative position relationship between the first three-dimensional feature descriptor and the centroid position of the object as the output of the neural network model, wherein the first three-dimensional feature descriptor is generated based on the object category, pixel information and depth information; determine the volume size of the target object based on the first target three-dimensional feature descriptor; determine the position information of the target object based on the centroid position of the target object; and determine the object posture of the target object based on the first target relative position relationship.

[0153] In one embodiment, the first information recognition subunit is specifically configured to determine the volume size of the target object according to a matching result of the first target three-dimensional feature descriptor in a global coordinate system, where the global coordinate system is a coordinate system of the preset area.

[0154] In one embodiment, the information identification unit includes:

[0155] a second information recognition subunit, configured to search a database for a target volume size corresponding to the target object category according to the target object category, and use the target volume size as the volume size of the target object, wherein the database includes volume sizes corresponding to different object categories;

[0156] The object posture and position information of the target object are determined according to the target pixel information and the target depth information.

[0157] In one embodiment, the second information recognition subunit is specifically used to generate a second target three-dimensional feature descriptor based on the target pixel information and the target depth information; input the second target three-dimensional feature descriptor into the trained posture recognition model corresponding to the target object category to obtain the centroid position of the target object output by the posture recognition model and the second target relative position relationship between the second target three-dimensional feature descriptor and the centroid position of the target object. The posture recognition model is obtained by training the second three-dimensional feature descriptor as the input of the neural network model and the centroid position of the object and the relative position relationship between the second three-dimensional feature descriptor and the centroid position of the object as the output of the neural network model for each object type. The second three-dimensional feature descriptor is generated based on pixel information and depth information; the position information of the target object is determined based on the centroid position of the target object; and the object posture of the target object is determined based on the second target relative position relationship.

[0158] In one embodiment, the segmentation processing module 720 is specifically configured to determine a segmentation position between the foreground and the background of the two-dimensional image according to the depth information;

[0159] Acquiring image position information of the target object in the two-dimensional image;

[0160] Segmentation processing is performed according to the segmentation position and the image position information to obtain the target depth information and the target pixel information.

[0161] Each module in the object information recognition device described above may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of the processor in the terminal in hardware form, or may be stored in the terminal's memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0162] Based on the same inventive concept, the present application also provides a path planning device for implementing the path planning method described above. The solution provided by this device is similar to the solution described in the method described above. Therefore, the specific limitations of one or more path planning device embodiments provided below can be found in the above-mentioned limitations of the path planning method and will not be repeated here.

[0163] refer to Figure 8 , Figure 8 A schematic diagram of the structure of a path planning device provided in an embodiment of the present application. In one embodiment, Figure 8 As shown, the object information recognition device includes an acquisition module 810, a segmentation processing module 820, an information recognition module 830 and a path planning module 840, wherein:

[0164] An acquisition module 810 is configured to acquire a two-dimensional image and depth information obtained by capturing an image of a preset area;

[0165] a segmentation processing module 820 configured to perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image;

[0166] an information recognition module 830 for performing information recognition on the target object based on the target pixel information and the target depth information to obtain object information of the target object, where the object information is used to indicate a space occupied by the target object in the preset area;

[0167] The path planning module 840 is used to plan the robot's travel path according to the object information.

[0168] Each module in the above-mentioned path planning device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the terminal in hardware form, or can be stored in the memory in the terminal in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0169] refer to Figure 9 , Figure 9 This is a diagram of the internal structure of a terminal provided in an embodiment of the present application. In one embodiment, Figure 9 As shown, the terminal includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the terminal is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an object information recognition method and a path planning method are implemented. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen, and the input device of the terminal can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the terminal housing, or an external keyboard, touchpad or mouse.

[0170] Those skilled in the art will understand that Figure 9The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0171] In one embodiment, a terminal is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0172] In one embodiment, a robot is provided, comprising a first image acquisition device, a second image acquisition device, and a processor, wherein the processor is configured to implement the steps in the above-mentioned method embodiments.

[0173] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0174] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0175] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0176] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0177] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for identifying object information, characterized in that: include: Acquire a two-dimensional image and depth information obtained by image acquisition of a preset area; performing segmentation processing according to the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image; Obtaining a target object category of the target object; performing information identification on the target object according to the target object category, the target pixel information, and the target depth information to obtain volume size, object posture, and position information of the target object, wherein the volume size, object posture, and position information of the target object are collectively used to indicate a space occupied by the target object in the preset area; The step of identifying the target object based on the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes: generating a first target three-dimensional feature descriptor according to the target object category, the target pixel information, and the target depth information, wherein the first target three-dimensional feature descriptor is used to determine a centroid position of the target object and a first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position, and the first target three-dimensional feature descriptor is used to characterize shape information of a local surface of the target object, as well as shape information and color information of a local curved surface of the target object; The volume size of the target object is determined according to the first target three-dimensional feature descriptor; the position information of the target object is determined according to the centroid position; and the object posture of the target object is determined according to the relative position relationship of the first target.

2. The method according to claim 1, characterized in that The acquiring the target object category of the target object includes: Inputting the target pixel information into a trained object category recognition model to obtain the object category and confidence level output by the object category recognition model, wherein the confidence level is used to indicate the credibility of the object category output by the object category recognition model; If the confidence level is higher than a preset confidence level, the object category output by the object category recognition model is used as the target object category of the target object.

3. The method according to claim 1, characterized in that The performing information recognition on the target object according to the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes: generating a first target three-dimensional feature descriptor according to the target object category, the target pixel information, and the target depth information; Inputting the first target three-dimensional feature descriptor into a trained object feature extraction model to obtain the centroid position of the target object and a first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position of the target object output by the object feature extraction model, wherein the object feature extraction model is obtained by training the first three-dimensional feature descriptor as input to a neural network model and using the centroid position of the object and the relative position relationship between the first three-dimensional feature descriptor and the centroid position of the object as outputs of the neural network model, wherein the first three-dimensional feature descriptor is generated based on object category, pixel information, and depth information; determining a volume size of the target object according to the first target three-dimensional feature descriptor; Determining the position information of the target object according to the centroid position of the target object; The object posture of the target object is determined according to the relative position relationship of the first target.

4. The method according to claim 3, characterized in that The determining the volume size of the target object according to the first target three-dimensional feature descriptor includes: The volume size of the target object is determined according to a matching result of the first target three-dimensional feature descriptor in a global coordinate system, where the global coordinate system is a coordinate system of the preset area.

5. The method according to claim 1, characterized in that The performing information recognition on the target object according to the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes: According to the target object category, searching a database for a target volume size corresponding to the target object category, and using the target volume size as the volume size of the target object, wherein the database includes volume sizes corresponding to different object categories; The object posture and position information of the target object are determined according to the target pixel information and the target depth information.

6. The method according to claim 5, characterized in that The determining the object posture and position information of the target object according to the target pixel information and the target depth information includes: generating a second target three-dimensional feature descriptor according to the target pixel information and the target depth information; Inputting the second target three-dimensional feature descriptor into a trained posture recognition model corresponding to the target object category to obtain the centroid position of the target object output by the posture recognition model and a second target relative position relationship between the second target three-dimensional feature descriptor and the centroid position of the target object, wherein the posture recognition model is trained for each object type by using the second three-dimensional feature descriptor as input to a neural network model and using the centroid position of the object and the relative position relationship between the second three-dimensional feature descriptor and the centroid position of the object as outputs of the neural network model, wherein the second three-dimensional feature descriptor is generated based on pixel information and depth information; Determining the position information of the target object according to the centroid position of the target object; The object posture of the target object is determined according to the relative position relationship of the second target.

7. The method according to any one of claims 1 to 6, characterized in that The performing segmentation processing according to the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image includes: determining a segmentation position between a foreground and a background of the two-dimensional image according to the depth information; Acquiring image position information of the target object in the two-dimensional image; Segmentation processing is performed according to the segmentation position and the image position information to obtain the target depth information and the target pixel information.

8. A path planning method, characterized in that: include: Acquire a two-dimensional image and depth information obtained by image acquisition of a preset area; performing segmentation processing according to the two-dimensional image and the depth information to obtain target depth information corresponding to the target object located in the preset area and target pixel information of the target object in the two-dimensional image; Obtaining a target object category of the target object; performing information identification on the target object according to the target object category, the target pixel information, and the target depth information to obtain volume size, object posture, and position information of the target object, wherein the volume size, object posture, and position information of the target object are collectively used to indicate a space occupied by the target object in the preset area; Planning the robot's travel path based on the volume, size, posture and position information of the target object; The step of identifying the target object based on the target object category, the target pixel information, and the target depth information to obtain the volume size, object posture, and position information of the target object includes: generating a first target three-dimensional feature descriptor according to the target object category, the target pixel information, and the target depth information, wherein the first target three-dimensional feature descriptor is used to determine a centroid position of the target object and a first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position, and the first target three-dimensional feature descriptor is used to characterize shape information of a local surface of the target object, as well as shape information and color information of a local curved surface of the target object; The volume size of the target object is determined according to the first target three-dimensional feature descriptor; the position information of the target object is determined according to the centroid position; and the object posture of the target object is determined according to the relative position relationship of the first target.

9. An object information recognition device, characterized in that: include: An acquisition module is used to acquire a two-dimensional image and depth information obtained by collecting images of a preset area; a segmentation processing module, configured to perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to a target object located in the preset area and target pixel information of the target object in the two-dimensional image; An information recognition module, configured to obtain a target object category of the target object; performing information identification on the target object according to the target object category, the target pixel information, and the target depth information to obtain volume size, object posture, and position information of the target object, wherein the volume size, object posture, and position information of the target object are collectively used to indicate a space occupied by the target object in the preset area; Among them, the information recognition module is also used to generate a first target three-dimensional feature descriptor based on the target object category, the target pixel information and the target depth information, wherein the first target three-dimensional feature descriptor is used to determine the centroid position of the target object, and determine the first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position, and the first target three-dimensional feature descriptor is used to characterize the shape information of the local surface of the target object, as well as the shape information and color information of the local curved surface of the target object; determine the volume size of the target object according to the first target three-dimensional feature descriptor; determine the position information of the target object according to the centroid position; and determine the object posture of the target object according to the first target relative position relationship.

10. A path planning device, characterized in that: include: An acquisition module is used to acquire a two-dimensional image and depth information obtained by collecting images of a preset area; a segmentation processing module, configured to perform segmentation processing based on the two-dimensional image and the depth information to obtain target depth information corresponding to a target object located in the preset area and target pixel information of the target object in the two-dimensional image; An information recognition module, configured to obtain a target object category of the target object; performing information identification on the target object according to the target object category, the target pixel information, and the target depth information to obtain volume size, object posture, and position information of the target object, wherein the volume size, object posture, and position information of the target object are collectively used to indicate a space occupied by the target object in the preset area; A path planning module is used to plan the robot's travel path based on the volume, size, posture and position information of the target object; Among them, the information recognition module is also used to generate a first target three-dimensional feature descriptor based on the target object category, the target pixel information and the target depth information, wherein the first target three-dimensional feature descriptor is used to determine the centroid position of the target object, and determine the first target relative position relationship between the first target three-dimensional feature descriptor and the centroid position, and the first target three-dimensional feature descriptor is used to characterize the shape information of the local surface of the target object, as well as the shape information and color information of the local curved surface of the target object; determine the volume size of the target object according to the first target three-dimensional feature descriptor; determine the position information of the target object according to the centroid position; and determine the object posture of the target object according to the first target relative position relationship.

11. A terminal comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A robot, characterized in that: include: A first image acquisition device is used to acquire an image of a preset area to obtain a two-dimensional image; The second image acquisition device is used to acquire images of a preset area to obtain depth information; A processor for implementing the steps of the method according to any one of claims 1 to 8.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Industrial part intelligent identification and sorting system based on computer vision

    CN111421539A