Object pose estimation, image processing method, and related devices
By employing a multilayer perceptron network for object instance mask segmentation in 3D space, the problem of long object pose estimation time in existing technologies is solved, achieving more efficient object pose estimation and improving the performance of augmented reality applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SAMSUNG TELECOM R&D CENT
- Filing Date
- 2020-12-08
- Publication Date
- 2026-05-12
AI Technical Summary
Existing object pose estimation methods are time-consuming and cannot meet the needs of scenarios with high real-time requirements, especially in cluster-based object instance segmentation and key point detection.
We employ an object instance segmentation and keypoint regression method based on 3D spatial location information. By using a multilayer perceptron network to perform instance mask segmentation based on point cloud information at the object center, we reduce clustering steps and improve computational efficiency.
It effectively reduces the computation time for object pose estimation and improves the system efficiency, accuracy, and robustness in augmented reality applications.
Smart Images

Figure CN114663502B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to object pose estimation, image processing methods, and related equipment. Background Technology
[0002] With the development of artificial intelligence technology, augmented reality, computer vision, map navigation, and other technologies are becoming increasingly important in people's lives and work. Among them, pose estimation, which can estimate the pose of objects in images captured by a camera, has applications in many artificial intelligence technologies.
[0003] Pose estimation involves instance segmentation and keypoint detection tasks. In existing technologies, clustering is generally used for these tasks. However, clustering is very time-consuming and cannot meet the needs of some scenarios with very high real-time requirements. Summary of the Invention
[0004] The purpose of this application is to provide an object pose estimation, image processing method, and related equipment to reduce the time required for image processing. The specific solutions provided in the embodiments of this application are as follows:
[0005] In a first aspect, this application provides a method for estimating the pose of an object, including:
[0006] Obtain the image features corresponding to the point cloud of the input image;
[0007] Based on the image features, semantic segmentation information, instance mask information, and key point information of the object are determined;
[0008] Object pose estimation is performed based on the semantic segmentation information, instance mask information, and key point information.
[0009] Secondly, this application provides an image processing method, including:
[0010] Obtain the image features corresponding to the point cloud of the input image;
[0011] Based on the image features, the instance mask information of the image is determined by a multilayer perceptron network for instance mask segmentation, using the point cloud corresponding to the object center as a reference.
[0012] Image processing is performed based on the instance mask information.
[0013] Thirdly, this application provides an object pose estimation device, comprising:
[0014] The first acquisition module is used to acquire the image features corresponding to the point cloud of the input image;
[0015] The first determining module is used to determine the semantic segmentation information, instance mask information, and key point information of the object based on the image features.
[0016] The pose estimation module is used to estimate the pose of an object based on the semantic segmentation information, instance mask information, and key point information.
[0017] Fourthly, this application provides an image processing apparatus, comprising:
[0018] The second acquisition module is used to acquire the image features corresponding to the point cloud of the input image;
[0019] The second determining module is used to determine the instance mask information of the image based on the image features and using a multilayer perceptron network for instance mask segmentation with the point cloud corresponding to the object center as a reference.
[0020] The processing module is used to perform image processing based on the instance mask information.
[0021] Fifthly, this application provides an electronic device including a memory and a processor; the memory stores a computer program; the processor is used to execute the object pose estimation method or image processing method provided in the embodiments of this application when running the computer program.
[0022] Sixthly, this application provides a computer-readable storage medium storing a computer program, which, when run by a processor, executes the object pose estimation method or image processing method provided in the embodiments of this application.
[0023] The beneficial effects of the technical solutions provided in this application will be described in detail in the following description of the specific implementation methods in conjunction with various optional embodiments, and will not be elaborated further here. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0025] Figure 1 A flowchart of an object pose estimation method provided in one embodiment of this application;
[0026] Figure 2 This is a flowchart of object pose estimation in one example of this application;
[0027] Figure 3 This is a network structure diagram for object pose estimation in an example of this application;
[0028] Figure 4 This is a flowchart of object pose estimation in one example of this application;
[0029] Figure 5 This is a network structure diagram for object pose estimation in an example of this application;
[0030] Figure 6 This is a flowchart of object pose estimation in one example of this application;
[0031] Figure 7 This is a flowchart of object pose estimation in one example of this application;
[0032] Figure 8 A schematic diagram of a location-aware instance mask segmentation method according to this application is shown;
[0033] Figure 9 This illustration shows another instance mask segmentation based on location awareness according to this application;
[0034] Figure 10 This is a network structure diagram for object pose estimation in an example of this application;
[0035] Figure 11 This is a flowchart of object pose estimation in one example of this application;
[0036] Figure 12 A schematic diagram illustrating one application scenario of this application is shown;
[0037] Figure 13 A flowchart illustrating an image processing method provided in one embodiment of this application;
[0038] Figure 14 This is a schematic diagram of the structure of an object pose estimation device provided in one embodiment of the present application;
[0039] Figure 15 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this application;
[0040] Figure 16 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation
[0041] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting the invention.
[0042] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0043] To better understand and explain the solutions provided in the embodiments of this application, the relevant technologies involved in this application will be described first below.
[0044] This application belongs to the field of augmented reality technology, specifically involving multimodal (color and depth) image processing and recognition technology based on learning algorithms, object recognition, instance segmentation, and 6-DOF object pose estimation technology.
[0045] Augmented Reality (AR) is a technology that calculates the position and angle of camera images in real time and adds corresponding images, overlaying computer-generated virtual objects or information about real objects onto real-world scenes to enhance the real world. Augmented Reality technology provides users with a realistic information experience by adding virtual content to the real-world scene in front of them. In three-dimensional space, augmented reality systems need high-precision real-time processing and understanding of the three-dimensional state of surrounding objects to achieve a high-quality virtual-real fusion effect in front of the user.
[0046] Pose estimation utilizes geometric models or structures to represent the structure and shape of an object. By extracting the object's features, a correspondence is established between the geometric model and the image. Then, the spatial pose of the object is estimated using geometric or other methods. The geometric model used can be a simple geometric shape, such as a plane or cylinder, or a certain geometric structure, or a three-dimensional model obtained through laser scanning or other methods. Pose estimation techniques can include instance segmentation tasks and keypoint detection tasks. The instance segmentation task includes a semantic segmentation task for separating the target object from the background and determining the category of the target object, and an instance mask segmentation task for determining the pixels belonging to each target object. The keypoint detection task is used to determine the position of the target object in the image. The pose estimation method provided in this application can be applied to the field of augmented reality technology.
[0047] 6DoF pose estimation of images: Given an image containing color and depth information, estimate the 6DoF pose of the target object in the image. 6DoF pose is also known as 6-dimensional pose, which includes 3-dimensional position and 3-dimensional spatial orientation.
[0048] In existing technologies, a 6DoF pose estimation method based on RGBD (Red+Green+Blue+depth, an image containing the three primary colors of red, green, and blue, and depth information) has been proposed for object pose estimation. This method uses a deep Hough voting network to predict voting information based on the object's keypoints, and finally employs clustering to determine the object's 3D keypoints and uses least-squares fitting to estimate the object's pose. This method is an extension of the 2D keypoint method, which can fully utilize the depth information in the depth image. However, this method requires object instance segmentation and keypoint detection through clustering, which is a very time-consuming method.
[0049] In existing technologies, object instance segmentation methods generally employ detection branches to extract objects or rely on grouping or clustering to group similar instance points to identify objects. However, detection branch-based methods cannot guarantee consistent instance labels for each point, while grouping methods require parameter tuning and are computationally very expensive. Currently, the fastest time-efficient solutions in object instance segmentation rely on 2D images; however, 2D projection onto the image may contain artifacts, such as occlusion and overlap between image regions of multiple objects.
[0050] To address at least one of the aforementioned problems, this application provides an object pose estimation method that performs instance segmentation and keypoint regression of objects based on their positional information in 3D space, effectively avoiding clustering methods and reducing the time required for object pose estimation. This application also provides an image processing method that uses a multilayer perceptron network for object instance mask segmentation to determine the object's instance mask information based on the point cloud corresponding to the object's center, thus avoiding clustering methods and reducing the time required for image processing.
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following will describe in detail, with reference to specific embodiments and accompanying drawings, the various optional implementation methods of this application and how the technical solutions of the embodiments of this application solve the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings. Figure 1 The image shows an embodiment of an object pose estimation method provided in this application, which includes the following steps S101-S103:
[0052] Step S101: Obtain the image features corresponding to the point cloud of the input image.
[0053] Specifically, an object is an object with six degrees of freedom in three-dimensional space, also known as a 6-dimensional object; an object is an object in the real world, such as a mouse, a fan, a desk, etc.
[0054] The input object image can be an RGBD image (Red+Green+Blue+depth, an image containing the three primary colors of red, green, and blue and depth information) acquired by an image acquisition device, including a color image and a depth image. In some embodiments, the object image can also be a grayscale image. Optionally, the object image can only include a depth image, or it can include a color image, a grayscale image, and a depth image. The color image, grayscale image, and depth image correspond to images containing the same object in the same scene. The same image can include multiple different types of objects, and the type and number of objects are not limited here. The acquisition method of the color image, grayscale image, and corresponding depth image is not limited in this application embodiment. For example, it can be a color image, grayscale image, and depth image acquired simultaneously by a general image acquisition device and a depth image acquisition device, or it can be acquired by an image acquisition device or video acquisition device that has the functions of acquiring depth images, color images, and grayscale images simultaneously.
[0055] Specifically, the choice between using color or grayscale images can be made based on actual application requirements. For example, a color image can be used to obtain better pose estimation results, while a grayscale image can be used if higher efficiency is required for pose estimation. A grayscale image is an image where each pixel has only one sampled color, and grayscale images have many levels of color depth between black and white.
[0056] In practical applications, many scenarios require continuous real-time object pose estimation. Optionally, an RGBD video of the object can be acquired using a video capture device. Each frame of this video is an RGBD image, and the object's color and depth images can be extracted from the same video frame. For each video frame, a corresponding color image and depth image can be obtained. The object in the color image and depth image of the same video frame is consistent, and there can be one or more objects. Based on the scheme provided in the embodiments of this application, real-time object pose estimation can be achieved using the color and depth images corresponding to each acquired video frame.
[0057] Point clouds can be a massive collection of points representing the surface characteristics of an object. Point clouds can be obtained based on laser measurement principles and photogrammetry principles. Point clouds obtained based on laser measurement principles include three-dimensional coordinates (XYZ) and laser reflection intensity; point clouds obtained based on photogrammetry principles include three-dimensional coordinates (XYZ) and color information (RGB); point clouds obtained by combining laser measurement and photogrammetry principles include three-dimensional coordinates (XYZ), laser reflection intensity, and color information (RGB). After acquiring the spatial coordinate information of each sampling point on the object's surface, a collection of points is obtained, referred to in this application as a "point cloud".
[0058] Step S102: Based on image features, determine the semantic segmentation information, instance mask information, and key point information of the object.
[0059] Specifically, the semantic segmentation task refers to separating objects from the background, obtaining a semantic segmentation probability map through input image features, and determining the predicted category of each point cloud based on the semantic segmentation probability map.
[0060] Instance mask segmentation (object detection) refers to determining the position of each object in an image. By inputting image features, a probability map of different positions of instance segmentation can be obtained. Based on the probability map, the position of the object corresponding to each point cloud can be determined.
[0061] Optionally, since instance mask segmentation involves processing object positions, this application primarily uses point cloud information provided in image features in conjunction with other image feature information for instance mask segmentation. Point cloud information may include information representing the spatial coordinate values of the point cloud.
[0062] Specifically, key point information can be determined based on the instance mask information of the object, or it can be determined based on the semantic segmentation information and instance mask information of the object; optionally, the spatial coordinate values of the key points are obtained in step S102.
[0063] The key points of an object can be manually defined key points or key points calculated using the farthest point sampling (FPS) method.
[0064] Step S103: Perform object pose estimation based on semantic segmentation information, instance mask information, and key point information.
[0065] Specifically, based on semantic segmentation information, instance mask information, key point information, and the geometric model of the same object (which can be a computer-aided design, CAD model), the six-degree-of-freedom pose of the object can be calculated by least squares fitting, using the three-dimensional rotation and three-dimensional translation transformations between two sets of three-dimensional key points.
[0066] Optionally, the above steps S101-S103 can be performed by... Figure 2 The network shown is implemented as follows. Figure 2 As shown, the system includes a feature extraction module, a semantic segmentation module, an instance mask segmentation module, a keypoint detection module, and a pose estimation module. Specifically, the object's image is input into the feature extraction module to extract image features corresponding to the point cloud; these image features are then input into the semantic segmentation module and the instance mask segmentation module to determine the object's semantic segmentation information (such as the semantic category of the point cloud) and instance mask information; the image features are input into the keypoint detection module and combined with the instance mask information or (instance mask information and semantic segmentation information) to obtain keypoints (such as the spatial coordinate information of the keypoints); finally, the keypoints or keypoints and semantic segmentation information are input into the pose estimation module to determine the object and its corresponding pose information. The semantic segmentation module, instance mask segmentation module, and keypoint detection module can be based on a Multi-Layer Perceptron (MLP) network, forming a position-aware semantic segmentation network, an instance mask segmentation network, and a position-aware 3D keypoint regression network, respectively.
[0067] This application can extract image features (including point cloud information) corresponding to point clouds from object images. Since points and surfaces on an object do not overlap or occlude each other in 3D space, using 3D position information to perceive the position and integrity of the object is very important. Based on semantic segmentation information and instance mask segmentation information of the object perceived by 3D position information, the key point information of each object can be quickly calculated by regression without complex clustering, thus reducing computation time and improving the efficiency of key point detection. The six degrees of freedom pose of the object can be calculated using the predicted key point information and the key point information of the object's 3D model through least squares fitting. Furthermore, the object pose estimation method provided by this application also helps improve the efficiency, accuracy, and robustness of systems in augmented reality applications; the instance segmentation and 3D key point regression networks can be connected to form an end-to-end trainable network, thus avoiding tedious parameter adjustments in multiple post-processing (voting clustering) steps.
[0068] The optional embodiments provided in this application are described in detail below.
[0069] In one optional embodiment of this application, step S101, based on the image of the object, extracts image features including point cloud information, including at least one of the following steps A1-A2:
[0070] Step A1: Extract point cloud features based on the input depth image, and confirm the extracted point cloud features as the image features corresponding to the point cloud.
[0071] Step A2: Extract first image features based on the input color image and / or grayscale image; extract point cloud features based on the input depth image; fuse the first image features and the point cloud features to obtain the image features corresponding to the point cloud.
[0072] The solution provided in this embodiment can extract image features based on depth images, or it can extract image features based on color images and / or grayscale images and depth images. The extraction of image features can be performed through image feature extraction networks, image feature extraction algorithms, etc.
[0073] One optional embodiment of this application provides a solution in which feature extraction is performed based on a color image and / or grayscale image and a depth image. Image feature extraction is performed through an image feature extraction network. Specifically, the color image and / or grayscale image and the depth image are input into the image feature extraction network. When the input image is a color image, the input to the image feature extraction network is obtained by pixel-by-pixel concatenation of the color image and the depth image. The image consists of four channels, where H is the image height, W is the image width, and the four channels are the three channels corresponding to the RGB data of the color image and the channel corresponding to the depth data of the depth image. The output of the image feature extraction network includes the image feature vector of each pixel.
[0074] Optionally, to extract point cloud features from a depth image, the depth image can first be converted into a point cloud to obtain the corresponding point cloud data (e.g., a 2D image converted into a 3D image). Then, point cloud features can be extracted based on the point cloud data to obtain the point cloud features corresponding to the depth image. Specifically, point cloud feature extraction can be performed using a point cloud feature extraction network. The input to the point cloud feature extraction network is the point cloud data, and the output of the network includes the point cloud feature vector for each 3D point, thus obtaining the point cloud feature vector for each pixel.
[0075] In one embodiment, the following can be used: Figure 3 The network structure shown and Figure 4 The flowchart shown illustrates step A2. The feature extraction network layer includes an image convolutional neural network (CNN), a point cloud feature extraction network, and a fusion network. Figure 4The dashed lines in the middle show other feasible network structures, such as instance location probability maps, which can also be used as input data for 3D keypoint offsets. Specifically, step A2 extracts first image features based on the input color image and / or grayscale image, including: Step A21: Extracting first image features using a convolutional neural network based on the input color image and / or grayscale image.
[0076] Optionally, the image convolutional network can be a deep learning network based on convolutional neurons. The input data of this network can be a 3-channel color image I of size H multiplied by W, where H is the image height, W is the image width, and the 3 channels are the RGB color channels in the color image. The output data of this network can include an image feature vector map (first image feature) for each pixel, with a size of H multiplied by W. In one embodiment, the network structure can be a multi-layer convolutional neural network built according to the actual scenario requirements, or it can be the network part before the fully connected layers of a neural network such as AlexNet, VGG net (Visual Geometry Group net), or ResNet (deep residual network).
[0077] Specifically, step A2 involves extracting point cloud features based on the input depth image, including: Step A22: Extracting point cloud features based on the input depth image using a multilayer perceptron network.
[0078] Optionally, the point cloud feature extraction network can be a multilayer perceptron (MLP) network, such as PointNet++ (a segmentation algorithm network). The input data to this network can be a point cloud obtained by projecting a depth image from two dimensions to three dimensions, or features containing other characteristics of each point cloud, such as RGB color, normal vectors, etc., with an input size of [missing information]. N is the number of point clouds, and M is the length of the feature vector for each point cloud. The output data of this network can include the point cloud feature vector for each 3D point, with a size of [missing information]. .
[0079] In one embodiment, step A2, which fuses the first image features and the point cloud features to obtain the image features corresponding to the point cloud, includes: step A23: fusing the first image features and the point cloud features pixel by pixel to obtain the image features corresponding to the point cloud.
[0080] Optionally, based on the one-to-one correspondence between the depth image and the 3D point cloud projection, the image pixels corresponding to each 3D point can be determined. For each pixel in the image, the image pixel feature vector (first image feature) obtained by passing the pixel through an image convolutional network and the point cloud feature vector (point cloud feature) obtained by passing the pixel through a point cloud feature extraction network are densely fused pixel by pixel through a fusion network. As an example, this fusion can be achieved by concatenating the image pixel feature vector and the point cloud feature vector, and the size of the fused feature (the image feature corresponding to the point cloud) is [missing information]. .
[0081] In this embodiment, the image features, including point cloud information, obtained through dense fusion are fed into multiple parallel multilayer perceptron (MLP) network structures to extract more expressive features before further processing. Figure 3 The task shown (can be adopted as follows) Figure 4 The network structure shown includes: semantic segmentation estimation of objects, instance mask segmentation estimation, offset estimation of 3D keypoints of objects, and coordinate prediction of 3D keypoints of objects (keypoint regression). Each of these four tasks has its own corresponding loss function during training. Furthermore, the semantic segmentation task, instance mask segmentation task, and keypoint detection (including offset estimation and keypoint regression) are parallel tasks. The following describes an embodiment of performing these tasks based on the image features corresponding to the point cloud obtained in step S101.
[0082] In one embodiment, step S102 determines the semantic segmentation information, instance mask information, and key point information of the object based on image features, including the following steps B1-B3:
[0083] Step B1: Determine the semantic segmentation information corresponding to the point cloud based on image features.
[0084] Specifically, semantic segmentation can be performed based on a multilayer perceptron 1. The semantic segmentation task of objects is used to estimate the semantic category of each point cloud. As an example, assuming there are C semantic categories (preset object categories), or C objects, the goal of this task is to output a semantic segmentation probability map of size [value missing]. Each value represents the probability that the nth point cloud belongs to the cth class. During network training, a real ground truth (i.e., the real semantic information of each point cloud) is constructed using real semantic annotation information. Then, based on the network output and annotation information, the value of a preset loss function (such as Dice loss - a metric function used to evaluate the similarity between two samples, softmax loss, etc.) is determined to update the network parameters. Through continuous supervision with data containing real annotation information, the network learns the network parameters for semantic segmentation upon convergence. In testing, RGBD images can be fed into a convolutional neural network and a point cloud feature extraction network to obtain the corresponding image features, which are then input into a multilayer perceptron to obtain a probability map for semantic segmentation. Finally, the class corresponding to the highest probability of each point cloud in this probability map is taken as the predicted semantic class of that point cloud, serving as the semantic segmentation information for that point cloud.
[0085] Step B2: Create a 3D mesh based on the point cloud information corresponding to the input image, and determine the instance mask information of the object based on the 3D mesh.
[0086] The creation of a 3D mesh can be understood as dividing the 3D space containing an object into very small 3D spatial cells, such as... Figure 8 As shown. If the center of an object falls within a cell of a 3D mesh, the relevant information obtained from that cell is used for object instance segmentation and 3D keypoint detection.
[0087] Specifically, a multilayer perceptron 2 can be used for instance mask segmentation, which involves predicting the position of each instance by sensing 3D location information. As an example, such as... Figure 3 As shown, firstly, based on all input point clouds (size: N represents the number of point clouds, and 3 represents the coordinates (x-axis, y-axis, z-axis) of each point cloud. A 3D space is created based on the position of the point cloud in the 3D spatial coordinates. A preset 3D meshing strategy is used to divide the corresponding 3D space to obtain a 3D mesh. For example, when using the 3D meshing strategy of equal spacing single branch division, during network training, the network node corresponding to the cell where the object center is located is set to 1, and the rest are set to 0. Then, the value of the loss function (such as Dice loss - a metric function used to evaluate the similarity between two samples, softmax loss, etc.) is determined based on the network output and annotation information to update the network parameters. Through continuous supervision with data containing real annotation information, the network parameters for instance segmentation are learned when the network converges. In testing, the image features obtained after inputting the RGBD image are input into the network to obtain the probability map of different positions of instance segmentation, i.e., the instance position probability map. Finally, the position corresponding to the highest probability of each point cloud in the probability map is determined as the object position of that point cloud as the instance mask information. Here, all point clouds at each cell position can be regarded as the same instance object.
[0088] In one embodiment, instance mask information represents the mesh information corresponding to the point cloud in a 3D mesh; wherein, the mesh information corresponding to each point cloud of the object is determined based on the mesh information corresponding to the point cloud at the center of the object.
[0089] Specifically, in a 3D mesh based on point cloud information, the mesh information corresponding to each point cloud of an object in the 3D mesh can be known (each point cloud of the object may be scattered across multiple cells of the 3D mesh). In this embodiment, the target cell containing the point cloud corresponding to the center of the object is first determined, and then each point cloud of the object is uniformly regarded as corresponding to the target cell. For example: In a 3D mesh (horizontal x, vertical y, depth z), the object corresponds to 3 point clouds (the actual number of point clouds may be large, this is only for illustrative purposes), point cloud 1, point cloud 2, and point cloud 3; point cloud 1 corresponds to cell (1,3,5), point cloud 2 corresponds to cell (6,7,8), and point cloud 3 corresponds to cell (2,4,7). Among them, the current center of the object corresponds to point cloud 2, so cell (6,7,8) is regarded as the cell corresponding to the 3 point clouds (the channel corresponding to this cell is used to predict the position of the object).
[0090] Step B3: Determine key point information based on instance mask information, or based on semantic segmentation information and instance mask information.
[0091] Specifically, after obtaining the corresponding semantic segmentation information (the predicted semantic category corresponding to each point cloud) and instance mask information (the object location of each point cloud) based on steps B1 and B2, key point information can be determined by combining image features with instance mask information, or by combining image features with semantic segmentation information and instance mask information. The specific process for determining key point information will be described in the following embodiments.
[0092] In one embodiment, such as Figure 9 As shown, the preset 3D mesh generation strategies include a single-branch method with equal spacing, a multi-branch method with varying cell sizes, and a multi-branch method with varying starting positions. Specifically, step B2, which creates a 3D mesh based on the point cloud information corresponding to the input image, may include at least one of the following steps B21-B23:
[0093] Step B21: Divide the point cloud information into three-dimensional space at equal intervals to obtain a three-dimensional mesh.
[0094] This can also be called a single-branch method. Specifically, the three-dimensional space corresponding to the point cloud information can be understood as the three-dimensional physical space where the object is located. When dividing the space, the three-dimensional space is divided into D equal parts at equal intervals in each direction, thus dividing the three-dimensional space into... There are [number] cells. At this point, the number of nodes in this layer of the multilayer perceptron 2 network is [number]. Where N is the number of point clouds. If the center of an object falls at cell (i,j,k), then the object corresponds to the first point in that layer of the neural network. Nodes In this partitioning method, to prevent any two objects from falling into the same cell, the size of the partitioned cells must be smaller than the size of any object. If the objects are very small, the value of D will be very large, thus increasing the network parameters. Therefore, to solve this problem, this application also proposes two methods implemented in steps B22 and B23.
[0095] Step B22: Divide the point cloud information into three-dimensional spaces based on multiple preset intervals to obtain multiple three-dimensional meshes.
[0096] This can also be called a multi-branch method with varying cell size. Specifically, in step 22, different intervals are set for each branch, thus obtaining different numbers of parts D. As an example, such as... Figure 9 As shown, two branches (or more) are set up, corresponding to two parallel position sensing layers. The spacing between each layer is different, resulting in different values of D. The number of network nodes in these two layers are respectively... , Where N is the number of point clouds. Combined with... Figure 9As can be seen, the first three-dimensional mesh formed includes sections equally spaced based on a preset spacing of 1. The resulting second 3D grid consists of cells; it includes cells equally spaced based on a preset spacing of 2. One cell.
[0097] Step B23: Based on the same spacing but different dividing starting points, divide the point cloud information to obtain multiple three-dimensional meshes in the corresponding three-dimensional space.
[0098] This can also be called a multi-branch method with varying starting positions. Specifically, in step B23, different starting points are set for each branch when dividing the intervals. As an example, such as... Figure 9 As shown, two branches (or more) are set up, corresponding to two parallel position sensing layers. The starting point of each layer is different, but the spacing is the same. The number of network nodes in both layers is [missing information]. Combining Figure 9 As can be seen, the first and second three-dimensional grids are both divided into multiple cells based on a preset spacing. A comparison shows that the cells formed by the two three-dimensional grids are different.
[0099] In one embodiment, before keypoint regression, a keypoint offset estimation task is included. Based on the estimated offsets and point cloud information, predicted values for keypoint regression can be obtained. The keypoint offset estimation task estimates the offset of each point cloud point relative to each 3D keypoint in the object. The 3D keypoints in the object can be defined based on empirical values or calculated using a correlation sampling algorithm. In the following embodiment, an example is given where each object in the image has K keypoints.
[0100] Specifically, step B3, which determines key point information based on instance mask information, includes at least one of the following steps C1-C2:
[0101] Step C1: Estimate the first offset of each key point corresponding to the point cloud based on image features; determine the key point information of the object by regression based on the first offset and instance mask information.
[0102] Specifically, such as Figure 3 and 4 As shown in this embodiment, when each object in the image corresponds to K key points, the network layer size used to perform the 3D key point offset estimation task is: Where N represents the number of point clouds, K represents the number of key points, and 3 represents the three coordinates in three-dimensional space. In the network (corresponding to...) Figure 4In the training process of the multilayer perceptron 3), the difference between the preset spatial coordinate information of the key points and the spatial coordinate value of each point cloud is used as the true label of the offset. Then, the loss function (which can be the Euclidean distance loss function) is determined based on the predicted label of the offset obtained by the network performing the task and the true label, so as to update the network parameters based on the loss function.
[0103] Optionally, in step C1, the key point information of the object is determined by regression based on the first offset and instance mask information, including: determining the initial predicted value of the key point based on the point cloud information; determining the target predicted value corresponding to the key point in the 3D mesh based on the initial predicted value and instance mask information; and determining the key point information of the object by regression based on the target predicted value.
[0104] Specifically, by adding the first offset predicted for each point cloud for each keypoint (in this embodiment, keypoints are all three-dimensional keypoints) to the original coordinate information of each point cloud (which can be seen from the point cloud information), the coordinate information (initial predicted value) of each keypoint predicted based on each point cloud can be obtained. That is, first offset + point cloud information = initial predicted value.
[0105] Specifically, since step C1 involves offset estimation across the entire 3D space, instance mask information needs to be incorporated into the keypoint regression step to calculate the spatial coordinates of the object's keypoints in each cell. Combined with... Figure 3 and Figure 4 As shown in the flowchart, the instance mask segmentation result and the offset estimation result are used together as input data for the regressor. In the network structure diagram, the regressor is connected to multilayer perceptron 3, as well as multilayer perceptrons 1 and 2. Optionally, based on the initial predicted value and the instance mask information, the target predicted value of each key point in the point cloud prediction of each cell can be determined.
[0106] Step C2: Based on image features and instance mask information, estimate the second offset of the key points corresponding to each point cloud in each cell of the 3D mesh; based on the second offset, determine the key point information of the object by regression.
[0107] Specifically, such as Figure 3 and Figure 4 As shown in this embodiment, when each object in the image corresponds to K key points, the network layer size used to perform the 3D key point offset estimation task is: (This corresponds to calculating the offset of each keypoint for each point cloud in each cell), where N represents the number of point clouds, K represents the number of keypoints, and 3 represents the three coordinates in 3D space. In the network (corresponding to...) Figure 4In the training process of the multilayer perceptron 3), the difference between the preset spatial coordinate information of the key points and the spatial coordinate value of each point cloud is used as the true label of the offset. Then, the loss function (which can be the Euclidean distance loss function) is determined based on the predicted label of the offset obtained by the network performing the task and the true label, so as to update the network parameters based on the loss function.
[0108] Unlike step C1, step C2 estimates the offset based on each point cloud contained in each cell, while step C1 estimates the offset based on the entire point cloud in the three-dimensional space where the object is located.
[0109] Optionally, in step C2, the key point information of the object is determined by regression based on the second offset, including: determining the target predicted value of the key points predicted based on the point cloud based on the second offset and the point cloud information; and determining the key point information (such as spatial coordinate values) of the object by regression based on the target predicted value.
[0110] Specifically, for each cell, the coordinate information (target predicted value) of each key point predicted based on each point cloud is obtained by adding the second offset predicted for each key point (in this embodiment, key points are all three-dimensional key points) to the original coordinate information of each point cloud (which can be seen from the point cloud information). That is, second offset + point cloud information = target predicted value.
[0111] Comparing steps C1 and C2, it can be seen that the information referred to by the first offset and the second offset is the predicted offset of each point cloud relative to each key point. The difference is that step C1 processes all point clouds in the 3D space, while step C2 processes the point cloud in each cell of the 3D space. That is, position-aware information (instance mask information) has been added when estimating the offset in step C2. Therefore, when regressing key points, the target prediction value determined by the second offset and the point cloud information can be directly processed.
[0112] Specifically, since step C2 involves offset estimation for each cell in 3D space, there is no need to include instance mask information in the keypoint regression step. Figure 4 The content shown by the dashed line is represented in the flowchart as the result of the 3D keypoint offset estimation as the input data of the regressor, and in the network structure diagram as the regressor being connected to the multilayer perceptron 3.
[0113] In one embodiment, such as Figure 5 , 6As shown in Figure 7, step B3, which determines key point information based on semantic segmentation information and instance mask information, includes the following steps C3: determining instance segmentation information based on semantic segmentation information and instance mask information; estimating the first offset of each key point corresponding to the point cloud based on image features; and determining the key point information of the object by regression based on the first offset and instance segmentation information.
[0114] Specifically, semantic segmentation information represents the category corresponding to each point cloud, and instance mask information can represent the cell (the cell corresponding to the object) of each point cloud in the 3D mesh. The instance segmentation information determined based on semantic segmentation information and instance mask information can be used to represent the object's category and position information. The semantic segmentation information can be obtained by using... Figure 7 The semantic segmentation computation shown is obtained by processing the fused image features using a multilayer perceptron; instance mask information can be obtained using... Figure 7 The position-aware 3D instance segmentation multilayer perceptron shown is obtained by processing the fused image features.
[0115] Optionally, the process of determining instance segmentation information based on semantic segmentation information and instance mask information can be understood as a process of eliminating redundant information, by using semantic segmentation information to eliminate cluttered information in the instance mask information. Specifically, the result corresponding to the semantic segmentation information can be of size [missing information]. The matrix, the result corresponding to the instance mask information can be of size The process of combining the two matrices can be understood as a matrix multiplication process, and the result is instance segmentation information.
[0116] Optionally, the specific process for determining the first offset can refer to the content shown in step C1 above. The first offset can be determined using... Figure 7 The multilayer perceptron shown is used to process the fused image features to calculate the 3D keypoint offset.
[0117] In one embodiment, step C3, which determines the key point information of the object based on the first offset and instance segmentation information through regression, may include the following steps: determining the initial predicted value of the key points predicted based on the point cloud information; determining the target predicted value of the key points in the 3D mesh based on the initial predicted value and instance segmentation information; and determining the key point information of the object through regression based on the target predicted value.
[0118] Specifically, determining the key point information of an object through regression can be achieved by using... Figure 7 The position-aware 3D keypoint regression multilayer perceptron shown is obtained by processing the first offset and instance segmentation information.
[0119] Specifically, the process for determining the initial and target predicted values can be referred to in step C1 above. The difference between step C3 and step C1 is that step C1 determines the target predicted value based on the initial predicted value and instance mask information; step C3 determines the target predicted value based on the initial predicted value and instance segmentation information. Relatively speaking, since the instance segmentation information has effectively eliminated some redundant information, the target predicted value determined in step C3 has higher accuracy.
[0120] In one embodiment, steps C1-C2, which determine the key point information of the object based on the target prediction value through regression, include at least one of the following steps C01-C04:
[0121] Step C01: For each key point of the object, the average of the target prediction value corresponding to the key point and each point cloud is determined as the key point information.
[0122] For each location (which can be understood as a cell), the target predicted value represents the coordinate information (predicted coordinate value) of each keypoint predicted based on each point cloud. Referring to Table 1 below, an example is provided (assuming there are currently 3 point clouds and 2 keypoints):
[0123] Table 1 (Target Predicted Values)
[0124]
[0125] As shown in Table 1 above, each point cloud corresponds to a specific key point with a corresponding predicted coordinate value (target predicted value).
[0126] In one embodiment, step C01 can be represented by the following formula (1):
[0127] y= ......(1)
[0129] Where i = 1, ..., N; i represents the i-th point cloud, and there are a total of N point clouds; x is the target prediction value; y is the key point information (such as the spatial coordinates of the key points).
[0130] Among them, the spatial coordinate information of key point 1 is the coordinate value obtained by calculating [(a1, b1, c1) + (a2, b2, c2) + (a3, b3, c3)] / 3.
[0131] Step C02: For each key point of the object, the weighted average of the target prediction value corresponding to the key point and each point cloud, and the probability value corresponding to each point cloud in the instance mask information, is determined as the key point information.
[0132] Optionally, step C02 can be understood as a weighted regression method that uses the confidence M of the location-aware instance mask segmentation as the weight (if the confidence contribution of the mask prediction is small, then the contribution to the key point prediction is small).
[0133] In one embodiment, step C02 can be represented by the following formula (2):
[0134] y= / ......(2)
[0136] Where i = 1, ..., N; i represents the i-th point cloud, and there are a total of N point clouds; x is the target prediction value; w is the mask confidence of each cell (the probability value of each cell corresponding to each point cloud), w = ;idx=1,...,D 3 W_ins is the network output probability map for instance mask segmentation, and idx is the label of the cell containing the object. Optionally, w can also be expressed as w = M_idx.
[0137] Specifically, the target prediction value can be represented by the example shown in Table 1 in step C01 above. The probability value corresponding to each point cloud in the instance mask information is the probability value of each point cloud corresponding to the current position (cell) (there are multiple cells in the 3D mesh, and the probability value of each point cloud corresponding to each cell can be predicted when segmenting the instance mask).
[0138] Among them, the spatial coordinate information of key point 1 is as follows:
[0139] The obtained coordinate values.
[0140] Step C03: For each key point of the object, the key point information of the object is determined by taking the weighted average of the target prediction value corresponding to the preset number of point clouds that are closest to the center point of the object and the probability value corresponding to the preset number of point clouds in the instance mask information.
[0141] In one embodiment, step C03 can be represented by the following formula (3):
[0142] y= / ......(3)
[0144] Where i = 1, ..., T; i represents the i-th point cloud, and there are a total of T point clouds. The T point clouds are determined by predicting the offset of all point clouds relative to the object center (keypoint), i.e. Where dx, dy, and dz are the predicted offset values in the x, y, and z directions (or the target predicted values), respectively. These are sorted in ascending order, and the top T point clouds are used as the point clouds for calculating keypoint information in step C03. x is the target predicted value; w is the mask confidence score for each cell (the probability value for each cell corresponding to each point cloud), w = ;idx=1,...,D 3 W_ins is the network output probability map for instance mask segmentation, and idx is the label of the cell containing the object. Optionally, w can also be expressed as w = M_idx.
[0145] Optionally, the representation of the target prediction value can refer to the example shown in Table 1 of step C01 above. Assume that the first T point clouds (T = 2) of the current three point clouds are point cloud 1 and point cloud 3. Then, the spatial coordinate information of key point 1 is: The obtained coordinate values.
[0146] Step C04: For each key point of the object, determine the key point information by taking the weighted average of the proximity value between the key point and each point cloud, the target prediction value between the key point and each point cloud, and the probability value between the key point and each point cloud in the instance mask information.
[0147] Specifically, the physical meaning represented by the proximity value between the key point and the point cloud is to quantify the distance between each point cloud and the key point of the object (the quantized value is between [0,1]), which can be expressed by the following formula (4):
[0148] p=(d_max-offset) / d_max ......(4)
[0150] Where d_max can be the farthest Euclidean distance between the keypoint and all point clouds on the object (which can be calculated from the target prediction value and the point cloud spatial coordinates, or it can be a set threshold constant); offset is the predicted offset (obtained by performing an offset estimation task), i.e. , where dx, dy, and dz are the predicted offset values in the x, y, and z directions, respectively.
[0151] Based on the distance proximity value shown in the above formula (4), it can be determined that the larger the distance proximity value p is, the closer the point cloud is to the key point.
[0152] In one embodiment, step C04 can be represented by the following formula (5):
[0153] y= / ......(5)
[0155] Where i = 1, ..., N; i represents the i-th point cloud, and there are N point clouds in total; x is the target prediction value; p is the distance proximity value shown in formula (4); w is the mask confidence of each cell (the probability value of each cell corresponding to each point cloud), w = ;idx=1,...,D 3 W_ins is the network output probability map for instance mask segmentation, and idx is the label of the cell containing the object. Optionally, w can also be expressed as w = M_idx.
[0156] Specifically, the target predicted value can be represented by the example shown in Table 1 in step C01 above.
[0157] Among them, the spatial coordinate information of key point 1 is as follows:
[0158] The obtained coordinate values.
[0159] By implementing step C04 above, compared to step C01, the influence of outliers in the offset predicted by the offset estimation task can be further removed. In a feasible embodiment, a module for outputting the proximity value p (the output value of this module is the predicted value) can be added before the regressor of the network to reduce the computational load in the regression calculation.
[0160] In one embodiment, the regressor in the network structure performs the keypoint regression task, specifically executing the steps shown in C01-C04 above. During network training, the regressor can be added to the overall network structure for training, or the regressor can be placed outside the network training, that is, during network training, only the network parameters of each task, including feature extraction, semantic segmentation, instance mask segmentation, and keypoint offset estimation, are trained and updated.
[0161] Optionally, considering that extracting only point cloud features including spatial coordinate information has limited expressive power when performing point cloud feature extraction based on depth images, other features such as RGB color and normal vectors are generally added. However, calculating other feature information is time-consuming. To improve the accuracy of the network's pose estimation while ensuring timeliness, the implementation network provided in this application also includes a feature mapping module, such as... Figure 10 or Figure 11 As shown (or in) Figure 3 and Figure 4 Add a feature mapping module; Figure 11 (The color image can also be replaced with a grayscale image). Steps A1 and A2, which extract point cloud features based on the input depth image, may also include the following steps A11-A12:
[0162] Step A11: Obtain the corresponding point cloud information based on the input depth image;
[0163] Step A12: Extract point cloud features based on point cloud information and at least one of color features and normal features.
[0164] Specifically, during network training, when performing point cloud feature extraction, the point cloud information obtained from the input depth image can first be input into the point cloud feature extraction network for feature extraction training. Then, the point cloud information and other feature information (such as RGB color, normal vector, etc.) are input into the point cloud feature extraction network, and the processed feature information is processed by minimizing Euclidean distance (or other distance metric methods can be used). This allows the image features corresponding to the extracted point cloud to have more other feature representations when the network is used for testing, thereby reducing the time required to extract other feature information.
[0165] As mentioned in the above embodiments Figure 4 , Figure 6 and Figure 11 In this context, the module boxes corresponding to sharp angles represent operation steps or network structures, while those corresponding to rounded angles represent processing results. For example... Figure 4 As shown, the block diagram of dense fusion represents the fusion of image pixel features and point cloud features, and the semantic segmentation probability map represents the output data of the multilayer perceptron 1.
[0166] To further illustrate the application of the object pose estimation method provided in the embodiments of this application in real-world scenarios, combined with... Figure 12 The corresponding augmented reality will be explained.
[0167] Figure 12 The illustration shows a scenario where the object pose estimation method provided in this application is applied to an augmented reality system. The object image is captured in real-time by the user's AR glasses, and the captured content is video data corresponding to the three-dimensional space in front of the user when using the augmented reality system. Object pose estimation can be understood as processing each frame of image data in the video data. Combined with... Figure 12 As can be seen, black represents real-world objects, and white represents virtual objects. By executing the method provided in this application's embodiments, virtual content can be aligned in real time to real-world objects with the correct poses. For moving objects in real-world scenes, real-time pose estimation ensures that virtual objects in the augmented reality system, which have few latency artifacts in real-world scenes, are updated promptly. Figure 12 Users can perceive that virtual objects are aligned with real objects in real time for display.
[0168] To better illustrate the effectiveness of the network provided in this application embodiment in implementing the object pose estimation method, the experimental data shown in Table 2 below are provided:
[0169] Table 2
[0170]
[0171] As shown in Table 2, the experimental data in this embodiment uses the YCB video dataset (pose estimation dataset) to compute the inference time of PVN3D (a 3D keypoint detection neural network based on Hough voting) and its various models on a server (object instance segmentation and object keypoint detection are performed using post-processing clustering). In the method provided in this application, a single-branch method with equal spacing is used for the position-aware network (here, D is set to 10). The least squares fitting is used to calculate the 6-DOF pose of the object. Experimental tests show that, compared with the PVN3D method, the method provided in this application achieves faster object pose estimation on different types of GPUs.
[0172] Based on the same inventive concept, embodiments of this application also provide an image processing method, such as... Figure 13 As shown, it includes the following steps S1-S3:
[0173] Step S1: Obtain the image features corresponding to the point cloud of the input image.
[0174] Specifically, the image features corresponding to the point cloud of the input image obtained in step S1 can refer to the content shown in step S101 of the above embodiment. The image can also be a depth image, a color image, or a grayscale image.
[0175] Step S2: Based on image features, the multilayer perceptron network for instance mask segmentation determines the instance mask information of the image using the point cloud corresponding to the object center as a reference.
[0176] Specifically, the image may include one or more objects, and each object may correspond to one or more point clouds. In this embodiment, the instance mask segmentation task is completed using a multilayer perceptron network. The process of this network processing image features can be referred to steps B1 and B2 in the above embodiments.
[0177] The following describes the specific process of training a multilayer perceptron network to complete the instance mask segmentation task. Specifically, the training steps of the multilayer perceptron network for instance mask segmentation include the following steps S21-S23:
[0178] Step S21: Obtain the training dataset; the training dataset includes multiple training images and object annotation information corresponding to each training image.
[0179] Step S22: Input the training image into the multilayer perceptron network for instance mask segmentation, so that the network outputs the predicted grid information corresponding to each point cloud of the object based on the three-dimensional network corresponding to the input training image and the cell where the point cloud corresponding to the center of the object is located.
[0180] Step S23: Based on the predicted mesh information and object annotation information, determine the parameters of the multilayer perceptron network for instance mask segmentation.
[0181] Alternatively, the network training method can include the following two approaches:
[0182] (1) Object annotation information can represent the actual location (grid) of each point cloud of an object in three-dimensional space. A training image can include multiple objects, each of which corresponds to its own annotation information. The multilayer perceptron network for instance mask segmentation can output the initial grid information corresponding to each point cloud of the object. Then, based on the initial grid information and the grid information of the point cloud corresponding to the center of the object, the initial grid information of the point cloud corresponding to the center of the object is finally used as the predicted grid information corresponding to each point cloud of the object. When determining the network parameters, the network parameters are updated based on the initial grid information and the annotation information. Optionally, the loss value of each iteration of the network can be determined based on a preset loss function such as dice loss or softmax loss, and then the network parameters are updated based on the loss value.
[0183] (2) Object annotation information can characterize the actual position (mesh) of the point cloud corresponding to the center of the object in three-dimensional space. The multilayer perceptron network of instance mask segmentation can output the mesh information (predicted mesh information) of the point cloud corresponding to the center of the object; and then update the network parameters based on the predicted mesh information and annotation information.
[0184] That is, during network training, training can be performed only on the grid information of the point cloud corresponding to the center of the object, and the grid information of other point clouds of the object can be regarded as consistent with the grid information of the point cloud corresponding to the center of the object.
[0185] Step S3: Perform image processing based on the instance mask information.
[0186] In one embodiment, the above image processing method further includes step S4: determining semantic segmentation information of the image based on the image features.
[0187] Specifically, the process of determining the semantic segmentation information of the image based on image features in step S4 can be referred to the content shown in step S102 in the above embodiment.
[0188] Step S3, which involves image processing based on the instance mask information, includes the following steps S31-S32:
[0189] Step S31: Determine the instance segmentation information of the object based on the semantic segmentation information and the instance mask information.
[0190] Specifically, the process of determining the instance segmentation information of the object based on semantic segmentation information and instance mask information in step S31 can be referred to the content shown in step C3 in the above embodiment.
[0191] Step S32: Perform image processing based on the instance segmentation information.
[0192] In one embodiment, the above image processing method can be applied to an object pose estimation method to perform object pose estimation.
[0193] Corresponding to the object pose estimation method provided in this application, this application embodiment also provides an object pose estimation device 1400, the structural schematic diagram of which is shown below. Figure 14 As shown, the object attitude estimation device 1400 includes: a first acquisition module 1401, a first determination module 1402, and an attitude estimation module 1403.
[0194] The first acquisition module 1401 is used to acquire image features corresponding to the point cloud of the input image; the first determination module 1402 is used to determine the semantic segmentation information, instance mask information and key point information of the object based on the image features; and the pose estimation module 1403 is used to estimate the pose of the object based on the semantic segmentation information, instance mask information and key point information.
[0195] Optionally, the first acquisition module 1401 is configured to perform at least one of the following:
[0196] Point cloud features are extracted from the input depth image, and the extracted point cloud features are confirmed as the image features corresponding to the point cloud.
[0197] First image features are extracted from the input color image and / or grayscale image; point cloud features are extracted from the input depth image; the first image features and point cloud features are fused to obtain the image features corresponding to the point cloud.
[0198] Optionally, when the first acquisition module 1401 performs point cloud feature extraction based on the input depth image, it is also used to perform:
[0199] Obtain the corresponding point cloud information based on the input depth image;
[0200] Point cloud features are extracted based on point cloud information and at least one of the following:
[0201] Color characteristics and normal characteristics.
[0202] Optionally, when the first acquisition module 1401 fuses the first image features and point cloud features to obtain the image features corresponding to the point cloud, it is also used to perform:
[0203] The first image features and the point cloud features are fused pixel by pixel to obtain the image features corresponding to the point cloud.
[0204] Optionally, when the first acquisition module 1401 is used to extract first image features based on the input color image and / or grayscale image, it is also used to: extract first image features based on the input color image and / or grayscale image using a convolutional neural network; and / or
[0205] When the first acquisition module 1401 is used to perform point cloud feature extraction based on the input depth image, it is also used to perform: extracting point cloud features based on the input depth image through a multilayer perceptron network.
[0206] Optionally, when the first determining module 1402 is used to determine the semantic segmentation information, instance mask information, and key point information of an object based on image features, it is also used to perform:
[0207] Based on image features, determine the semantic segmentation information corresponding to the point cloud;
[0208] A 3D mesh is created based on the point cloud information corresponding to the input image, and the instance mask information of the object is determined based on the 3D mesh.
[0209] Key point information is determined based on instance mask information, or based on semantic segmentation information and instance mask information.
[0210] Optionally, the instance mask information represents the mesh information corresponding to the point cloud in the 3D mesh, wherein the mesh information corresponding to each point cloud of the object is determined based on the mesh information corresponding to the point cloud at the center of the object.
[0211] Optionally, when the first determining module 1402 is used to create a 3D mesh based on the point cloud information corresponding to the input image, it is also used to perform at least one of the following:
[0212] The point cloud information is divided into three-dimensional spaces at equal intervals to obtain a three-dimensional mesh.
[0213] The three-dimensional space corresponding to the point cloud information is divided based on multiple preset intervals to obtain multiple three-dimensional meshes;
[0214] Based on the same spacing but different dividing starting points, multiple three-dimensional meshes are obtained in the three-dimensional space corresponding to the dividing point cloud information.
[0215] Optionally, when the first determining module 1402 is used to determine key point information based on instance mask information, it is also used to perform at least one of the following:
[0216] The first offset of each key point in the point cloud is estimated based on image features; the key point information of the object is determined by regression based on the first offset and instance mask information.
[0217] Based on image features and instance mask information, the second offset of the key points corresponding to each point cloud is estimated in each cell of the 3D mesh; based on the second offset, the key point information of the object is determined by regression.
[0218] Optionally, when the first determining module 1402 is used to determine key point information based on semantic segmentation information and instance mask information, it is also used to perform:
[0219] Instance segmentation information is determined based on semantic segmentation information and instance mask information;
[0220] Estimate the first offset of each key point corresponding to each point cloud based on image features;
[0221] Based on the first offset and instance segmentation information, the key point information of the object is determined by regression.
[0222] Optionally, when the first determining module 1402 is used to determine the key point information of an object by regression based on the first offset and instance mask information, it is also used to: determine the initial predicted value of the key points predicted based on the point cloud based on the first offset and point cloud information; determine the target predicted value of the key points in the 3D mesh based on the initial predicted value and instance mask information; and determine the key point information of the object by regression based on the target predicted value.
[0223] Optionally, when the first determining module 1402 is used to determine the key point information of the object based on the second offset by regression, it is also used to: determine the target predicted value of the key points predicted based on the point cloud based on the second offset and the point cloud information; and determine the key point information of the object based on the target predicted value by regression.
[0224] Optionally, when the first determining module 1402 is used to determine the key point information of the object by regression based on the first offset and instance segmentation information, it is also used to: determine the initial predicted value of the key point predicted based on the point cloud based on the first offset and point cloud information; determine the target predicted value of the key point in the three-dimensional mesh based on the initial predicted value and instance segmentation information; and determine the key point information of the object by regression based on the target predicted value.
[0225] Optionally, when the first determining module 1402 is used to determine the key point information of the object based on the target prediction value through regression, it is also used for at least one of the following:
[0226] For each key point of an object, the average of the target prediction values corresponding to that key point and each point cloud is determined as the key point information;
[0227] For each key point of an object, the key point information is determined by the weighted average of the target prediction value corresponding to the key point and the probability value corresponding to each point cloud in the instance mask information.
[0228] For each key point of an object, the key point information is determined by the weighted average of the target prediction value corresponding to the preset number of point clouds that are closest to the object's center point and the probability value corresponding to the preset number of point clouds in the instance mask information.
[0229] For each key point of an object, the key point information is determined by the weighted average of the proximity value between the key point and each point cloud, the target prediction value between the key point and each point cloud, and the probability value between the key point and each point cloud in the instance mask information.
[0230] Corresponding to the image processing method provided in this application, this application embodiment also provides an image processing apparatus 1500, the structural schematic diagram of which is shown below. Figure 15 As shown, the image processing device 1500 includes: a second acquisition module 1501, a second determination module 1502, and a processing module 1503.
[0231] The second acquisition module 1501 is used to acquire image features corresponding to the point cloud of the input image; the second determination module 1502 is used to determine the instance mask information of the image based on the image features and using the point cloud corresponding to the center of the object as a reference through a multilayer perceptron network for instance mask segmentation; and the third processing module 1503 is used to perform image processing based on the instance mask information.
[0232] Optionally, the training steps for the multilayer perceptron network for instance mask segmentation include:
[0233] Obtain the training dataset; the training dataset includes multiple training images and the object annotation information corresponding to each training image;
[0234] The training image is input into the multilayer perceptron network for instance mask segmentation, so that the network outputs the predicted grid information corresponding to each point cloud of the object based on the 3D network corresponding to the input training image and the cell where the point cloud corresponding to the center of the object is located.
[0235] Based on predicted mesh information and object annotation information, the parameters of a multilayer perceptron network for instance mask segmentation are determined.
[0236] Optionally, the device 1500 further includes a third determining module for determining semantic segmentation information of the image based on image features.
[0237] Optionally, the processing module 1503 is further configured to: determine the instance segmentation information of the object based on the semantic segmentation information and the instance mask information; and perform image processing based on the instance segmentation information.
[0238] The apparatus in this application embodiment can execute the method provided in the embodiments of this application, and the implementation principle is similar. The actions performed by each module in the apparatus in each embodiment of this application correspond to the steps in the method in each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.
[0239] This application also provides an electronic device including a memory and a processor; wherein the memory stores a computer program; and the processor is used to execute the method provided in any optional embodiment of this application when running the computer program.
[0240] This application also provides a computer-readable storage medium storing a computer program that, when run by a processor, executes the methods provided in any optional embodiment of this application.
[0241] As an alternative, Figure 16 A schematic diagram of the structure of an electronic device to which this application is applicable is shown, such as... Figure 16 As shown, the electronic device 1600 may include a processor 1601 and a memory 1603. The processor 1601 and the memory 1603 are connected, for example, via a bus 1602. Optionally, the electronic device 1600 may also include a transceiver 1604. It should be noted that in practical applications, the transceiver 1604 is not limited to one, and the structure of the electronic device 1600 does not constitute a limitation on the embodiments of this application.
[0242] Processor 1601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1601 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0243] Bus 1602 may include a pathway for transmitting information between the aforementioned components. Bus 1602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 1602 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 16 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0244] The memory 1603 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0245] The memory 1603 is used to store application code that executes the scheme of this application, and its execution is controlled by the processor 1601. The processor 1601 is used to execute the application code (computer program) stored in the memory 1603 to implement the content shown in any of the foregoing method embodiments.
[0246] In the embodiments provided in this application, the above-described object pose estimation method performed by an electronic device can be executed using an artificial intelligence model.
[0247] According to embodiments of this application, the method executed in an electronic device can obtain output data that identifies images or image features within images by using image data or video data as input data for an artificial intelligence model. The artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operating rule or artificial intelligence model configured to perform desired features (or objectives) by training a basic artificial intelligence model with multiple training data using a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and neural network computation is performed by calculating the results of the previous layer with the multiple weight values.
[0248] Visual understanding is a technology used to identify and process things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.
[0249] The object pose estimation device provided in this application can implement at least one of multiple modules through an AI model. AI-related functions can be executed using non-volatile memory, volatile memory, and a processor.
[0250] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors (such as central processing unit (CPU), application processor (AP), etc.), or pure graphics processing units (such as graphics processing unit (GPU), visual processing unit (VPU), and / or AI-specific processors (such as neural processing unit (NPU)).
[0251] The one or more processors control the processing of input data based on predefined operating rules or artificial intelligence (AI) models stored in non-volatile and volatile memory. These predefined operating rules or AI models are provided through training or learning.
[0252] Here, "providing through learning" refers to obtaining predefined operating rules or an AI model with desired characteristics by applying a learning algorithm to multiple learning datasets. This learning can be performed within the device itself, in which the AI is executed according to the embodiment, and / or can be implemented via a separate server / system.
[0253] This AI model can consist of multiple neural network layers. Each layer has multiple weight values, and the computation of a layer is performed using the results of the previous layer and the multiple weights of the current layer. Examples of neural networks include, but are not limited to, Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), Generative Adversarial Networks (GANs), and Deep Q-Networks.
[0254] A learning algorithm is a method of training a predetermined target device (e.g., a robot) using multiple learning data sets to enable, allow, or control the target device to make determinations or predictions. Examples of such learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0255] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0256] The above description is only a partial embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method performed by an electronic device, characterized in that, include: Determine image features corresponding to the point cloud of an input image, wherein the input image includes a depth image, and determining image features corresponding to the point cloud of the input image includes: extracting point cloud features based on the depth image, and obtaining image features based on the point cloud features; Based on the image features, semantic segmentation information, instance mask information, and key point information of the object are determined; Object pose estimation is performed based on the semantic segmentation information, instance mask information, and key point information. The determination of the object's instance mask information includes: A 3D mesh is generated based on the point cloud information corresponding to the input image, and the instance mask information of the object is determined based on the 3D mesh. The instance mask information is used to characterize the mesh information corresponding to the point cloud in the 3D mesh, and the mesh information corresponding to each point cloud of the object is determined based on the mesh information corresponding to the point cloud at the center of the object.
2. The method according to claim 1, characterized in that, The process of obtaining image features based on the point cloud features includes: The point cloud features are determined to be image features.
3. The method according to claim 1, characterized in that, Extracting point cloud features based on the depth image includes: Determine the point cloud information corresponding to the depth image; Point cloud features are extracted based on at least one of color features and normal features, as well as the point cloud information.
4. The method according to claim 1, characterized in that, The input image also includes at least one of a color image and a grayscale image; Determine the image features corresponding to the point cloud of the input image, including: Based on at least one of the color image and grayscale image, extract the first image features; Image features are determined by fusing the point cloud features and the first image features.
5. The method according to claim 4, characterized in that, Image features are determined by fusing the point cloud features and the first image features, including: Image features are determined by fusing the point cloud features and the first image features pixel by pixel.
6. The method according to any one of claims 1-5, characterized in that, The determination of semantic segmentation information, instance mask information, and key point information of objects based on the image features includes: Based on the image features, determine the semantic segmentation information corresponding to the point cloud; Key point information is determined based on the instance mask information, or based on the semantic segmentation information and the instance mask information.
7. The method according to any one of claims 1 to 5, characterized in that, Generate a 3D mesh based on the point cloud information corresponding to the input image, including any one or any combination of two or more of the following: A 3D mesh is determined by dividing the 3D space corresponding to the point cloud information at equal intervals. The 3D space corresponding to the point cloud information is divided based on multiple preset intervals to determine multiple 3D meshes; Based on the same starting point with different spacing, the 3D space corresponding to the point cloud information is divided to determine multiple 3D meshes.
8. The method according to claim 6, characterized in that, Determining key point information based on the instance mask information includes: Estimate the first offset of each key point corresponding to the point cloud based on the image features; determine the key point information by regression based on the first offset and instance mask information; and / or Based on the image features and instance mask information, the second offset of each key point corresponding to each point cloud is estimated in each cell of the 3D grid. Based on the second offset, the key point information is determined by regression.
9. The method according to claim 6, characterized in that, Determining key point information based on the instance mask information includes: Based on the image features and the instance mask information, the second offset of each key point corresponding to each point cloud is estimated in each cell of the 3D mesh; Based on the second offset, key point information is determined through regression, including: Based on the second offset and point cloud information, determine the target predicted value of the key points based on point cloud prediction; and Based on the predicted target value, key point information is determined through regression.
10. The method according to claim 9, characterized in that, The determination of key point information based on the target predicted value through regression includes any one or any combination of two or more of the following: For each key point of an object, the average of the target prediction values corresponding to that key point and each point cloud is determined as the key point information; For each key point of an object, the key point information is determined by the weighted average of the target prediction value corresponding to the key point and each point cloud, and the probability value corresponding to each point cloud in the instance mask information. For each key point of an object, the key point information is determined by the weighted average of the target prediction value corresponding to the preset number of point clouds that are closest to the object's center point and the probability value corresponding to the preset number of point clouds in the instance mask information. For each key point of an object, the key point information is determined by the weighted average of the proximity value between the key point and each point cloud, the target prediction value between the key point and each point cloud, and the probability value between the key point and each point cloud in the instance mask information.
11. The method according to claim 6, characterized in that, Key point information is determined based on the semantic segmentation information and instance mask information, including: Based on the semantic segmentation information and the instance mask information, the instance segmentation information is determined; Estimate the first offset of each key point corresponding to each point cloud based on the image features; Based on the first offset and the instance segmentation information, key point information is determined by regression.
12. The method according to claim 11, characterized in that, The step of determining key point information based on the first offset and the instance segmentation information through regression includes: Based on the first offset and the point cloud information, determine the initial predicted value of the key points based on the point cloud prediction; Based on the initial predicted values and instance segmentation information, the target predicted values of key points in the 3D mesh are determined; Based on the predicted target value, key point information is determined through regression.
13. The method according to claim 8, characterized in that, The step of determining key point information based on the first offset and instance mask information through regression includes: Based on the first offset and the point cloud information, determine the initial predicted value of the key points based on the point cloud prediction; Based on the initial predicted values and instance mask information, determine the target predicted values of key points in the 3D mesh; Based on the predicted target value, key point information is determined through regression.
14. The method according to claim 1, characterized in that, Object pose estimation includes estimating the object's orientation.
15. The method according to claim 14, characterized in that, The estimated attitude is a 6-DOF attitude.
16. The method according to claim 1, characterized in that, The key point information refers to the object's position information.
17. The method according to claim 1, characterized in that, The determination of the key point information includes: Based on the instance segmentation information of the identified objects, key point information is determined.
18. The method according to claim 17, characterized in that, Also includes: The instance segmentation information is determined based on the semantic segmentation information and the instance mask information.
19. The method according to claim 1, characterized in that, The instance mask information indicates the region in three-dimensional space that contains point cloud information of the object and / or the region in three-dimensional space where point cloud information of the object does not exist.
20. A method performed by an electronic device, characterized in that, include: Based on point cloud information corresponding to the input image, image features are determined, wherein the input image includes a depth image, and the determination of image features includes: extracting point cloud features based on the depth image, and obtaining image features based on the point cloud features; Based on the image features, semantic segmentation information is determined; Determine instance mask information based on point cloud information; Based on the instance mask information, key point information is determined; Based on the semantic segmentation information, instance mask information, and key point information, the pose of the object is determined; The determination of instance mask information includes: A 3D mesh is generated based on the point cloud information corresponding to the input image, and the instance mask information of the object is determined based on the 3D mesh. The instance mask information is used to characterize the mesh information corresponding to the point cloud in the 3D mesh, and the mesh information corresponding to each point cloud of the object is determined based on the mesh information corresponding to the point cloud at the center of the object.
21. The method according to claim 20, characterized in that, The determination of key information includes: For each point cloud in the point cloud information, the offset of the key point is estimated based on image features; Key point information is determined through regression based on offset.
22. An electronic device, characterized in that, Including memory and processor; The memory contains computer programs; The processor is configured to execute the method of any one of claims 1 to 19 or 20 to 21 when running the computer program.
23. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, performs the method according to any one of claims 1 to 19 or 20 to 21.
24. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 19 or 20 to 21.