Object estimation method, electronic equipment and storage medium
By combining color images and text data to generate normalized object coordinates, the universality problem of object posture and size estimation is solved, and efficient estimation across object categories is achieved.
Patent Information
- Application Number
- CN202410109875.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
Existing methods for estimating objects posture and size lack versatility and cannot effectively estimate across object categories.
By obtaining the color image, depth image and text data of the target object, normalized object coordinates are generated based on the color image and text data, and estimating it in combination with the depth image to generate the size and pose estimation results of the target object.
Position and size estimation across object categories is realized, improving the versatility and efficiency of estimation.
Smart Images

Figure CN120374716A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large model technology and robot control. Specifically, it relates to an object estimation method, an electronic device, and a storage medium. Background Art
[0002] Currently, existing object estimation solutions can be divided into three categories: instance-level, generalized instance-level, and category-level. Among them, instance-level object pose estimation can usually only perform pose recognition on a limited number of objects. Generalized instance-level object pose estimation methods can handle objects not seen in the training stage, but these methods usually require a complete 3D object model or a large number of image templates with object poses in the application stage, which limits the practical application value of such methods. Category-level object pose estimation defines that the same category of objects share a normalized reference coordinate system, and estimates the pose by finding the relationship between the camera coordinate system and this reference coordinate system. Although this method can be generalized to new objects in the same category, it cannot achieve cross-object category estimation, thus making the generality of the estimation method poor when estimating the pose and size of an object.
[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present application provide an object estimation method, an electronic device, and a storage medium to at least solve the technical problem of poor generality of the method for estimating the pose and size of an object.
[0005] According to one aspect of the embodiments of the present application, an object estimation method is provided, including: obtaining a color image, a depth image, and text data of a target object, where the text data is used to describe the target object; generating a normalized object coordinate of the target object based on the color image and the text data; estimating the target object based on the normalized object coordinate and the depth image to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object.
[0006] According to another aspect of the embodiments of the present application, there is also provided an object estimation method, including: in response to an input instruction acting on an operation interface, displaying a color image, a depth image, and text data of a target object on the operation interface, where the text data is used to describe the target object; in response to an estimation instruction acting on the operation interface, displaying an estimation result of the target object on the operation interface, where the estimation result is a result obtained by estimating the target object based on the normalized object coordinates of the target object and the depth image, and the normalized object coordinates are generated based on the color image and the text data.
[0007] According to another aspect of the embodiments of the present application, there is also provided a robot control method, including: acquiring a color image, a depth image, and text data of an object to be grasped, where the text data is used to describe the object to be grasped; generating normalized object coordinates of the object to be grasped based on the color image and the text data; estimating the object to be grasped based on the normalized object coordinates and the depth image to obtain an estimation result of the object to be grasped, where the estimation result is used to characterize the size and posture of the object to be grasped; and controlling the robot to grasp the object to be grasped based on the estimation result.
[0008] According to another aspect of the embodiments of the present application, there is also provided an object estimation method, including: obtaining a color image, a depth image, and text data of a target object by calling a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the color image, the depth image, and the text data, and the text data is used to describe the target object; generating normalized object coordinates of the target object based on the color image and the text data; estimating the target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object; and outputting the estimation result by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the estimation result.
[0009] In an embodiment of the present application, a color image, a depth image, and text data of a target object are obtained, where the text data is used to describe the target object; based on the color image and the text data, normalized object coordinates of the target object are generated; based on the normalized object coordinates and the depth image, an estimation of the target object is performed to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object. It is easy to notice that the normalized object coordinates of the target object can be generated based on the color image and the text data of the target object, and the pose and size of the target object can be estimated by using the normalized object coordinates and the depth image to obtain the estimation result of the target object. That is to say, in the process of estimating the pose and size of the target object, the pose and size of the target object are determined based on the normalized object coordinates, so that not only can the same type of objects share a normalized reference coordinate system, but also the normalized object coordinates are generated based on the color image of the target object and the text data describing the target object. That is to say, when generating the normalized object coordinates, not only the color image of the target object is considered, but also the description text of the target object is considered. Through the text data, different types of target objects can be described and the normalized object coordinates can be generated, thereby realizing the estimation of the pose and size across object categories, improving the estimation efficiency when estimating the pose and size of the target object, and solving the technical problem of poor generality of the method for estimating the pose and size of an object.
[0010] It is easy to notice that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. Brief Description of the Drawings
[0011] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0012] Figure 1 is a schematic diagram of an application scenario of an object estimation method according to an embodiment of the present application;
[0013] Figure 2 is a flowchart of an object estimation method according to Embodiment 1 of the present application;
[0014] Figure 3 is a schematic diagram of an estimation of a target object according to an embodiment of the present application;
[0015] Figure 4 is a flowchart of an object estimation method according to Embodiment 2 of the present application;
[0016] Figure 5It is a flowchart of an object estimation method according to Embodiment 3 of the present application;
[0017] Figure 6 It is a flowchart of a robot control method according to Embodiment 4 of the present application;
[0018] Figure 7 It is a schematic diagram of an object estimation device according to Embodiment 5 of the present application;
[0019] Figure 8 It is a schematic diagram of an object estimation device according to Embodiment 6 of the present application;
[0020] Figure 9 It is a schematic diagram of an object estimation device according to Embodiment 7 of the present application;
[0021] Figure 10 It is a schematic diagram of a robot control device according to Embodiment 8 of the present application;
[0022] Figure 11 It is a structural block diagram of a computer terminal according to an embodiment of the present application. Detailed implementation manners
[0023] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] The technical solution provided by this application is mainly implemented using large model technology. Here, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be called a foundation model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one hundred million parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0026] It should be noted that in actual applications, the pre-trained model can be fine-tuned with a small number of samples so that the large model can be applied to different tasks. For example, large models can be widely applied in the fields of natural language processing (NLP), computer vision, speech processing, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., and can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. Therefore, the main application scenarios of large models include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiments of this application, the data processing by the basic vision-language large model in the robot control scenario is taken as an example for explanation.
[0027] First, some nouns or terms that appear in the process of describing the embodiments of this application are applicable to the following explanations:
[0028] CLIP: Contrastive Language-Image Pre-Training, a model for matching the relevance between text and images.
[0029] DinoLearning Robust Visual Features without Supervision: A general visual feature extraction module.
[0030] VQVAE: Vector Quantization Variational Auto Encoder, a VAE model with vector quantization.
[0031] NOCS: Normalized Object Coordinate Space, the normalized object coordinate space.
[0032] UNet network: Also known as the U-shaped network, it is a deep learning network structure for image segmentation. Its characteristic is that it has a U-shaped network structure, which can effectively extract features in the image and perform pixel-level segmentation.
[0033] ReLU: Rectified Linear Unit, a non-linear activation function commonly used in the hidden layers of neural networks.
[0034] Umeyama algorithm: A method for calculating the pose matrix by minimizing the root mean square deviation between two point sets, used to calculate the rigid transformation (rotation, translation, and scaling) between two sets of point sets, and can also be used for matching between two point clouds. For example, it is commonly used in computer vision and 3D reconstruction.
[0035] Embodiment 1
[0036] According to an embodiment of the present application, an object estimation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0037] Considering that the number of model parameters of the large model is huge and the computing resources of the mobile terminal are limited, Figure 1 It is a schematic diagram of an application scenario of an object estimation method according to an embodiment of the present application. The above object estimation method provided by the embodiments of the present application can be applied to such as Figure 1 the application scenario shown, but not limited to this. In the application scenario shown in Figure 1 the large model is deployed in the server 10. The server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include but are not limited to: smart phones, tablets, laptops, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. The client device 20 can interact with the user through a graphical user interface to call the large model, and thus implement the method provided by the embodiments of the present application.
[0038] In the embodiments of the present application, the system composed of a client device and a server may perform the following steps: The client 20 executes to give an estimation instruction based on the graphical user interface on the client 20 and the original image containing the target object. The server 10 executes to obtain the color image, depth image, and text data of the target object based on the estimation instruction given by the client 20 and the original image containing the target object provided by the user, where the text data is used to describe the target object; generate the normalized object coordinates of the target object based on the color image and the text data; estimate the target object based on the normalized object coordinates and the depth image to obtain the estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object. It should be noted that in the case where the operating resources of the client device can meet the deployment and running conditions of the large model, the embodiments of the present application can be performed in the client device.
[0039] In the above operating environment, the present application provides an Figure 2 object estimation method as shown below. Figure 2 It is a flowchart of the object estimation method according to Embodiment 1 of the present application. As Figure 2 shown, the method may include the following steps:
[0040] Step S202: Obtain the color image, depth image, and text data of the target object, where the text data is used to describe the target object.
[0041] The above target object may be an object for which pose and size estimation are required, where the target object may be represented in the form of a picture.
[0042] The above color image may be used to represent the color information contained in the target object. Optionally, the color image may be a Red Green Blue image (abbreviated as RGB image).
[0043] The above depth image may be the depth information of the target object in space. Optionally, the depth image of an object may be obtained by using a depth sensor or structured light technology. Such an image helps to understand the three-dimensional shape and position of the object. In the fields of computer vision and robotics, depth images are widely used in tasks such as object recognition, pose and size estimation, and environmental perception. Optionally, the above depth image may be a Red Green Blue Depth image (abbreviated as RGBD image).
[0044] In an optional embodiment, the user may provide the color image and depth image of the target object, and obtain the text data corresponding to the target object from the database, where the above database may be a database on the Internet, and there is a large amount of text descriptions of images in this database.
[0045] In another alternative embodiment, a camera or a video camera can also be used to capture a color image of the target object, and a 3D camera or a depth sensor can be used to obtain a depth image of the target object. Further, speech recognition technology or manual input can be used to obtain text data. Among them, speech recognition technology can convert speech into text, and manual input is to input text descriptions through a keyboard or a touch screen.
[0046] Step S204: Generate the normalized object coordinates of the target object based on the color image and the text data.
[0047] The above-mentioned normalized object coordinates of the target object can be the coordinate representation of the target object in the reference coordinate system.
[0048] In an alternative embodiment, after obtaining the color image and text data of the target object, the text data can be input into a pre-trained CLIP model to obtain the text features corresponding to the target object, and the color image can be passed into the encoder module of VQVAE to extract the corresponding image features. Further, the UNet network in the text-to-image stable diffusion model (also known as Stable Diffusion) can be used to fuse the obtained text features and image features, so as to extract the key information of the target object, and use the key information of the target object to generate the normalized object coordinates of the target object. Optionally, since it is necessary to predict the normalized object coordinates corresponding to each pixel of the target object in the image, the extracted image features have different responses in different object parts.
[0049] In another alternative embodiment, target detection can also be performed using the color image to find the region containing the target object in the image. Specifically, a deep learning model can be used for target detection. Further, object recognition can be performed, that is, on the basis of target detection, the detected target object is recognized using the text data to obtain the specific category information of the target object. Then, the normalized object coordinate calculation can be performed, that is, according to the position information of the target object region obtained by target detection and combined with the size of the image, the normalized object coordinates of the target object are calculated. Optionally, the specific calculation method can be adjusted according to actual requirements and data formats. For example, the center point coordinates of the target object are normalized to the range of [0,1]. Finally, the generated normalized object coordinates and the category information of the target object can be integrated into a data structure for subsequent applications and analyses.
[0050] Step S206: Estimate the target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object.
[0051] In an alternative embodiment, after obtaining the normalized object coordinates, the pose and size of the target object can be estimated based on the normalized object coordinates and the depth image to obtain the estimation result of the target object. Optionally, when performing pose and size estimation, the size and pose of the target object can be fitted by means of least squares. Optionally, the depth map can be converted into a point cloud in the camera coordinate system by using the camera internal parameters. Further, according to the position of the pixel, the one-to-one correspondence between the normalized object coordinates of the target object and the point cloud in the camera coordinate system can be found, that is, the position of each point in the normalized object coordinates in the observed point cloud can be found. Optionally, through these two correspondences, the size and pose regression of the target object can be performed, that is, the pose and size of the target object can be estimated to obtain the estimation result of the target object.
[0052] In another alternative embodiment, the pose and size of the target object can also be estimated based on the normalized object coordinates and the depth image by using computer vision algorithms. Optionally, common methods can include model-based methods, feature point-based methods, etc. Among them, model-based methods usually require building a three-dimensional model of the target object and then estimating the pose by matching the model and the depth image, while feature point-based methods can estimate the pose by detecting feature points in the depth image and combining the normalized object coordinates. Further, the estimation result of the target object can be determined according to the output result of the estimation algorithm. Specifically, it can include information such as the rotation angle and translation vector of the target object.
[0053] Optionally, assuming that the above method of the present application is applied in a robot grasping and object manipulation scenario, first, for a specific object, the grasp pose of the manipulator that can grasp the object can be predefined, and then the pose of the same type of object in the scene can be recognized by using this method, and finally, the predefined grasp pose can be mapped to the scene according to the recognized pose, so as to obtain the grasp pose of the manipulator that can grasp the object in the scene.
[0054] Optionally, the augmented reality scenario is similar to the robot grasping and object manipulation scenario. After the pose of the object in the actual scene is recognized, the object in the virtual scene can be mapped to the real scene. For the task of object reconstruction, since this method has a consistent object reference coordinate system for the same object, that is, the reference coordinate systems corresponding to the poses estimated at different angles of the object are consistent, this solution can be used to realize object reconstruction.
[0055] In an embodiment of the present application, by obtaining a color image, a depth image, and text data of a target object, where the text data is used to describe the target object; based on the color image and the text data, generating normalized object coordinates of the target object; and estimating the target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object. It is easy to notice that the normalized object coordinates of the target object can be generated based on the color image and the text data of the target object, and the pose and size of the target object can be estimated by using the normalized object coordinates and the depth image to obtain the estimation result of the target object. That is, in the process of estimating the pose and size of the target object, the pose and size of the target object are determined based on the normalized object coordinates, so that not only can the same type of objects share a normalized reference coordinate system, but also the normalized object coordinates are generated based on the color image of the target object and the text data describing the target object. That is, when generating the normalized object coordinates, not only the color image of the target object is considered, but also the description text of the target object is considered. The text data can be used to describe different types of target objects and generate normalized object coordinates, thereby realizing the estimation of the pose and size across object categories, improving the estimation efficiency when estimating the pose and size of the target object, and further solving the technical problem of poor generality of the method for estimating the pose and size of an object.
[0056] In the above embodiment of the present application, generating the normalized object coordinates of the target object based on the color image and the text data includes: performing feature fusion on the color image and the text data to obtain a fusion feature, where the fusion feature carries semantic information of the target object; using a first image feature extraction module to extract features from the color image to obtain a first image feature, where the first image feature is used to characterize different parts of the target object; and generating the normalized object coordinates based on the fusion feature and the first image feature.
[0057] The above first image feature extraction module can be used to extract features from the color image, so that the extracted image features have different responses at different object parts. Optionally, in the present application, no specific limitation is imposed on the first image feature extraction module. In the present application, the first image feature extraction module is taken as a visual feature extraction module (also called the DinoV2 module) for illustration.
[0058] In an alternative embodiment, after obtaining the color image and text data of the target object, the color image can be input into the DinoV2 module to extract the corresponding image feature, i.e., the Dino Feature Map (also referred to as the Dino feature map), which is the first image feature. Further, the UNet network in the stable diffusion model from text to image (also referred to as StableDiffusion) can be used to fuse the obtained text feature with the first image feature, thereby extracting the key information of the target object and generating the normalized object coordinates of the target object using the key information of the target object. Optionally, since it is necessary to predict the normalized object coordinates corresponding to each pixel of the target object in the image, the first image feature extracted has different responses at different parts of the target object.
[0059] In the above embodiment of the present application, feature fusion is performed on the color image and text data to obtain a fusion feature, including: using a second image feature extraction module to extract features from the color image to obtain a second image feature; using a text feature extraction module to extract features from the text data to obtain a text feature; using a fusion module to perform feature fusion on the second image feature and the text feature to obtain a fusion feature.
[0060] The above-mentioned second image feature extraction module can be used to extract features from the color image. Optionally, there is no specific limitation on the second image feature extraction module in the present application. In the present application, the encoder module of VQVAE is taken as an example to illustrate.
[0061] The above-mentioned text feature extraction module can be used to extract features from the text data. Optionally, there is no specific limitation on the text feature extraction module in the present application. In the present application, the pre-trained CLIP model is taken as an example to illustrate.
[0062] In an alternative embodiment, in order to improve the accuracy of the obtained fusion feature, the encoder module of VQVAE can also be used to perform additional feature extraction on the RGB image to obtain a second image feature, and the text data is input into the pre-trained CLIP model to obtain the text feature corresponding to the target object. Further, the UNet network in Stable Diffusion can be used to fuse the obtained text feature with the second image feature, thereby obtaining a Stable Diffusion feature map (also referred to as the SD feature map) with the semantic information of the target object, i.e., obtaining a fusion feature.
[0063] In the above embodiments of the present application, generating normalized object coordinates based on the fused feature and the first image feature includes: concatenating the fused feature and the first image feature to obtain a concatenated feature; using a decoder to decode the concatenated feature to obtain normalized object coordinates.
[0064] In an alternative embodiment, after obtaining the fused feature and the first image feature, the two features, that is, the fused feature and the first image feature, can be concatenated to obtain a concatenated feature. Further, the concatenated feature can be fed into a decoder module. Optionally, the decoder can include a 5-layer convolutional neural network. Each layer of the convolutional neural network can include a Rectified Linear Unit (ReLU) activation function and a batch normalization module. The feature can be linearly operated through the convolutional neural network to obtain an estimated NOCS map, that is, to obtain normalized object coordinates.
[0065] In the above embodiments of the present application, concatenating the fused feature and the first image feature to obtain a concatenated feature includes: performing bilinear interpolation on the first image feature to obtain an interpolated feature, where the size of the interpolated feature is the same as that of the fused feature; concatenating the fused feature and the interpolated feature to obtain a concatenated feature.
[0066] In an alternative embodiment, since the scales of the above-extracted SD feature map and Dino feature map may be inconsistent, that is, the size of the SD feature map is 1 / 32 of the original RGB image, while the Dino feature map is 1 / 14 of the original RGB image. Therefore, before estimating the NOCS map using the SD feature map and the Dino feature map, bilinear interpolation can be performed on the Dino feature map to make the size of the Dino feature map consistent with that of the SD feature map. Then, the fused feature and the interpolated feature with the same size can be concatenated to obtain a concatenated feature.
[0067] In the above embodiments of the present application, during the training of the fusion module and the decoder, the network parameters of the first image feature extraction module, the second image feature extraction module, and the text feature extraction module are controlled to remain unchanged.
[0068] In an alternative embodiment, before using the fusion module to perform feature fusion on the color image and text data, the fusion module needs to be trained to ensure that the fusion module can be used to perform feature fusion on the color image and text data normally. Optionally, during the training of the fusion module, it is necessary to ensure that the network parameters of the first image feature extraction module, the second image feature extraction module, and the text feature extraction module remain unchanged.
[0069] Further, before decoding the spliced features using the decoder, the decoder also needs to be trained. Optionally, when training the decoder, it is also necessary to ensure that the network parameters of the first image feature extraction module, the second image feature extraction module, and the text feature extraction module remain unchanged.
[0070] In the above embodiments of the present application, estimating a target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object includes: converting the depth image into point cloud data of the target object; determining the corresponding positions of the points of the normalized object coordinates in the point cloud data; and performing an estimation based on the point cloud data and the corresponding positions to obtain the estimation result.
[0071] In an alternative embodiment, the camera intrinsic parameters can be used to convert the depth map into point cloud data of the target object in the camera coordinate system. Further, the corresponding relationship between the object normalized coordinates and the point cloud data of the target object can be found according to the positions of the pixels, that is, the positions of each point in the normalized object coordinates in the observed point cloud can be found. Optionally, to avoid the influence of noise and other outliers, the Random Sample Consensus algorithm (also known as the RANSAC algorithm) can be combined with the Umeyama algorithm, so that the size and pose of the object can be estimated more robustly, that is, an estimation is performed based on the point cloud data and the corresponding positions to obtain the estimation result.
[0072] In the above embodiments of the present application, performing an estimation based on the point cloud data and the corresponding positions to obtain the estimation result includes: selecting a first point from the points of the normalized object coordinates; performing an estimation based on the first point in the point cloud data and the first corresponding position of the first point in the point cloud data to obtain an initial estimation result; selecting a second point from the remaining points of the normalized object coordinates based on the initial estimation result, where the remaining points are used to represent the points in the normalized object coordinates other than the first point; adding the second point to the first point, and repeating the steps of obtaining the initial estimation result and selecting the second point until a preset condition is satisfied; and determining the initial estimation result as the estimation result.
[0073] The above preset condition can be that all the points in the point cloud data have been used in the estimation process, or a large part of the points in the point cloud data have been used in the estimation process, that is, the obtained initial estimation result can represent the estimation result of the target object.
[0074] In an alternative embodiment, when making an estimation based on point cloud data and corresponding positions, a first point may be selected from the points in the normalized object coordinates first. Herein, the first point may be set by those skilled in the art according to requirements or randomly selected. Thus, an initial estimation result may be obtained based on the first point and its corresponding position in the point cloud data, that is, the first corresponding position. Further, a second point may be selected from the remaining points in the normalized object coordinates except the first point based on the initial estimation result, and the second point may be added to the first point. Further, the above operation process may be repeatedly executed until the obtained initial estimation result can represent the estimation result of the target object, that is, until a preset condition is satisfied, and the initial estimation result may be determined as the estimation result. When selecting the second point, a point may be randomly selected from the remaining points in the normalized object coordinates except the first point and used as the second point, or a point closest to the first point may be selected as the second point. In the present application, there is no specific limitation on the selection method of the second point. Optionally, by performing a loop operation through the above process, different points in the point cloud data can be quickly traversed, and thus the estimation result can be obtained.
[0075] In the above embodiment of the present application, adding the second point to the first point includes: determining the number of the second points; and adding the second point to the first point when the number is greater than or equal to a preset number.
[0076] The above preset number may be set by those skilled in the art according to requirements. Optionally, there is no specific limitation on the number of the second points in the present application.
[0077] In an alternative embodiment, when adding the second point to the first point, the data of the second point needs to be considered, that is, the second point can be added to the first point only when the number of the second points is greater than or equal to a preset data. Optionally, by applying the point clouds in the point cloud data that are greater than or equal to the preset number to the estimation process, the accuracy of the estimation result can be improved.
[0078] In the above embodiment of the present application, when the number is less than the preset number, the method further includes: reselecting a first point from the points in the normalized object coordinates, and repeatedly executing the steps of obtaining the initial estimation result and selecting the second point until the preset condition is satisfied; and determining the initial estimation result as the estimation result.
[0079] In an alternative embodiment, when the number of second points is less than a preset number, it is necessary to reselect the first point from the normalized object coordinates and perform an estimation based on the first point and the first corresponding position to obtain an initial estimation result. Further, reselect the second point, add the second point to the first point, and repeat the estimation process until the determined initial estimation result can represent the estimation result of the target object, that is, until a preset condition is met, and determine the initial estimation result as the estimation result. Optionally, when the number of second points is less than the preset number, it can be considered that the accuracy of the obtained initial estimation result is not high enough. Therefore, more point clouds are needed for estimation. By selecting a higher number of point clouds for estimation, the accuracy of the estimation result can be improved.
[0080] In the above embodiments of the present application, obtaining a color image of a target object includes: obtaining an original image containing the target object; performing object detection on the original image to obtain a mask image of the target object; and cropping the original image based on the mask image to obtain a color image.
[0081] The above original image of the target object may be an image containing the target object.
[0082] The above mask image may be an image processing technique that changes or hides certain parts of an image by using a mask. Mask images are commonly used in the fields of image processing and computer vision and can be used for image segmentation, filtering, enhancement, or hiding.
[0083] In an alternative embodiment, after obtaining the original image of the target object, a relatively mature object detection and segmentation model can be used to process the original image of the target object to obtain a mask image of the target object. Further, the original image can be cropped based on the mask image to obtain a color image.
[0084] In the above embodiments of the present application, obtaining text data of a target object includes: obtaining an original image containing the target object; performing object classification on the original image to obtain the target category of the target object; and obtaining the text data corresponding to the target category from a text database.
[0085] In an alternative embodiment, after obtaining the original image containing the target object, the text data of the target object can be obtained by looking up the corresponding text database based on the object category of the target object predicted by the detection and segmentation model, that is, performing object classification on the original image to obtain the target category of the target object, and obtaining the text data corresponding to the target category from the text database.
[0086] Optionally, compared with the features required for learning pose estimation from scratch in a smaller pose estimation dataset, the features extracted by the basic large model have stronger generalization performance and are more suitable for cross-category object estimation. Therefore, a pre-trained basic vision-language large model can be used to extract features of image texts. Optionally, this application gives full play to the alignment relationship learned in the pre-trained basic vision-language large model and uses this alignment relationship to guide the training of the model. When using this alignment relationship to guide the training of the model, the following steps can be taken:
[0087] Data preparation: Prepare a dataset similar to the pre-trained model, including images and corresponding labels. These images can be images containing different object categories, and the labels can be the names of object categories.
[0088] Fine-tune the pre-trained model: Use the prepared dataset to fine-tune the pre-trained model to adapt to the new task. The method of transfer learning can be adopted, freeze the underlying weights of the model, and add new fully connected layers or classifiers at the top layer, and then train on the new dataset.
[0089] Learn the alignment relationship: During the fine-tuning process, the model will learn the alignment relationship between the image and the corresponding label, which means that the model will learn the similarities and differences between different object categories, so as to better understand the relationship between different categories.
[0090] Furthermore, this alignment relationship can be applied again in the application stage to transfer the knowledge of object estimation to new categories. Specifically, the following steps can be taken:
[0091] Feature extraction: Use the fine-tuned model to extract the feature vectors of the image. These feature vectors contain the knowledge of the alignment relationship.
[0092] Similarity comparison: Compare the similarity between the images of the new category and the learned categories, and judge the relationship between the new category and the learned categories by calculating the distance or similarity between the feature vectors.
[0093] Knowledge transfer: According to the results of the similarity comparison, the knowledge of the learned categories can be transferred to the new category, so as to improve the accuracy of object estimation of the new category.
[0094] That is, perform object classification on the original image to obtain the target category of the target object, and obtain the text data corresponding to the target category from the text database. Optionally, by using this alignment relationship, the knowledge of object estimation is transferred to the new category, realizing the size and estimation across object categories. Compared with the given object 3D model or a large number of template guiding information in the existing work, using text as a prior has a lower cost and can make full use of the information of the existing large model, and its generalization performance has more advantages.
[0095] Figure 3 A schematic diagram of the estimation of a target object according to an embodiment of the present application is shown as Figure 3 shown. The user can provide an original image containing the target object, perform object detection on the original image to obtain a mask image of the target object, crop the original image based on the mask image to obtain a color image. Optionally, a visual feature extraction module can be used to extract features from the color image to obtain a first image feature, that is, a Dino feature map. In addition, the encoder of VQVAE can be used to extract features from the color image to obtain a second image feature. A text feature can be obtained by using a text-image correlation matching model, that is, a CLIP model to extract features from text data. Optionally, the UNet network in the Stable Diffusion model (also known as StableDiffus ion) can be used to fuse the obtained text feature and the second image feature to obtain a fused feature, that is, an SD feature map. Then, the fused feature and the first image feature are concatenated to obtain a concatenated feature. Thus, a decoder can be used to decode the concatenated feature to obtain normalized object coordinates. Optionally, the camera internal parameters can be used to convert the depth map into point cloud data of the target object in the camera coordinate system. Optionally, in order to avoid the influence of noise and other outliers, the Random Sample Consensus algorithm (also known as the RANSAC algorithm) can be combined with the Umeyama algorithm to estimate the target object using the depth image and the normalized object coordinates, and an estimation result can be obtained, enabling a more robust estimation of the size and pose of the object.
[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0097] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0099] Embodiment 2
[0100] According to an embodiment of the present application, there is also provided an object estimation method. Figure 4 is a flowchart of the object estimation method according to Embodiment 2 of the present application, as Figure 4 shown, the method includes the following steps:
[0101] Step S402: In response to an input instruction acting on the operation interface, display a color image, a depth image, and text data of the target object on the operation interface, where the text data is used to describe the target object;
[0102] Step S404: In response to an estimation instruction acting on the operation interface, display an estimation result of the target object on the operation interface, where the estimation result is the result obtained by estimating the target object based on the normalized object coordinates and the depth image of the target object, and the normalized object coordinates are generated based on the color image and the text data.
[0103] In an optional embodiment, when the user gives an input instruction on the operation interface of the client 40, the server 41 can respond to the input instruction and control the display of the color image, the depth image, and the text data of the target object on the operation interface. Further, when the user gives an estimation instruction on the operation interface of the client 40, the server 41 can respond to the estimation instruction and control the display of the estimation result of the target object on the operation interface, where the estimation result is the result obtained by estimating the target object based on the normalized object coordinates and the depth image of the target object, and the normalized object coordinates are generated based on the color image and the text data.
[0104] Embodiment 3
[0105] According to an embodiment of the present application, there is also provided an object estimation method. Figure 5 is a flowchart of the object estimation method according to Embodiment 3 of the present application, as Figure 5 shown, the method includes the following steps:
[0106] Step S502: Obtain the color image, depth image, and text data of the target object by calling the first interface. The first interface includes a first parameter, and the parameter value of the first parameter includes the color image, depth image, and text data. The text data is used to describe the target object.
[0107] Step S504: Generate the normalized object coordinates of the target object based on the color image and the text data.
[0108] Step S506: Estimate the target object based on the normalized object coordinates and the depth image to obtain the estimation result of the target object.
[0109] Step S508: Output the estimation result by calling the second interface. The second interface includes a second parameter, and the parameter value of the second parameter includes the estimation result.
[0110] In an optional embodiment, the server 51 can control the client 50, and obtain the color image, depth image, and text data of the target object through the first interface on the client 50. Further, the server 51 can generate the normalized object coordinates of the target object based on the color image and the text data, estimate the target object based on the normalized object coordinates and the depth image to obtain the estimation result of the target object, and control the second interface on the client 50 to output the estimation result.
[0111] Embodiment 4
[0112] According to another aspect of the embodiments of the present application, a robot control method is further provided. Figure 6 is a flowchart of the robot control method according to Embodiment 4 of the present application, as Figure 6 shown, the method includes:
[0113] Step S602: Obtain the color image, depth image, and text data of the object to be grasped. The text data is used to describe the object to be grasped.
[0114] Step S604: Generate the normalized object coordinates of the object to be grasped based on the color image and the text data.
[0115] Step S606: Estimate the object to be grasped based on the normalized object coordinates and the depth image to obtain the estimation result of the object to be grasped. The estimation result is used to characterize the size and posture of the object to be grasped.
[0116] Step S608: Control the robot to grasp the object to be grasped based on the estimation result.
[0117] In an alternative embodiment, when controlling a robot to grasp an object, a color image, a depth image, and text data of the object to be grasped may be acquired first. Further, based on the color image and the text data, a normalized object coordinate of the object to be grasped may be generated, and based on the normalized object coordinate and the depth image, an estimation of the object to be grasped may be performed to obtain an estimation result of the object to be grasped, so as to control the robot to grasp the object to be grasped.
[0118] Embodiment 5
[0119] According to an embodiment of the present application, there is also provided an apparatus for implementing the above object estimation method. Figure 7 is a schematic diagram of an object estimation apparatus according to Embodiment 5 of the present application, as Figure 7 shown, the apparatus includes: an acquisition module 702, a generation module 704, and an estimation module 706.
[0120] Among them, the acquisition module 702 is configured to acquire a color image, a depth image, and text data of a target object, where the text data is used to describe the target object; the generation module 704 is configured to generate a normalized object coordinate of the target object based on the color image and the text data; the estimation module 706 is configured to estimate the target object based on the normalized object coordinate and the depth image to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object.
[0121] In the above embodiment of the present application, the generation module 704 includes: a fusion unit configured to perform feature fusion on the color image and the text data to obtain a fusion feature, where the fusion feature carries semantic information of the target object; an extraction unit configured to extract features from the color image by using a first image feature extraction module to obtain a first image feature, where the first image feature is used to characterize different parts of the target object; a generation unit configured to generate a normalized object coordinate based on the fusion feature and the first image feature.
[0122] In the above embodiment of the present application, the fusion unit includes: a first extraction subunit configured to extract features from the color image by using a second image feature extraction module to obtain a second image feature; a second extraction subunit configured to extract features from the text data by using a text feature extraction module to obtain a text feature; a fusion subunit configured to perform feature fusion on the second image feature and the text feature by using a fusion module to obtain a fusion feature.
[0123] In the above embodiment of the present application, the generation unit includes: a splicing subunit configured to splice the fusion feature and the first image feature to obtain a spliced feature; a decoding subunit configured to decode the spliced feature by using a decoder to obtain a normalized object coordinate.
[0124] In the above embodiments of the present application, the splicing subunit is further configured to perform bilinear interpolation on the first image feature to obtain an interpolation feature, where the size of the interpolation feature is the same as that of the fusion feature; and splice the fusion feature and the interpolation feature to obtain a spliced feature.
[0125] In the above embodiments of the present application, the estimation module 706 includes: a conversion unit configured to convert a depth image into point cloud data of a target object; a determination unit configured to determine the corresponding position of a point with normalized object coordinates in the point cloud data; and an estimation unit configured to perform an estimation based on the point cloud data and the corresponding position to obtain an estimation result.
[0126] In the above embodiments of the present application, the estimation unit includes: a first selection subunit configured to select a first point from the points with normalized object coordinates; an estimation subunit configured to perform an estimation based on the first point in the point cloud data and the first corresponding position of the first point in the point cloud data to obtain an initial estimation result; a second selection subunit configured to select a second point from the remaining points of the normalized object coordinates based on the initial estimation result, where the remaining points are used to represent the points other than the first point in the normalized object coordinates; an addition subunit configured to add the second point to the first point and repeat the steps of obtaining the initial estimation result and selecting the second point until a preset condition is satisfied; and a determination subunit configured to determine the initial estimation result as the estimation result.
[0127] In the above embodiments of the present application, the addition subunit is further configured to determine the number of the second points; and add the second point to the first point when the number is greater than or equal to a preset number.
[0128] In the above embodiments of the present application, the apparatus further includes: a selection module configured to reselect a first point from the points with normalized object coordinates and repeat the steps of obtaining the initial estimation result and selecting the second point until a preset condition is satisfied; and a determination module configured to determine the initial estimation result as the estimation result.
[0129] In the above embodiments of the present application, the acquisition module 702 includes: a first acquisition unit configured to acquire an original image including a target object; a detection unit configured to perform target detection on the original image to obtain a mask image of the target object; and a cropping unit configured to crop the original image based on the mask image to obtain a color image.
[0130] In the above embodiments of the present application, the acquisition module 702 further includes: a second acquisition unit configured to acquire an original image including a target object; a classification unit configured to perform target classification on the original image to obtain the target category of the target object; and a second acquisition unit configured to acquire text data corresponding to the target category from a text database.
[0131] It should be noted here that the above-mentioned acquisition module 702, generation module 704, and estimation module 706 correspond to steps S202 to S206 in Embodiment 1. The examples and application scenarios implemented by the module and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned module or unit may be a hardware component or software component stored in a memory and processed by one or more processors, and the above-mentioned module may also be part of a device and may run in the server 10 provided in Embodiment 1.
[0132] It should be noted that the preferred implementation schemes involved in the above-mentioned embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0133] Embodiment 6
[0134] According to an embodiment of the present application, there is also provided another device for implementing the above object estimation method. Figure 8 It is a schematic diagram of an object estimation device according to Embodiment 6 of the present application, as Figure 8 shown. The device includes: a first display module 802 and a second display module 804.
[0135] Among them, the first display module 802 is configured to respond to an input instruction acting on an operation interface and display a color image, a depth image, and text data of a target object on the operation interface, where the text data is used to describe the target object; the second display module 804 is configured to respond to an estimation instruction acting on the operation interface and display an estimation result of the target object on the operation interface, where the estimation result is a result obtained by estimating the target object based on the normalized object coordinates and the depth image of the target object, and the normalized object coordinates are generated based on the color image and the text data.
[0136] It should be noted here that the above-mentioned first display module 802 and second display module 804 correspond to steps S402 to S404 in Embodiment 2. The examples and application scenarios implemented by the module and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned module or unit may be a hardware component or software component stored in a memory and processed by one or more processors, and the above-mentioned module may also be part of a device and may run in the server 10 provided in Embodiment 1.
[0137] It should be noted that the preferred implementation schemes involved in the above-mentioned embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 2, but are not limited to the schemes provided in Embodiment 2.
[0138] Embodiment 7
[0139] According to an embodiment of the present application, there is also provided another apparatus for implementing the above object estimation method. Figure 9 is a schematic diagram of an object estimation apparatus according to Embodiment 7 of the present application, as Figure 9 shown, the apparatus includes: an acquisition module 902, a generation module 904, an estimation module 906, and an output module 908.
[0140] Among them, the acquisition module 902 is configured to obtain a color image, a depth image, and text data of a target object by calling a first interface. The first interface includes a first parameter, and the parameter value of the first parameter includes the color image, the depth image, and the text data. The text data is used to describe the target object.
[0141] The generation module 904 is configured to generate normalized object coordinates of the target object based on the color image and the text data.
[0142] The estimation module 906 is configured to estimate the target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object.
[0143] The output module 908 is configured to output the estimation result by calling a second interface. The second interface includes a second parameter, and the parameter value of the second parameter includes the estimation result.
[0144] It should be noted here that the above acquisition module 902, generation module 904, estimation module 906, and output module 908 correspond to steps S502 to S508 in Embodiment 3. The examples and application scenarios implemented by the module and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above module or unit can be a hardware component or a software component stored in a memory and processed by one or more processors. The above module can also be a part of the apparatus and can run in the server 10 provided in Embodiment 1.
[0145] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 3, but are not limited to the schemes provided in Embodiment 3.
[0146] Embodiment 8
[0147] According to an embodiment of the present application, there is also provided another apparatus for implementing the above robot control method. Figure 10 is a schematic diagram of a robot control apparatus according to Embodiment 8 of the present application, as Figure 10 shown, the apparatus includes: an acquisition module 1002, a generation module 1004, an estimation module 1006, and a control module 1008.
[0148] Among them, the acquisition module 1002 is used to acquire the color image, depth image and text data of the object to be grasped, where the text data is used to describe the object to be grasped. The generation module 1004 is used to generate the normalized object coordinates of the object to be grasped based on the color image and the text data. The estimation module 1006 is used to estimate the object to be grasped based on the normalized object coordinates and the depth image, and obtain the estimation result of the object to be grasped, where the estimation result is used to characterize the size and pose of the object to be grasped. The control module 1008 is used to control the robot to grasp the object to be grasped based on the estimation result.
[0149] It should be noted here that the above acquisition module 1002, generation module 1004, estimation module 1006, and control module 1008 correspond to steps S602 to S608 in Embodiment 4. The instances and application scenarios implemented by the module and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above module or unit can be a hardware component or software component stored in the memory and processed by one or more processors, and the above module can also be a part of the device and can run in the server 10 provided in Embodiment 1.
[0150] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 4, but are not limited to the schemes provided in Embodiment 4.
[0151] Embodiment 9
[0152] An embodiment of the present application can provide a computer terminal, and the computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal can also be replaced with a terminal device such as a mobile terminal.
[0153] Optionally, in this embodiment, the above computer terminal can be located in at least one network device among multiple network devices of a computer network.
[0154] In this embodiment, the above computer terminal can execute the program code of the following steps in the object estimation method: acquire the color image, depth image and text data of the target object, where the text data is used to describe the target object; generate the normalized object coordinates of the target object based on the color image and the text data; estimate the target object based on the normalized object coordinates and the depth image, and obtain the estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object.
[0155] Optionally, Figure 11 is a structural block diagram of a computer terminal according to an embodiment of the present application. As Figure 11As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 1102, a memory 1104, a storage controller, and a peripheral interface. The peripheral interface is connected to a radio frequency module, an audio module, and a display.
[0156] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the object estimation method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above object estimation method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal A through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0157] The processor can call the information and application programs stored in the memory through a transmission device to perform the following steps: obtain a color image, a depth image, and text data of a target object, where the text data is used to describe the target object; generate a normalized object coordinate of the target object based on the color image and the text data; estimate the target object based on the normalized object coordinate and the depth image to obtain an estimation result of the target object.
[0158] Optionally, the above processor may further execute the program code of the following steps: perform feature fusion on the color image and the text data to obtain a fusion feature, where the fusion feature carries semantic information of the target object; use a first image feature extraction module to extract features from the color image to obtain a first image feature, where the first image feature is used to characterize different parts of the target object; generate a normalized object coordinate based on the fusion feature and the first image feature.
[0159] Optionally, the above processor may further execute the program code of the following steps: use a second image feature extraction module to extract features from the color image to obtain a second image feature; use a text feature extraction module to extract features from the text data to obtain a text feature; use a fusion module to perform feature fusion on the second image feature and the text feature to obtain a fusion feature.
[0160] Optionally, the above processor may further execute the program code of the following steps: splice the fusion feature and the first image feature to obtain a spliced feature; use a decoder to decode the spliced feature to obtain a normalized object coordinate.
[0161] Optionally, the above-mentioned processor may also execute the program code of the following steps: perform bilinear interpolation on the first image feature to obtain an interpolated feature, where the size of the interpolated feature is the same as that of the fusion feature; splice the fusion feature and the interpolated feature to obtain a spliced feature.
[0162] Optionally, the above-mentioned processor may also execute the program code of the following steps: during the training of the fusion module and the decoder, control the network parameters of the first image feature extraction module, the second image feature extraction module, and the text feature extraction module to remain unchanged.
[0163] Optionally, the above-mentioned processor may also execute the program code of the following steps: convert the depth image into point cloud data of the target object; determine the corresponding positions of the points of the normalized object coordinates in the point cloud data; perform estimation based on the point cloud data and the corresponding positions to obtain an estimation result.
[0164] Optionally, the above-mentioned processor may also execute the program code of the following steps: select a first point from the points of the normalized object coordinates; perform estimation based on the first point in the point cloud data and the first corresponding position of the first point in the point cloud data to obtain an initial estimation result; select a second point from the remaining points of the normalized object coordinates based on the initial estimation result, where the remaining points are used to represent the points other than the first point in the normalized object coordinates; add the second point to the first point and repeat the steps of obtaining the initial estimation result and selecting the second point until a preset condition is met; determine the initial estimation result as the estimation result.
[0165] Optionally, the above-mentioned processor may also execute the program code of the following steps: determine the number of second points; in the case where the number is greater than or equal to the preset number, add the second point to the first point.
[0166] Optionally, the above-mentioned processor may also execute the program code of the following steps: reselect a first point from the points of the normalized object coordinates and repeat the steps of obtaining the initial estimation result and selecting the second point until a preset condition is met; determine the initial estimation result as the estimation result.
[0167] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtain an original image containing the target object; perform object detection on the original image to obtain a mask image of the target object; perform cropping on the original image based on the mask image to obtain a color image.
[0168] In the above embodiments of the present application, obtaining the text data of the target object includes: obtaining an original image containing the target object; performing object classification on the original image to obtain the target category of the target object; obtaining the text data corresponding to the target category from the text database.
[0169] In an embodiment of the present application, a color image, a depth image, and text data of a target object are obtained, where the text data is used to describe the target object; based on the color image and the text data, normalized object coordinates of the target object are generated; and based on the normalized object coordinates and the depth image, the target object is estimated to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object. It is easy to notice that the normalized object coordinates of the target object can be generated based on the color image and the text data of the target object, and the pose and size of the target object can be estimated by using the normalized object coordinates and the depth image to obtain the estimation result of the target object. That is, in the process of estimating the pose and size of the target object, the pose and size of the target object are determined based on the normalized object coordinates, so that not only can the same type of objects share a normalized reference coordinate system, but also the normalized object coordinates are generated based on the color image of the target object and the text data describing the target object. That is, when generating the normalized object coordinates, not only the color image of the target object is considered, but also the description text of the target object is considered. The text data can be used to describe different types of target objects and generate normalized object coordinates, thereby realizing the estimation of the pose and size across different object categories, improving the estimation efficiency when estimating the pose and size of the target object, and further solving the technical problem of poor generality of the method for estimating the pose and size of an object.
[0170] Those of ordinary skill in the art can understand that the structure shown in the figure is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices. Figure 11 It does not limit the structure of the above electronic device. For example, the computer terminal A may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 11 the figure, or have a different configuration from that shown in Figure 11 the figure.
[0171] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0172] Embodiment 10
[0173] Embodiments of the present application also provide a storage medium. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the object estimation method provided in the first embodiment above.
[0174] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0175] Optionally, in this embodiment, the storage medium is set to store program code for performing the following steps: obtaining a color image, a depth image, and text data of a target object, where the text data is used to describe the target object; generating normalized object coordinates of the target object based on the color image and the text data; estimating the target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object.
[0176] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0177] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0178] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.
[0179] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0182] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. An object estimation method, characterized in that, Including: Obtain a color image, a depth image, and text data of a target object, where the text data is used to describe the target object; Generate normalized object coordinates of the target object based on the color image and the text data; Estimate the target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object, where the estimation result is used to characterize the size and pose of the target object.
2. The method according to claim 1, wherein Generating the normalized object coordinates of the target object based on the color image and the text data includes: Perform feature fusion on the color image and the text data to obtain a fused feature, where semantic information of the target object is carried in the fused feature; Use a first image feature extraction module to extract features from the color image to obtain a first image feature, where the first image feature is used to characterize different parts of the target object; Generate the normalized object coordinates based on the fused feature and the first image feature.
3. The method according to claim 2, characterized in that, Performing feature fusion on the color image and the text data to obtain a fused feature includes: Use a second image feature extraction module to extract features from the color image to obtain a second image feature; Use a text feature extraction module to extract features from the text data to obtain a text feature; Use a fusion module to perform feature fusion on the second image feature and the text feature to obtain the fused feature.
4. The method according to claim 2, wherein Generating the normalized object coordinates based on the fused feature and the first image feature includes: Concatenate the fused feature and the first image feature to obtain a concatenated feature; Use a decoder to decode the concatenated feature to obtain the normalized object coordinates.
5. The method according to claim 4, wherein Concatenating the fused feature and the first image feature to obtain a concatenated feature includes: Perform bilinear interpolation on the first image feature to obtain an interpolated feature, where the size of the interpolated feature is the same as that of the fused feature; Concatenate the fused feature and the interpolated feature to obtain the concatenated feature.
6. The method according to claim 3 or 4, characterized in that, During the training process of the fusion module and the decoder, control the network parameters of the first image feature extraction module, the second image feature extraction module, and the text feature extraction module to remain unchanged.
7. The method according to claim 1, characterized in that, Estimating the target object based on the normalized object coordinates and the depth image to obtain an estimation result of the target object includes: Convert the depth image into point cloud data of the target object; Determine the corresponding positions of the points of the normalized object coordinates in the point cloud data; Perform an estimation based on the point cloud data and the corresponding positions to obtain the estimation result.
8. The method according to claim 7, wherein Performing an estimation based on the point cloud data and the corresponding positions to obtain the estimation result includes: Select a first point from the points of the normalized object coordinates; Perform an estimation based on the first point in the point cloud data and the first corresponding position of the first point in the point cloud data to obtain an initial estimation result; Select a second point from the remaining points of the normalized object coordinates based on the initial estimation result, where the remaining points are used to represent the points in the normalized object coordinates other than the first point; Add the second point to the first point, and repeat the steps of obtaining the initial estimation result and selecting the second point until a preset condition is met; Determine the initial estimation result as the estimation result.
9. The method according to claim 8, wherein Adding the second point to the first point includes: Determine the number of the second points; When the number is greater than or equal to a preset number, add the second point to the first point.
10. The method according to claim 9, characterized in that, When the number is less than the preset number, the method further includes: Re-select the first point from the points of the normalized object coordinates, and repeat the steps of obtaining the initial estimation result and selecting the second point until the preset condition is met; Determine the initial estimation result as the estimation result.
11. The method according to claim 1, characterized in that, Obtain the text data of the target object, including: Obtain the original image containing the target object; Perform target classification on the original image to obtain the target category of the target object; Obtain the text data corresponding to the target category from the text database.
12. An object estimation method, characterized in that, Include: In response to an input instruction acting on the operation interface, display a color image, a depth image, and text data of the target object on the operation interface, where the text data is used to describe the target object; In response to an estimation instruction acting on the operation interface, display the estimation result of the target object on the operation interface, where the estimation result is the result obtained by estimating the target object based on the normalized object coordinates and the depth image of the target object, and the normalized object coordinates are generated based on the color image and the text data.
13. An object estimation method, characterized in that, Include: Obtain a color image, a depth image, and text data of the target object by calling a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the color image, the depth image, and the text data, and the text data is used to describe the target object; Generate the normalized object coordinates of the target object based on the color image and the text data; Estimate the target object based on the normalized object coordinates and the depth image to obtain the estimation result of the target object; Output the estimation result by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the estimation result.
14. A robot control method, characterized in that, Include: Obtain a color image, a depth image, and text data of the object to be grasped, where the text data is used to describe the object to be grasped; Generate the normalized object coordinates of the object to be grasped based on the color image and the text data; Estimate the object to be grasped based on the normalized object coordinates and the depth image to obtain the estimation result of the object to be grasped, where the estimation result is used to represent the size and pose of the object to be grasped; Control the robot to grasp the object to be grasped based on the estimation result.
15. An electronic device, characterized in that, Include: A memory that stores an executable program; A processor for running the program, wherein when the program runs, it executes the method according to any one of claims 1 to 14.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 14.
17. A computer program product, characterized in that, It includes a computer program that implements the method according to any one of claims 1 to 14 when executed by a processor.