Pose determination method and apparatus based on large language model, device, storage medium, and product
By acquiring and aligning color and depth images, and using a large language model to determine the mask image of the target object and filter sub-point clouds, the problem of low accuracy in target object pose determination is solved, achieving higher accuracy in recognition and pose determination.
Patent Information
- Application Number
- PCT/CN2024/125754
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-11
- Filing Date
- 2024-10-18
- Publication Date
- 2026-04-16
AI Technical Summary
The current method for determining the pose of the target object suffers from low accuracy.
Acquire and align the color and depth images of the target region, determine the target mask image of the target object using a pre-defined large language model, filter out sub-point clouds, and determine pose information based on the sub-point clouds.
By ensuring spatial consistency of image data and utilizing large language models to identify category information, the accuracy of target object recognition and pose determination is improved.
Smart Images

Figure CN2024125754_16042026_PF_FP_ABST
Abstract
Description
Pose determination methods, devices, equipment, storage media, and products based on large language models
[0001] This application claims priority to Chinese Patent Application No. 202411416427.X, filed on October 11, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer vision technology, and for example to a pose determination method, apparatus, device, storage medium, and product based on a large language model. Background Technology
[0003] With the rapid advancement of image processing technology, especially breakthroughs in deep learning, computer vision, and 3D reconstruction, extracting the pose information of target objects from image or video data has become possible. The pose information of a target object refers to its position and orientation in three-dimensional space, which is crucial for fields such as automated robotic operations, augmented reality, autonomous driving, and intelligent logistics.
[0004] However, the current method for determining the pose of the target object suffers from low accuracy.
[0005] Summary of the Invention
[0006] This application provides a pose determination method, apparatus, device, storage medium, and product based on a large language model to solve the problem of low accuracy in current target object pose determination methods.
[0007] According to one aspect of this application, a pose determination method based on a large language model is provided, comprising:
[0008] Acquire a color image and a depth image of a target region, wherein the color image is aligned with the depth image, and the target region includes a set of target objects;
[0009] Based on the color image, using a preset large language model, a target mask image of at least one target object in the target object set is determined, wherein the target mask image has category information;
[0010] Using a target mask image of at least one target object, a sub-point cloud of at least one target object is selected from an original point cloud, wherein the original point cloud is determined based on the depth image;
[0011] Based on the sub-point cloud of the at least one target object, determine the pose information of the at least one target object.
[0012] According to another aspect of this application, a pose determination device based on a large language model is provided, comprising:
[0013] An image acquisition module is used to acquire a color image and a depth image of a target region, wherein the color image is aligned with the depth image, and the target region includes a set of target objects;
[0014] A mask image determination module is used to determine, based on the color image and using a preset large language model, a target mask image of at least one target object in the target object set, wherein the target mask image has category information;
[0015] A sub-point cloud determination module is used to filter at least one sub-point cloud of a target object from an original point cloud using a target mask image of at least one target object, wherein the original point cloud is determined based on the depth image;
[0016] The pose determination module is used to determine the pose information of at least one target object based on the sub-point cloud of the at least one target object.
[0017] According to another aspect of this application, an electronic device is provided, the electronic device comprising:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the pose determination method described in any embodiment of this application.
[0021] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the pose determination method described in any embodiment of this application.
[0022] According to another aspect of this application, a computer program product is provided, the computer program product including a computer program that, when executed by a processor, implements the pose determination method described in any embodiment of this application.
[0023] The technical solution provided in this application embodiment acquires a color image and a depth image of a target region, wherein the color image and the depth image are aligned, and the target region includes a set of target objects; based on the color image, a preset large language model is used to determine a target mask image of at least one target object in the target object set, wherein the target mask image has category information; using the target mask image of at least one target object, at least one sub-point cloud of the target object is selected from the original point cloud, wherein the original point cloud is determined based on the depth image; based on the sub-point cloud of the at least one target object, the pose information of at least one target object is determined. By acquiring and aligning the color image and the depth image of the target region, spatial consistency of the image data is ensured; based on the color image, the category information of the target object is identified using a preset large language model, and at least one target mask image with category information is generated, effectively improving the accuracy of target object recognition; using the target mask image of at least one target object, at least one sub-point cloud of the target object is selected from the original point cloud, and the pose information of at least one target object is determined based on these sub-point clouds, effectively improving the accuracy of target object pose determination. Attached Figure Description
[0024] Figure 1 is a flowchart of a pose determination method based on a large language model provided in Embodiment 1 of this application;
[0025] Figure 2 is a flowchart of a pose determination method based on a large language model provided in Embodiment 2 of this application;
[0026] Figure 3 is a structural diagram of a pose determination method provided in an embodiment of this application;
[0027] Figure 4 is a schematic diagram of a pose determination device based on a large language model provided in Embodiment 3 of this application;
[0028] Figure 5 is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of this application. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 is a flowchart of a pose determination method provided in Embodiment 1 of this application. This embodiment is applicable to the situation of determining the pose of a target object. The method can be executed by a pose determination device, which can be implemented in hardware and / or software and can be configured in an electronic device. As shown in Figure 1, the method includes:
[0033] S110. Obtain a color image and a depth image of the target region, wherein the color image is aligned with the depth image, and the target region includes a set of target objects.
[0034] In this embodiment, the target region can be understood as a portion of a specific scene captured by the image acquisition device, which includes multiple target objects. The image acquisition device may include a Red Green Blue-Depth (RGB-D) camera, or a color camera and a depth camera. The RGB-D camera can simultaneously acquire color and depth images, and these two types of images are typically aligned. Target objects can be understood as objects whose pose needs to be determined. The set of target objects is determined by at least one target object.
[0035] A color image can be understood as an image that includes color information, and can be represented as:
[0036] Where (x,y) represents the coordinates of a pixel in the color image, and I(x,y) represents the color value of the pixel (x,y), which is a three-dimensional vector whose components correspond to the values of the red, green, and blue color channels, respectively. This represents a three-dimensional real number space, specifically a three-dimensional color space consisting of three color channels: red, green, and blue. The value range for each channel is [0, 255].
[0037] Depth images include distance information from each point in the target region to the image acquisition device, and can be represented as:
[0038] Where d(x,y) represents the depth value (floating-point type) of pixel (x,y) in the depth image. This represents the set of real numbers, where the depth value can be any real number.
[0039] Specifically, to accurately identify the pose of a target object, it is necessary to acquire color and depth images of the target region and align the two types of images to ensure that each pixel in the color image matches its corresponding pixel in the depth image in spatial location. It is worth noting that if an RGB-D camera is used to acquire both types of images, since the color and depth images acquired by this camera are aligned, no further alignment processing is required. However, if separate color and depth cameras are used, the acquired color and depth images need to be aligned. This embodiment does not limit the specific method of alignment processing; for example, it can be achieved through a calibration process. The embodiments of this application use color and depth images acquired by an RGB-D camera as an example for illustration.
[0040] S120. Based on the color image, using a preset large language model, determine the target mask image of at least one target object in the target object set, wherein the target mask image has category information.
[0041] In this embodiment, the pre-trained large language model can be understood as a pre-trained deep learning model used to determine the category information of an image. The sample data volume and model parameters of the pre-trained large language model can be on a massive scale, for example, having billions or even hundreds of billions of parameters, trained on large-scale text and image data. The large language model can, for example, employ a Chat Generative Pre-trained Transformer (ChatGPT) network structure or a Contrastive Language–Image Pre-training (CLIP) network structure. In this embodiment, the images for classification can be mask images corresponding to each target object.
[0042] Specifically, color images can be processed using algorithms such as image segmentation, edge detection, and binarization to generate color images corresponding to each target object and mask images containing the mask regions corresponding to each target object. Then, the color image corresponding to each target object is input into a preset large language model. The preset large language model will combine the input prompt words and output the category information corresponding to each target object. Then, based on the mask regions and corresponding category information of each target object in the mask image, the target mask image of each target object can be obtained.
[0043] The cue words provide the pre-defined large language model with additional information about the image content, guiding it to correctly understand and classify the masked image. For example, the cue words can be descriptions of the target object's attributes or uses, such as "red cup" and "tool for writing," or specific instructions, such as "Please output the object category name in this image."
[0044] S130. Using a target mask image of at least one target object, filter at least one sub-point cloud of the target object from the original point cloud, wherein the original point cloud is determined based on the depth image.
[0045] In this embodiment, the original point cloud can be understood as a set of three-dimensional points containing depth and color information. The sub-point cloud is a part of the original point cloud.
[0046] Specifically, firstly, based on the depth image and the intrinsic parameter matrix K of the RGB-D camera, the original point cloud Pc is calculated:
[0047] Among them, f x f y These are the horizontal and vertical focal lengths of the RGB-D camera, respectively. x c y The focal length and optical center coordinates of the RGB-D camera can both be obtained through the camera software development kit (SDK).
[0048] The coordinates of each pixel in the depth image can be represented as (u i ,v i ), depth value The coordinates (x, y) of each point in the point cloud are calculated using the following method. i ,y i ,z i ):
[0049] Furthermore, the original point cloud Pc can be represented as: Pc = {P1, P2, ..., P} N}
[0050] Among them, Pi =(x i ,y i ,z i (x) represents the i-th point in the original point cloud. i ,y i ,z i ) represents the coordinates of the i-th point, and N is the total number of points in the original point cloud.
[0051] Since the color image and the depth image are aligned, and the original point cloud is determined based on the depth image, it can be concluded that the target mask image and the original point cloud are aligned.
[0052] Furthermore, for each target object with at least one target object, the sub-point cloud of the current target object can be selected from the original point cloud in the following ways: First, feature extraction is performed on the target mask image, and preprocessing (such as filtering, downsampling, etc.) and feature extraction (such as surface normals, curvature, etc.) are performed on the original point cloud. Then, the image features are matched with the point cloud features using a feature matching algorithm, and the sub-point cloud of the target object is selected based on the matching results. Alternatively, the points in the point cloud can be projected onto the same two-dimensional plane as the target mask image of the current target object to obtain a set of projection points. From the set of projection points, projection points with the same pixel position as the target object in the target mask image of the current target object are found. The set of points in the original point cloud corresponding to these projection points is determined as the sub-point cloud of the current target object.
[0053] S140. Based on the sub-point cloud of the at least one target object, determine the pose information of the at least one target object.
[0054] In this embodiment, the pose information includes the position and orientation information of the target object in three-dimensional space.
[0055] Specifically, for each target object with at least one target object, the pose information of the current target object can be determined based on the sub-point cloud of the current target object using a preset pose algorithm. The preset pose algorithm can be a deep learning algorithm, such as using a neural network model to extract features from the sub-point cloud and then estimating the pose based on the extracted features; the preset pose algorithm can also be a point cloud registration algorithm, such as the Iterative Closest Point (IPC) algorithm, etc. This embodiment does not limit the preset pose algorithm.
[0056] The technical solution provided in Embodiment 1 of this application acquires a color image and a depth image of a target region, wherein the color image and the depth image are aligned, and the target region includes a set of target objects; based on the color image, a preset large language model is used to determine a target mask image of at least one target object in the target object set, wherein the target mask image has category information; using the target mask image of at least one target object, at least one sub-point cloud of the target object is selected from the original point cloud, wherein the original point cloud is determined based on the depth image; based on the sub-point cloud of the at least one target object, the pose information of at least one target object is determined. By acquiring and aligning the color image and the depth image of the target region, spatial consistency of the image data is ensured; based on the color image, the category information of the target object is identified using a preset large language model, and at least one target mask image with category information is generated, effectively improving the accuracy of target object recognition; using the target mask image of at least one target object, at least one sub-point cloud of the target object is selected from the original point cloud, and based on these sub-point clouds, the pose information of at least one target object is determined, effectively improving the accuracy of target object pose determination.
[0057] In some embodiments, the preset large language model includes a first preset large language model and a second preset large language model, wherein the first preset large language model and the second preset large language model are different; based on the color image, using the preset large language model, determining the target mask image of at least one target object in the target object set includes: segmenting the color image using a preset image segmentation model to obtain a first mask image, wherein the first mask image includes a plurality of first mask regions; clustering the first mask image using a preset clustering algorithm to obtain a second mask image, wherein the second mask image includes a second mask region corresponding to the target object; Based on the second mask image and the color image, a mask region segmentation image for each target object is determined. For each target object in the target object set, based on the mask region segmentation image of the current target object and the input prompt word, the first category information of the current target object is determined using the first preset large language model, and the second category information of the current target object is determined using the second preset large language model. When the first category information and the second category information are the same, the target mask image of the current target object is determined based on the target category information and the second mask region corresponding to the current target object, wherein the target category information is either the first category information or the second category information.
[0058] In this embodiment, both the first and second preset large language models are pre-trained deep learning models. The first and second preset large language models are not the same; for example, they may have different structures, parameters, training data, or model input data. This embodiment uses the difference in structure between the first and second preset large language models as an example for illustration.
[0059] A preset image segmentation model can be understood as a pre-trained deep learning model used for fine-grained segmentation of color images. A preset clustering algorithm is a pre-configured algorithm used to cluster mask regions.
[0060] Specifically, a preset image segmentation model is used to perform mask segmentation on the color image to obtain a first mask image including multiple first mask regions, represented as: O0={mask1,mask2,…,mask m}
[0061] Where O0 represents the first mask image, mask i Let m represent the first mask region of the i-th binarized (pixel intensity value is either non-zero or 255), and m represent the number of first mask regions. It's important to note that the first mask region can be understood as an image region obtained after fine-grained segmentation of a color image. These regions represent different details in the color image, each with unique features and labels. A single first mask region cannot represent a complete target object. For example, in an image containing a mobile phone, fine-grained segmentation identifies different parts of the phone, such as the screen, camera, and buttons. Understandably, the first mask region (mask)... i There are two corresponding pixel intensity values: the pixel intensity value inside the first mask area is 255, and the pixel intensity value of the rest (i.e. the background area) is 0.
[0062] Furthermore, a pre-defined clustering algorithm is used to cluster the first mask image, so that the first mask region and its neighboring first mask regions are stitched together to form a mask region corresponding to the target object, i.e., the second mask region, thus obtaining the second mask image, which can be represented as:
[0063] Where O1 represents the second mask image, where, This is represented as the k-th binarized (pixel intensity value is either non-zero or 255) second mask region. Here, k represents the number of second mask regions, and k is less than m. It's worth noting that the second mask regions obtained through clustering can represent the target object. It can be understood that the second mask region... There are two corresponding pixel intensity values: the pixel intensity value inside the second mask area is 255, and the pixel intensity value of the rest (i.e. the background part) is 0.
[0064] Furthermore, the mask region segmentation image of each target object is determined in the following way:
[0065] Among them, maskImgs i This represents the mask region segmentation image corresponding to the i-th target object. It's important to note that the mask region segmentation image is a color image containing a single target object. I(x,y) represents the color value of pixel (x,y) in the color image. The mask region segmentation images of each target object constitute the mask region segmentation image set: O2={maskImgs1,…,maskImgs} k}
[0066] Furthermore, the target mask image of at least one target object in the target object set can be determined in the following way:
[0067] For each target object in the target object set, the mask region segmentation image of the current target object is input into a first preset large language model. The first preset large language model combines the input prompt words to perform classification prediction and outputs the first category information cls of the current target object. i (1) Furthermore, the segmented image of the mask region of the current target object is input into the second preset large language model. The second preset large language model combines the input prompt words to perform classification prediction and outputs the second category information cls of the current target object. i (2) In this embodiment, the inputs to the first preset large language model and the second preset large language model are the same. Then, it is determined whether the first category information and the second category information are the same. If the two category information are the same, the first category information or the second category information is used as the target category information. The target category information is then used to mark the second mask region corresponding to the current target object, thereby obtaining the target mask image of the current target object.
[0068] Optionally, if the two category information are different, the second mask region may not be marked.
[0069] Therefore, the target mask images can be filtered out to obtain a set of target mask images containing target category information:
[0070] in, This represents the l-th target mask image. l represents the number of target mask images in the set, where l is less than or equal to k. cls l This represents the target category information of the l-th target mask image.
[0071] By employing the above technical solutions, which comprehensively utilize image segmentation techniques, clustering algorithms, and two pre-defined large language models, accurate identification and classification of target objects in color images are achieved. This effectively improves the accuracy, reliability, and flexibility of the identification process, providing strong support for pose determination.
[0072] In some embodiments, the pose determination method further includes: when the first category information and the second category information are different, determining the target mask image of the current target object using a preset identifier and the second mask region corresponding to the current target object.
[0073] In this embodiment, the preset identifier is a pre-set marker used to indicate that the classification of the second mask area is abnormal.
[0074] Specifically, when the first category information and the second category information are different, a preset identifier can be used to mark the second mask region corresponding to the current target object to determine the target mask image of the current target object. Furthermore, when filtering target mask images, only target mask images containing target category information can be selected, resulting in a set of target mask images containing target category information. This technical solution lays the foundation for subsequently filtering target mask images containing target category information, further improving the accuracy of pose determination.
[0075] In some embodiments, the preset image segmentation model includes a segmentation model of everything (SAM), and / or the preset clustering algorithm includes a region growing algorithm. The above technical solutions can effectively improve the efficiency of color image segmentation.
[0076] In this embodiment, the core of the Segment Anything Model (SAM) lies in its ability to guide the model to segment specific objects in an image through prompts, such as clicks, bounding boxes, or text descriptions. This model exhibits excellent zero-shot learning capabilities, multimodal input processing capabilities, and real-time interactive segmentation performance, demonstrating high flexibility and efficiency. Region growing is an image segmentation method based on pixel similarity. It starts from a pre-selected seed point and gradually merges similar pixels into the region containing the seed point until a termination condition is met. The advantage of this algorithm is its simplicity and intuitiveness, its ability to preserve the natural boundaries of regions, and its requirement for prior knowledge of the number or shape of the targets.
[0077] Example 2
[0078] Figure 2 is a flowchart of a pose determination method provided in Embodiment 2 of this application. This embodiment is an optimization and extension based on the above optional embodiments.
[0079] Optionally, using a target mask image of at least one target object, a sub-point cloud of at least one target object is selected from the original point cloud. This includes: for each target object, projecting points in the original point cloud onto the same two-dimensional plane as the target mask image of the current target object to obtain a set of projection points; identifying projection points in the projection point set whose pixel positions are the same as the pixel positions of the target values in the target mask image of the current target object, and using these as target projection points to form a set of target projection points; and determining the set of points in the original point cloud corresponding to the set of target projection points as the sub-point cloud of the current target object. Through this technical solution, the sub-point cloud corresponding to the target object can be accurately identified, effectively improving the efficiency and accuracy of point cloud processing, and providing accurate three-dimensional spatial information for determining the pose of the target object.
[0080] Optionally, determining the pose information of at least one target object based on the sub-point cloud of the at least one target object includes: for each target object in the at least one target object, calculating the centroid and feature vector of the sub-point cloud corresponding to the current target object, and determining the pose information of the current target object based on the centroid and feature vector of the sub-point cloud. Through the above technical solution, the position and orientation of the target object in three-dimensional space can be accurately determined, thereby achieving precise estimation of the target object's pose.
[0081] As shown in Figure 2, the method includes:
[0082] S210. Obtain a color image and a depth image of the target region, wherein the color image is aligned with the depth image, and the target region includes a set of target objects.
[0083] S220. Based on the color image, using a preset large language model, determine the target mask image of at least one target object in the target object set, wherein the target mask image has category information.
[0084] S230. For each target object in at least one target object, project the points in the point cloud onto the same two-dimensional plane as the target mask image of the current target object to obtain a set of projection points. Determine the projection points in the set of projection points that are at the same position as the pixel points in the target mask image of the current target object with the pixel value of the target value, and take them as target projection points to form a set of target projection points. Determine the set of points in the original point cloud corresponding to the set of target projection points as the sub-point cloud of the current target object, wherein the original point cloud is determined based on the depth image.
[0085] The target value is the pixel value of a pixel in the target mask image. In this embodiment, since the mask image is a binary image, the pixel value corresponding to the background is 0, while the pixel value in the region corresponding to the target object is 1. Therefore, the target value is 1.
[0086] Specifically, first, the original point cloud is determined based on the depth image and camera parameters; then, for each target object in at least one target object, the sub-point cloud of the current target object can be determined in the following way:
[0087] The points in the point cloud are projected onto the same two-dimensional plane as the target mask image of the current target object, resulting in a set of projection points. At this point, the points in the projection point set correspond one-to-one with the pixels in the target mask image. Then, the projection points in the projection point set that are at the same position as the pixel with a pixel value of 1 in the target mask image of the current target object are found and used as target projection points. These target projection points form a set of target projection points. Finally, the set of points in the original point cloud corresponding to the set of target projection points is determined as the sub-point cloud of the current target object.
[0088] S240. For each target object in at least one target object, calculate the centroid and feature vector of the sub-point cloud corresponding to the current target object, and determine the pose information of the current target object based on the centroid and feature vector of the sub-point cloud.
[0089] In this embodiment, the centroid of the sub-point cloud represents the average position of the sub-point cloud, providing translation information. The eigenvectors of the sub-point cloud are used to represent its principal orientation. Specifically, the principal orientation of the sub-point cloud represents the main trend direction of its distribution in three-dimensional space.
[0090] Specifically, for each target object in at least one target object, the pose information of the current target object can be determined in the following ways:
[0091] First, calculate the centroid of the sub-point cloud of the current target object.
[0092] in, This indicates the category cls of the sub-point cloud. l The number of points in (x i ,y i ,z i ) represents the coordinates of the i-th point in the sub-point cloud.
[0093] Then, the feature vector of the sub-point cloud of the current target object is calculated.
[0094] The rotation matrix of the sub-point cloud is obtained by calculating the covariance matrix and decomposing the eigenvectors.
[0095] in This represents the first principal axis (X direction) direction component vector of the sub-point cloud principal direction; This represents the direction component vector of the second principal axis (Y direction) of the main direction of the sub-point cloud. This represents the component vector of the third principal axis (Z direction) of the main direction of the sub-point cloud, thus forming a 3×3 dimension matrix.
[0096] Furthermore, based on the centroid and feature vector of the sub-point cloud corresponding to the current target object, the pose information of the current target object is determined. As shown below:
[0097] The technical solution provided in Embodiment 2 of this application, by comprehensively utilizing the alignment technology of color images and depth images, target mask recognition of preset large language models, accurate projection screening of point clouds, and calculation of the centroid and feature vectors of sub-point clouds, achieves accurate determination of the pose of the target object in three-dimensional space, significantly improving the accuracy and efficiency of pose recognition.
[0098] For example, Figure 3 is a structural diagram of pose determination provided by an embodiment of this application. As shown in Figure 3: the SAM model is used to perform mask segmentation on the color image RGB to obtain a first mask image O0 including multiple first mask regions. The region growing algorithm is used to cluster the first mask image O0 to obtain a second mask image O1. Then, multiple mask region segmentation images are obtained based on the second mask image O1 and the color image RGB segmentation, forming a mask region segmentation image set O2. Each mask region segmentation image in the mask region segmentation image set O2 is combined with the corresponding prompt word Q: "Please output the category name of the object in this image", and input into the first preset large language model LLM_model1 and the second preset large language model LLM_model2 respectively to obtain the first category information and the second category information. It is further determined whether the two category information are the same. If they are the same, the first category information or the second category information can be determined as the target. Category information, such as the first category information cls4 and the second category information cls4 in the figure, can determine the target category information as cls4. Then, the target category information is used to mark the second mask region in the second mask image O1 to obtain the target mask image. Based on the target mask image, a target mask image set O3 is obtained. The original point cloud (not shown in the figure) is determined based on the depth image (Depth) and the color image (RGB). Based on the original point cloud and the target mask image set O3, the sub-point cloud (not shown in the figure) corresponding to each target object is determined. Then, the pose information (x, y, z, rx, ry, rz) of each target object can be obtained from the sub-point cloud. Here, (x, y, z) represents the position of the center of the bounding box corresponding to the target object in the world coordinate system, and (rx, ry, rz) represents the rotation angle of the target object around the x, y, and z axes of the world coordinate system (not shown in the figure). Each target object is projected onto the same two-dimensional image. The pose information of each target object in the two-dimensional image is represented by the border and the central axis.
[0099] Example 3
[0100] Figure 4 is a schematic diagram of a pose determination device provided in Embodiment 3 of this application. As shown in Figure 4, the device includes:
[0101] Image acquisition module 31 is used to acquire a color image and a depth image of a target region, wherein the color image is aligned with the depth image, and the target region includes a set of target objects;
[0102] The mask image determination module 32 is used to determine the target mask image of at least one target object in the target object set based on the color image and using a preset large language model, wherein the target mask image has category information;
[0103] The sub-point cloud determination module 33 is used to filter at least one sub-point cloud of a target object from the original point cloud using a target mask image of at least one target object, wherein the original point cloud is determined based on the depth image.
[0104] The pose determination module 34 is used to determine the pose information of at least one target object based on the sub-point cloud of the at least one target object.
[0105] The technical solution provided in Embodiment 3 of this application effectively improves the accuracy of target object pose determination.
[0106] Optionally, the preset large language model includes a first preset large language model and a second preset large language model, wherein the first preset large language model and the second preset large language model are different;
[0107] Optionally, the mask image determination module 32 includes:
[0108] The first image determination unit is used to segment the color image using a preset image segmentation model to obtain a first mask image, wherein the first mask image includes a plurality of first mask regions;
[0109] The second image determination unit is used to cluster the first mask image using a preset clustering algorithm to obtain a second mask image, wherein the second mask image includes a second mask region corresponding to the target object;
[0110] An initial image determination unit is used to determine a mask region segmentation image for each target object based on the second mask image and the color image;
[0111] The target image determination unit is configured to, for each target object in the target object set, determine the first category information of the current target object based on the mask region segmentation image of the current target object and the input prompt word, using the first preset large language model, and determine the second category information of the current target object based on the second preset large language model. When the first category information and the second category information are the same, the unit determines the target mask image of the current target object based on the target category information and the second mask region corresponding to the current target object, wherein the target category information is either the first category information or the second category information.
[0112] Optionally, the sub-point cloud determination module 33 is specifically used to, for each of at least one target object, project the points in the original point cloud onto a two-dimensional plane that is the same as the target mask image of the current target object to obtain a set of projection points, determine the projection points in the set of projection points that are at the same position as the pixel points with the pixel value of the target value in the target mask image of the current target object as target projection points, form a set of target projection points, and determine the set of points in the original point cloud corresponding to the set of target projection points as the sub-point cloud of the current target object.
[0113] Optionally, the pose determination module 34 is specifically used to calculate the centroid and feature vector of the sub-point cloud corresponding to the current target object for each target object in at least one target object, and determine the pose information of the current target object based on the centroid and feature vector of the sub-point cloud.
[0114] Optionally, when the first category information and the second category information are different, the target mask image of the current target object is determined by using a preset identifier and the second mask region corresponding to the current target object.
[0115] Optionally, the preset image segmentation model includes the Segmentation All Model (SAM), and / or the preset clustering algorithm includes a region growing algorithm.
[0116] The pose determination device provided in this application embodiment can execute the pose determination method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method execution.
[0117] Example 4
[0118] Figure 5 is a schematic diagram of an electronic device provided in Embodiment 4 of this application. This electronic device can represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0119] As shown in Figure 5, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor 11. The processor 11 can perform various appropriate actions and processes based on the computer programs stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0120] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0121] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as pose determination methods.
[0122] In some embodiments, the pose determination method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the pose determination method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the pose determination method by any other suitable means (e.g., by means of firmware).
[0123] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0125] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. The storage medium can be a non-transitory storage medium.
[0126] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0127] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0128] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0129] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0130] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
[0131] This application also provides a computer program product, including a computer program and / or instructions, which, when executed by a processor, implements the pose determination method as provided in any embodiment of this application.
[0132] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
Claims
1. A pose determination method based on a large language model, comprising: Acquire a color image and a depth image of a target region, wherein the color image is aligned with the depth image, and the target region includes a set of target objects; Based on the color image, using a preset large language model, a target mask image of at least one target object in the target object set is determined, wherein the target mask image has category information; Using a target mask image of at least one target object, a sub-point cloud of at least one target object is selected from an original point cloud, wherein the original point cloud is determined based on the depth image; Based on the sub-point cloud of the at least one target object, determine the pose information of the at least one target object.
2. The method according to claim 1, wherein, The preset large language model includes a first preset large language model and a second preset large language model, and the first preset large language model is different from the second preset large language model. Based on the color image, using a preset large language model, determine the target mask image of at least one target object in the target object set, including: The color image is segmented using a preset image segmentation model to obtain a first mask image, wherein the first mask image includes multiple first mask regions; The first mask image is clustered using a preset clustering algorithm to obtain a second mask image, wherein the second mask image includes a second mask region corresponding to the target object; Based on the second mask image and the color image, a mask region segmentation image for each target object is determined; For each target object in the target object set, based on the mask region segmentation image of the current target object and the input prompt word, the first category information of the current target object is determined using the first preset large language model, and the second category information of the current target object is determined using the second preset large language model. When the first category information and the second category information are the same, the target mask image of the current target object is determined based on the target category information and the second mask region corresponding to the current target object, wherein the target category information is either the first category information or the second category information.
3. The method according to claim 1, wherein, Using a target mask image of at least one target object, filter a sub-point cloud of at least one target object from the original point cloud, including: For each of the at least one target objects, the points in the original point cloud are projected onto the same two-dimensional plane as the target mask image of the current target object to obtain a set of projection points. The projection points in the set of projection points that are at the same position as the pixels with the target value in the target mask image of the current target object are identified as target projection points, forming a set of target projection points. The set of points in the original point cloud corresponding to the set of target projection points is identified as the sub-point cloud of the current target object.
4. The method according to claim 1, wherein, Based on the sub-point cloud of the at least one target object, determine the pose information of the at least one target object, including: For each target object in at least one target object, calculate the centroid and feature vector of the sub-point cloud corresponding to the current target object, and determine the pose information of the current target object based on the centroid and feature vector of the sub-point cloud.
5. The method according to claim 2, further comprising: When the first category information and the second category information are different, the target mask image of the current target object is determined by using the preset identifier and the second mask area corresponding to the current target object.
6. The method according to claim 2, wherein, The preset image segmentation model includes the Segmentation Model All (SAM), and / or the preset clustering algorithm includes the Region Growing Algorithm.
7. A pose determination device based on a large language model, comprising: The image acquisition module is configured to acquire a color image and a depth image of a target region, wherein the color image is aligned with the depth image, and the target region includes a set of target objects; The mask image determination module is configured to determine the target mask image of at least one target object in the target object set based on the color image and using a preset large language model, wherein the target mask image has category information; The sub-point cloud determination module is configured to filter at least one sub-point cloud of a target object from the original point cloud using a target mask image of at least one target object, wherein the original point cloud is determined based on the depth image. The pose determination module is configured to determine the pose information of at least one target object based on the sub-point cloud of the at least one target object.
8. An electronic device, wherein, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the pose determination method based on a large language model as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by the processor, the program implements the pose determination method based on a large language model as described in any one of claims 1-6.
10. A computer program product, wherein, The computer program product includes a computer program that, when executed by a processor, implements the pose determination method based on a large language model according to any one of claims 1-6.
Citation Information
Patent Citations
Language description-based class-level target object 6D pose acquisition method and storage medium
CN114399515A
Target object space positioning method and system, storage medium and computer equipment
CN114897974A
Object-conditioned plane segmentation
GB2628342A
Cited By
Human pose estimation methods, devices, terminals, and media based on multimodal perception
CN122290218A