Pipeline pose estimation method, device, equipment, medium and program product
By generating observation point clouds, dividing semantic primitives, extracting shape feature vectors and performing feature matching methods, the problem of difficult to estimate complex spatial pipeline poses in the prior art is solved, and efficient pose estimation and automatic assembly of class-level pipelines are realized.
Patent Information
- Application Number
- CN202510090954.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art lacks the pose estimation technology of class-level pipelines, making it difficult to effectively estimate the pose and shape of complex spatial pipelines, especially in the absence of high-precision prior models.
By generating the observed point cloud of the target pipeline in the target image, semantic primitives are divided according to the semantic label classification information in the point cloud, shape feature vectors are extracted to generate geometric models, and feature matching is performed to obtain pose parameters.
The position estimation of category-level pipelines is realized, which meets the automated assembly needs of complex space pipelines in aerospace products, and improves the efficiency and accuracy of robot assembly.
Smart Images

Figure CN120014026A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a pipeline posture estimation method, device, equipment, medium and program product. Background Art
[0002] In the field of aerospace, pipelines are an important part of the whole machine pressure system, power system, cooling system and control system. They are important carriers of medium and energy transmission in the system. The assembly quality of the pipeline system will directly affect the reliability and life of aerospace products. At present, the assembly of pipeline systems is mainly manual, which has the problems of cumbersome manual operation, low efficiency and easy errors. Robotic assembly is an effective means to achieve automated, high-precision and high-efficiency pipeline system assembly, and pipeline posture estimation is an important prerequisite for robotic assembly.
[0003] Existing instance-level object pose estimation methods require the provision of high-fidelity and high-precision CAD (Computer Aided Design) models. However, the pipelines in the pipeline system are all produced in small batches of single pieces. It is too expensive to obtain high-fidelity CAD models for all pipelines. In addition, there may be thousands of pipelines in a system, and each one is different. It is difficult to obtain CAD models for all of them. In the absence of a high-precision prior model, existing pose estimation methods cannot effectively estimate pose or shape. In addition, pipelines of the same category in a pipeline system may have large geometric differences, such as differences in bending angles, branch shapes, diameters, and lengths. Existing pose estimation methods usually assume that objects of the same category have similar shapes, and it is difficult to generalize to objects with large shape differences. In summary, faced with complex spatial pipelines with many models and large shape differences within the same category, the existing technology lacks category-level pipeline pose estimation technology. Summary of the invention
[0004] The purpose of the technical solution of the present invention is to provide a pipeline posture estimation method, device, equipment, medium and program product, which are used to solve the problem of lack of category-level pipeline posture estimation technology in the prior art.
[0005] To achieve the above object, the present invention is achieved by:
[0006] In a first aspect, an embodiment of the present invention provides a pipeline posture estimation method, comprising:
[0007] Generate an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image;
[0008] Dividing the observation point cloud into at least one semantic primitive according to semantic label classification information of each point in the observation point cloud, wherein the semantic label classification information is used to indicate the semantic primitive category to which the point belongs;
[0009] generating a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives;
[0010] The geometric model is feature matched with at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline.
[0011] Optionally, the pipeline pose estimation method, wherein generating an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image, comprises:
[0012] Obtaining a segmentation mask of a target pipeline in the target image according to a target detection network, a target image, and a depth image corresponding to the target image;
[0013] According to the depth image and the camera internal parameters, the pixel points in the segmentation mask of the target pipeline are mapped to the three-dimensional space to generate the observation point cloud of the target pipeline.
[0014] Optionally, in the pipeline pose estimation method, before dividing the observation point cloud into at least one semantic primitive according to the semantic label classification information of each point in the observation point cloud, the method further comprises:
[0015] A semantic label classification is performed on each point in the observation point cloud through a three-dimensional image convolutional network to obtain a semantic label for each point in the observation point cloud.
[0016] Optionally, the pipeline pose estimation method, wherein generating a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives, comprises:
[0017] For each of the semantic primitives, extract the geometric features of the semantic primitive to obtain a shape feature vector corresponding to the semantic primitive;
[0018] A geometric model of the target pipeline is generated according to a geometric generation network and a shape feature vector corresponding to at least one of the semantic primitives.
[0019] Optionally, the pipeline pose estimation method, wherein the geometric features of the semantic primitives are extracted to obtain the shape feature vectors corresponding to the semantic primitives, comprises:
[0020] Centralizing the semantic primitive to obtain a centralized point cloud of the semantic primitive after processing;
[0021] According to the centralized point cloud of the semantic primitive, the geometric features of the semantic primitive are extracted to obtain a shape feature vector corresponding to the semantic primitive.
[0022] Optionally, in the pipeline pose estimation method, the geometric generation network includes a coarse-grained generator and a fine-grained generator, the coarse-grained generator is used to generate the geometric shape of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives, and the fine-grained generator is used to refine the geometric information based on the geometric shape and image of the target pipeline to generate a geometric model of the target pipeline.
[0023] Optionally, the pipeline pose estimation method, wherein the geometric model is feature matched with at least one of the semantic primitives to obtain the pose parameters of the target pipeline, comprises:
[0024] Perform feature matching between the geometric model and at least one of the semantic primitives to obtain a correspondence between a point in the geometric model and a point in at least one of the semantic primitives;
[0025] According to a preset point cloud matching algorithm and a correspondence between a point in the geometric model and a point in at least one of the semantic primitives, the position and posture parameters of the target pipeline are acquired.
[0026] Optionally, in the pipeline pose estimation method, the target detection network is obtained based on training of a pipeline image set, and each pipeline image in the pipeline image set is generated by establishing a pipeline model and simulating an environment.
[0027] In a second aspect, an embodiment of the present invention provides a pipeline posture estimation device, comprising:
[0028] A first generating module, configured to generate an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image;
[0029] A division module, used for dividing the observation point cloud into at least one semantic primitive according to the semantic label classification information of each point in the observation point cloud, wherein the semantic label classification information is used for indicating the semantic primitive category to which the point belongs;
[0030] A second generating module, configured to generate a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives;
[0031] A matching module is used to perform feature matching between the geometric model and at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline.
[0032] In a third aspect, an embodiment of the present invention provides a pipeline posture estimation device, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the pipeline posture estimation method as described in the first aspect is implemented.
[0033] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, on which a program is stored, and when the program is executed by a processor, the pipeline posture estimation method as described in the first aspect is implemented.
[0034] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the pipeline posture estimation method as described in the first aspect.
[0035] The beneficial effects of the above technical solution of the present invention are as follows:
[0036] In an embodiment of the present invention, an observation point cloud of a target pipeline in the target image is generated based on a target image and a depth image corresponding to the target image; the observation point cloud is divided into at least one semantic primitive according to the semantic label classification information of each point in the observation point cloud, and the semantic label classification information is used to indicate the semantic primitive category to which the point belongs; a geometric model of the target pipeline is generated according to a shape feature vector corresponding to at least one of the semantic primitives; the geometric model is feature matched with at least one of the semantic primitives to obtain the pose parameters of the target pipeline. In this way, pose estimation of category-level pipelines is achieved, the pose estimation requirements of complex spatial pipelines in the automated assembly of pipeline systems in aerospace products are met, and the efficiency and accuracy of robot assembly are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of a pipeline posture estimation method according to an embodiment of the present invention;
[0038] Figure 2 A schematic diagram of the position and posture parameters of the target pipeline according to an embodiment of the present invention;
[0039] Figure 3 Schematic diagram of the structure of the pipeline posture estimation device according to an embodiment of the present invention;
[0040] Figure 4 The hardware block diagram of the pipeline posture estimation device described in an embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0042] In various embodiments of the present invention, it should be understood that the sequence numbers of the following processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0043] Additionally, the terms "system" and "network" are often used interchangeably herein.
[0044] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0045] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a pipeline posture estimation method according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0046] S101, generating an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image.
[0047] In an embodiment of the present invention, the pixel points of the target pipeline in the target image are converted from an image coordinate system into a three-dimensional observation point cloud.
[0048] It should be noted that, in order to improve the quality of the observation point cloud, the embodiment of the present invention can perform denoising and filtering on the observation point cloud to remove cluster points and noise to obtain a high-quality observation point cloud.
[0049] In one implementation manner, optionally, generating an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image includes:
[0050] Obtaining a segmentation mask of a target pipeline in the target image according to a target detection network, a target image, and a depth image corresponding to the target image;
[0051] According to the depth image and the camera internal parameters, the pixel points in the segmentation mask of the target pipeline are mapped to the three-dimensional space to generate the observation point cloud of the target pipeline.
[0052] In an embodiment of the present invention, the target detection network may be Mask R-CNN (Mask Region–based Convolutional Neural Network, a deep learning model for target detection and instance segmentation); the target image may be an RGB image of a target pipeline. First, the target image and the depth image are combined, and the category, position, and segmentation area of the target pipeline are identified through the target detection network, thereby obtaining a segmentation mask of the target pipeline. Then, using the pixel values in the depth image and the camera intrinsic parameters (including focal length and principal point coordinates), the pixels (valid pixels) in the segmentation mask of the target pipeline are mapped to a three-dimensional space, and the three-dimensional coordinates of each pixel are calculated to generate an observation point cloud of the target pipeline.
[0053] In one embodiment, optionally, the target detection network is obtained by training based on a pipeline image set, and each pipeline image in the pipeline image set is generated by establishing a pipeline model and simulating an environment.
[0054] In the embodiment of the present invention, firstly, a pipeline model (eg, a CAD model of the pipeline) is established based on the production and processing YBC data of the pipeline. For example, a total of 20 pipeline models of different shapes are established, with lengths ranging from 0.5m to 1m.
[0055] In order to increase the robustness of the pipeline pose estimation method described in the embodiment of the present invention to interference objects, interference objects that are also texture-free like pipelines (such as metal parts, mugs, etc.) are selected from texture-free object datasets (such as T-LESS) and three-dimensional model datasets (such as ShapeNetCore).
[0056] Furthermore, 7 common backgrounds in industrial scenes were selected to create virtual scenes, including rubber mats, wooden desktops, concrete floors, etc., to increase the richness of the background. 3 to 6 pipes and 1 to 4 interference objects were placed in the created virtual scenes, and the color, roughness and metallicity of the pipes and interference objects were randomly set within a certain range.
[0057] In addition, based on the pipeline model, a simulation environment is established, including:
[0058] Lighting is an important factor affecting the appearance of objects in images. It is affected by the environment, time, etc. By performing domain randomization on lighting to cover more lighting conditions, generalization can be improved.
[0059] Set up multiple directional lights to simulate artificial lighting equipment such as ceiling lights, workshop lights, and direct exposure to natural light in the factory, and randomly adjust the position and intensity of the light source to simulate the lighting characteristics under different times and conditions;
[0060] Set appropriate shadow sharpness and softness to simulate shadows in the scene.
[0061] Finally, to simulate more complex situations such as pipe stacking, we let each pipe fall freely under the action of gravity in turn to create different pipe occlusion situations. Due to the complex geometric structure of the pipe, it presents large shape differences under different viewing angles. Randomizing the viewing angle to generate diverse data allows the model to learn more comprehensive and robust feature expressions.
[0062] Therefore, the embodiment of the present invention uses domain randomization technology and mixed reality technology to generate the pipeline image set and annotate the pipeline posture information. Through virtual simulation scenes and physical rendering technology, it is ensured that the generated pipeline image set has diversity and authenticity in terms of lighting, material, background, layout and posture, thereby reducing the annotation cost.
[0063] S102, dividing the observation point cloud into at least one semantic primitive according to semantic label classification information of each point in the observation point cloud, wherein the semantic label classification information is used to indicate a semantic primitive category to which the point belongs.
[0064] In one implementation manner, optionally, before dividing the observation point cloud into at least one semantic primitive according to the semantic label classification information of each point in the observation point cloud, the method further includes:
[0065] The semantic label classification of each point in the observation point cloud is performed through 3D-GCN (Three-Dimensionally Graph Convolutional Network) to obtain the semantic label classification information of each point in the observation point cloud.
[0066] In an embodiment of the present invention, the semantic label classification information is used to indicate the semantic primitive category to which the point in the observation point cloud belongs. Optionally, the semantic primitive category includes a straight line segment, an arc segment or a connector.
[0067] The semantic label classification is performed on each point in the observation point cloud through 3D-GCN, so that the observation point cloud is divided into different semantic primitives according to the semantic label classification information obtained by the semantic classification.
[0068] It should be noted that, for each point in the observation point cloud, the semantic label is calculated using the following formula (1).
[0069] l i =argmin j=1,2,…,Ng ‖p i -g j ‖2(1)
[0070] Among them, pi represents the three-dimensional coordinates of the i-th point in the observation point cloud, g j N represents the three-dimensional coordinates of the center point of the jth semantic primitive; g represents the total number of semantic primitives corresponding to the observation point cloud; l i It's point p i The corresponding semantic label; ‖·‖2 represents the Euclidean norm, the distance between two points.
[0071] S103: Generate a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives.
[0072] In one implementation manner, optionally, generating a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives includes:
[0073] For each of the semantic primitives, extract the geometric features of the semantic primitive to obtain a shape feature vector corresponding to the semantic primitive;
[0074] A geometric model of the target pipeline is generated according to a geometric generation network and a shape feature vector corresponding to at least one of the semantic primitives.
[0075] It should be noted that the shape feature vector can also be called a shape invariant descriptor, which is used to accurately express the local geometric form of the semantic primitive and has invariance (e.g., robustness to scaling, rotation, and translation). The shape invariant descriptor is a high-dimensional feature representation that can keep the shape unchanged under scaling, rotation, and translation transformations, and can decouple shape features and posture features, thereby avoiding the error transmission problem caused by the coupling of shape features and posture features.
[0076] All shape-invariant descriptors are mapped into a unified shape feature space, and primitives of the same category but with different geometric forms (such as straight line segments, arc segments, connectors, and other different shapes) are mapped into adjacent areas in the space, so that shape-invariant descriptors with similar geometric shapes are clustered in the feature space, and shape-invariant descriptors with large differences are distinguished from each other. This constructed shape feature space provides a standardized representation for subsequent shape matching and optimization, helps to accurately describe the different shape features of complex pipelines, and supports generalization capabilities for large intra-class shape differences. Here, pipelines of the same category with different shapes are encoded into a unified shape feature space, capturing the main structural information of the pipeline shape while ignoring subtle shape differences, thereby achieving unified modeling of diverse shapes within the class.
[0077] In the embodiment of the present invention, firstly, for each of the semantic primitives, the geometric features of the semantic primitive are extracted to obtain a shape feature vector of the semantic primitive.
[0078] In one implementation manner, optionally, extracting the geometric features of the semantic primitive to obtain the shape feature vector corresponding to the semantic primitive includes:
[0079] Centralizing the semantic primitive to obtain a centralized point cloud of the semantic primitive after processing;
[0080] According to the centralized point cloud of the semantic primitive, the geometric features of the semantic primitive are extracted to obtain a shape feature vector corresponding to the semantic primitive.
[0081] It should be noted that the semantic primitive is centralized to obtain a centralized point cloud of the processed semantic primitive, so as to transform the semantic primitive into the same reference coordinate system, and then the geometric features are extracted based on the centralized point cloud of the semantic primitive to obtain a shape feature vector.
[0082] Furthermore, a shape feature vector corresponding to at least one of the semantic primitives corresponding to the observation point cloud is input into a geometry generation network to generate a geometric model of the target pipeline.
[0083] In one embodiment, optionally, the geometry generation network includes a coarse-grained generator and a fine-grained generator, the coarse-grained generator is used to generate the geometry of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives, and the fine-grained generator is used to refine the geometry information based on the geometry and image of the target pipeline to generate a geometry model of the target pipeline.
[0084] It can be understood that, first, the shape feature vector corresponding to at least one of the semantic primitives is input into the coarse-grained generator in the geometry generation network, so that the semantic primitive can be quickly generated, that is, the geometric shape of the target pipeline, such as a straight line segment, an arc segment or the overall outline of a connector. Then, the geometric shape and image of the target pipeline are input into the fine-grained generator in the geometry generation network, and the geometric information is refined according to the shape feature vector corresponding to at least one of the semantic primitives, and the local details of the target pipeline, such as supplementary curvature, surface texture, etc., are generated, so as to finally generate the geometric model of the target pipeline.
[0085] It should be noted that, after generating the geometric model of the target pipeline, the pipeline pose estimation method according to the embodiment of the present invention further includes:
[0086] Performing shape optimization on the geometric model to obtain an optimized geometric model specifically includes:
[0087] By matching the generated geometric model with the observed point cloud, the geometric error between the two is calculated, and the shape feature vector mentioned above is iteratively adjusted based on this as the optimization target.
[0088] Through gradual optimization, the generated geometric model continuously approaches the actual observed point cloud in morphology, and finally a high-precision geometric model that is highly consistent with the actual pipeline is obtained, providing a reliable shape basis for subsequent pose estimation and measurement.
[0089] Shape optimization mainly focuses on the geometric structure of the pipeline and is independent of the position and posture information of the pipeline in three-dimensional space. This decoupling design of shape and posture can effectively avoid mutual interference between the two, reduce error transmission, and improve the accuracy and stability of subsequent posture estimation.
[0090] S104, performing feature matching between the geometric model and at least one of the semantic primitives to obtain position and posture parameters of the target pipeline.
[0091] It should be noted that the posture parameters of the target pipeline can accurately describe the position, direction and size of the target pipeline in three-dimensional space, and provide precise spatial information for robot grasping and assembly tasks.
[0092] The geometric model may be the optimized geometric model mentioned above.
[0093] In one implementation manner, optionally, performing feature matching between the geometric model and at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline includes:
[0094] Perform feature matching between the geometric model and at least one of the semantic primitives to obtain a correspondence between a point in the geometric model and a point in at least one of the semantic primitives;
[0095] According to a preset point cloud matching algorithm and a correspondence between a point in the geometric model and a point in at least one of the semantic primitives, the position and posture parameters of the target pipeline are acquired.
[0096] In an embodiment of the present invention, feature matching is performed between points in the geometric model and points in at least one of the semantic primitives, and the correspondence between the points in the geometric model and the points in at least one of the semantic primitives is determined by calculating the minimum Euclidean distance between the two.
[0097] Optionally, the preset point cloud matching algorithm is Umeyama algorithm.
[0098] The preset point cloud matching algorithm and the correspondence between the points in the geometric model and the points in at least one of the semantic primitives are used to calculate the pose parameters of the target pipeline according to the following formula (2). Optionally, the pose parameters include a scaling factor Rotation and translation
[0099]
[0100] Among them, argmin represents the minimization of the objective function, s represents the scaling factor, R represents the rotation amount, t represents the translation amount, N0 represents the number of observation point clouds, and p i represents the three-dimensional coordinates of the i-th point in the observation point cloud, g i Indicates that p i The three-dimensional coordinates corresponding to the center point of the i-th semantic primitive, ‖·‖2 represents the Euclidean norm, the distance between the two points.
[0101] The pose parameters of the target pipeline can be represented on the target image or the depth image, specifically, as Figure 2 As shown, the pose parameters are represented by three-dimensional coordinate axes, and the size is represented by a three-dimensional bounding box.
[0102] In addition, an embodiment of the present invention further provides a pipeline pose estimation model, which can implement the above pipeline pose estimation method when applied, and the application scenario is such as pose estimation of the above target pipeline. The pipeline pose estimation model can still effectively generalize and output high-quality estimation results when facing pipelines with large shape differences within the class, and has good robustness for environments with different backgrounds and lighting conditions.
[0103] The pipeline pose estimation model includes: a point cloud extraction module, a semantic primitive module, a shape description module, a shape optimization module and a pose estimation module. In the application process, the point cloud extraction module is used to generate an observation point cloud of the target pipeline in the target image according to the target image and the depth image corresponding to the target image; the semantic primitive module is used to divide the observation point cloud into at least one semantic primitive according to the semantic label classification information of each point in the observation point cloud; the shape description module is used to generate a shape feature vector corresponding to the semantic primitive, and generate a geometric model of the target pipeline according to the shape feature vector corresponding to at least one semantic primitive; the shape optimization module is used to; the pose estimation module is used to optimize the geometric model; the pose estimation module is used to perform feature matching between the optimized geometric model and at least one semantic primitive to obtain the pose parameters of the target pipeline.
[0104] It should be noted that the pipeline pose estimation model needs to be trained before application, that is, the above modules in the pipeline pose estimation model need to be trained. During the training process, the point cloud extraction module is used to train based on the above pipeline image set to extract the observation point cloud of the pipeline in each pipeline image in the pipeline image set.
[0105] In summary, the pipeline pose estimation method described in the embodiment of the present invention can achieve accurate category-level pipeline pose estimation. Accurate category-level pipeline pose estimation enables the robot to perform accurate grasping actions as planned, thereby improving the efficiency and accuracy of automatic assembly.
[0106] See also Figure 3 , Figure 3 Schematic diagram of the structure of the pipeline posture estimation device according to an embodiment of the present invention. Figure 3 As shown, the device comprises:
[0107] A first generating module 301 is used to generate an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image;
[0108] A division module 302, configured to divide the observation point cloud into at least one semantic primitive according to semantic label classification information of each point in the observation point cloud, wherein the semantic label classification information is used to indicate a semantic primitive category to which the point belongs;
[0109] A second generating module 303 is used to generate a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives;
[0110] The matching module 304 is used to perform feature matching between the geometric model and at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline.
[0111] Optionally, in the pipeline posture estimation device, the first generating module 301 is specifically used for:
[0112] Obtaining a segmentation mask of a target pipeline in the target image according to a target detection network, a target image, and a depth image corresponding to the target image;
[0113] According to the depth image and the camera internal parameters, the pixel points in the segmentation mask of the target pipeline are mapped to the three-dimensional space to generate the observation point cloud of the target pipeline.
[0114] Optionally, the pipeline posture estimation device further comprises:
[0115] The acquisition module is used to perform semantic label classification on each point in the observation point cloud through a three-dimensional image convolution network to obtain semantic label classification information of each point in the observation point cloud.
[0116] Optionally, in the pipeline posture estimation device, the second generation module 303 includes:
[0117] An extraction unit, configured to extract the geometric features of each semantic primitive, and obtain a shape feature vector corresponding to the semantic primitive;
[0118] A generating unit is used to generate a geometric model of the target pipeline according to a geometric generating network and a shape feature vector corresponding to at least one of the semantic primitives.
[0119] Optionally, in the pipeline posture estimation device, the extraction unit is specifically used to:
[0120] Centralizing the semantic primitive to obtain a centralized point cloud of the semantic primitive after processing;
[0121] According to the centralized point cloud of the semantic primitive, the geometric features of the semantic primitive are extracted to obtain a shape feature vector corresponding to the semantic primitive.
[0122] Optionally, in the pipeline pose estimation device, the geometric generation network includes a coarse-grained generator and a fine-grained generator, the coarse-grained generator is used to generate the geometric shape of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives, and the fine-grained generator is used to refine the geometric information based on the geometric shape and image of the target pipeline to generate a geometric model of the target pipeline.
[0123] Optionally, in the pipeline posture estimation device, the matching module 304 is specifically used for:
[0124] Perform feature matching between the geometric model and at least one of the semantic primitives to obtain a correspondence between a point in the geometric model and a point in at least one of the semantic primitives;
[0125] According to a preset point cloud matching algorithm and a correspondence between a point in the geometric model and a point in at least one of the semantic primitives, the position and posture parameters of the target pipeline are acquired.
[0126] Optionally, in the pipeline pose estimation device, the target detection network is obtained based on training of a pipeline image set, and each pipeline image in the pipeline image set is generated by establishing a pipeline model and simulating an environment.
[0127] The pipeline posture estimation device provided in the embodiment of the present invention can execute the above-mentioned pipeline posture estimation method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.
[0128] like Figure 4As shown, an embodiment of the present invention provides a pipeline posture estimation device, including: a processor 401; and a memory 402 connected to the processor 401 through a bus interface, the memory 402 is used to store programs and data used by the processor 401 when performing operations, and the processor 401 calls and executes the programs and data stored in the memory 402.
[0129] The processor 401 is used to read the program in the memory 402 and execute the following steps:
[0130] Generate an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image;
[0131] Dividing the observation point cloud into at least one semantic primitive according to semantic label classification information of each point in the observation point cloud, wherein the semantic label classification information is used to indicate the semantic primitive category to which the point belongs;
[0132] generating a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives;
[0133] The geometric model is feature matched with at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline.
[0134] Among them, Figure 4 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically one or more processors represented by processor 401 and various circuits of memory represented by memory 402 are linked together. The bus architecture may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface. The transceiver 403 may be a plurality of components, namely, a transmitter and a transceiver, providing a unit for communicating with various other devices on a transmission medium. For different user devices, the user interface 404 may also be an interface capable of externally and internally connecting required devices, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, and the like.
[0135] The processor 401 is responsible for managing the bus architecture and general processing, and the memory 402 can store data used by the processor 401 when performing operations.
[0136] Optionally, the processor 401 is specifically configured to read a program in the memory 402 and execute the following steps:
[0137] Obtaining a segmentation mask of a target pipeline in the target image according to a target detection network, a target image, and a depth image corresponding to the target image;
[0138] According to the depth image and the camera internal parameters, the pixel points in the segmentation mask of the target pipeline are mapped to the three-dimensional space to generate the observation point cloud of the target pipeline.
[0139] Optionally, the processor 401 is further configured to read a program in the memory 402 and execute the following steps:
[0140] Semantic label classification is performed on each point in the observation point cloud through a three-dimensional image convolutional network to obtain semantic label classification information of each point in the observation point cloud.
[0141] Optionally, the processor 401 is specifically configured to read a program in the memory 402 and execute the following steps:
[0142] For each of the semantic primitives, extract the geometric features of the semantic primitive to obtain a shape feature vector corresponding to the semantic primitive;
[0143] A geometric model of the target pipeline is generated according to a geometric generation network and a shape feature vector corresponding to at least one of the semantic primitives.
[0144] Optionally, the processor 401 is specifically configured to read a program in the memory 402 and execute the following steps:
[0145] Centralizing the semantic primitive to obtain a centralized point cloud of the semantic primitive after processing;
[0146] According to the centralized point cloud of the semantic primitive, the geometric features of the semantic primitive are extracted to obtain a shape feature vector corresponding to the semantic primitive.
[0147] Optionally, the geometry generation network includes a coarse-grained generator and a fine-grained generator, the coarse-grained generator is used to generate the geometry of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives, and the fine-grained generator is used to refine the geometry information based on the geometry and image of the target pipeline to generate a geometry model of the target pipeline.
[0148] Optionally, the processor 401 is specifically configured to read a program in the memory 402 and execute the following steps:
[0149] Perform feature matching between the geometric model and at least one of the semantic primitives to obtain a correspondence between a point in the geometric model and a point in at least one of the semantic primitives;
[0150] According to a preset point cloud matching algorithm and a correspondence between a point in the geometric model and a point in at least one of the semantic primitives, the position and posture parameters of the target pipeline are acquired.
[0151] Optionally, the target detection network is obtained based on training of a pipeline image set, and each pipeline image in the pipeline image set is generated by establishing a pipeline model and simulating an environment.
[0152] A specific embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the above-mentioned pipeline posture estimation method are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0153] In addition, an embodiment of the present invention further provides a computer program product, including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0154] In the several embodiments provided by the present invention, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0155] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may be physically included separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0156] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform some steps of the sending and receiving methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.
[0157] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A pipeline posture estimation method, characterized in that: include: Generate an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image; Dividing the observation point cloud into at least one semantic primitive according to semantic label classification information of each point in the observation point cloud, wherein the semantic label classification information is used to indicate the semantic primitive category to which the point belongs; generating a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives; The geometric model is feature matched with at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline.
2. The pipeline posture estimation method according to claim 1, characterized in that: Generating an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image, including: Obtaining a segmentation mask of a target pipeline in the target image according to a target detection network, a target image, and a depth image corresponding to the target image; According to the depth image and the camera internal parameters, the pixel points in the segmentation mask of the target pipeline are mapped to the three-dimensional space to generate the observation point cloud of the target pipeline.
3. The pipeline posture estimation method according to claim 1, characterized in that: Before dividing the observation point cloud into at least one semantic primitive according to the semantic label classification information of each point in the observation point cloud, the method further includes: Semantic label classification is performed on each point in the observation point cloud through a three-dimensional image convolutional network to obtain semantic label classification information of each point in the observation point cloud.
4. The pipeline posture estimation method according to claim 1, characterized in that: Generating a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives, comprising: For each of the semantic primitives, extract the geometric features of the semantic primitive to obtain a shape feature vector corresponding to the semantic primitive; A geometric model of the target pipeline is generated according to a geometric generation network and a shape feature vector corresponding to at least one of the semantic primitives.
5. The pipeline posture estimation method according to claim 4, characterized in that: Extracting the geometric features of the semantic primitives to obtain the shape feature vector corresponding to the semantic primitives includes: Centralizing the semantic primitive to obtain a centralized point cloud of the semantic primitive after processing; According to the centralized point cloud of the semantic primitive, the geometric features of the semantic primitive are extracted to obtain a shape feature vector corresponding to the semantic primitive.
6. The pipeline posture estimation method according to claim 4, characterized in that: The geometric generation network includes a coarse-grained generator and a fine-grained generator. The coarse-grained generator is used to generate the geometric shape of the target pipeline according to the shape feature vector corresponding to at least one of the semantic primitives, and the fine-grained generator is used to refine the geometric information based on the geometric shape and image of the target pipeline to generate a geometric model of the target pipeline.
7. The pipeline posture estimation method according to claim 1, characterized in that: Performing feature matching between the geometric model and at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline includes: Perform feature matching between the geometric model and at least one of the semantic primitives to obtain a correspondence between a point in the geometric model and a point in at least one of the semantic primitives; According to a preset point cloud matching algorithm and a correspondence between a point in the geometric model and a point in at least one of the semantic primitives, the position and posture parameters of the target pipeline are acquired.
8. The pipeline posture estimation method according to claim 2, characterized in that: The target detection network is obtained by training based on a pipeline image set, and each pipeline image in the pipeline image set is generated by establishing a pipeline model and simulating an environment.
9. A pipeline posture estimation device, characterized in that: include: A first generating module, configured to generate an observation point cloud of a target pipeline in the target image according to a target image and a depth image corresponding to the target image; A division module, used for dividing the observation point cloud into at least one semantic primitive according to the semantic label classification information of each point in the observation point cloud, wherein the semantic label classification information is used for indicating the semantic primitive category to which the point belongs; A second generating module, configured to generate a geometric model of the target pipeline according to a shape feature vector corresponding to at least one of the semantic primitives; A matching module is used to perform feature matching between the geometric model and at least one of the semantic primitives to obtain the position and posture parameters of the target pipeline.
10. A pipeline posture estimation device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the pipeline posture estimation method as described in any one of claims 1 to 8 is implemented.
11. A readable storage medium, characterized in that: The readable storage medium stores a program, and when the program is executed by the processor, the pipeline posture estimation method according to any one of claims 1 to 8 is implemented.
12. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the pipeline posture estimation method as described in any one of claims 1 to 8.