Three-dimensional point cloud data processing and pose estimation method, device, equipment, medium and program product
By building a fusion multi-scale feature extraction network and pose estimation framework, feature extraction and pose estimation of three-dimensional point cloud data is solved, and the problem of relying on manual and low accuracy in the existing technology is realized, and high-precision pose estimation and automated data acquisition are achieved.
Patent Information
- Application Number
- CN202510102691.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing three-dimensional object detection technology relies on manual labor and has problems such as low accuracy and poor practicality. Especially when dealing with complex backgrounds and occlusions, it is difficult to achieve high-precision pose estimation.
By collecting three-dimensional point cloud data and preprocessing, a feature fusion model integrating a multi-scale feature extraction network is constructed, feature extraction is performed on three-dimensional point cloud data, pose estimation framework is constructed, pose estimation is performed on high-dimensional features, pose information is obtained, and the robot control system is controlled for grabbing operations.
Through automated data acquisition and labeling, the detection accuracy and robustness are improved, the accuracy and application of pose estimation are improved, and high-precision pose estimation is achieved under complex backgrounds and occlusions.
Smart Images

Figure CN120014050A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of three-dimensional point cloud data processing, and more specifically, to a three-dimensional point cloud data processing and posture estimation method, device, equipment, medium and program product. Background Art
[0002] Existing 3D point cloud data processing technologies include 3D target detection technology, pose estimation technology, and data acquisition and processing technology. 3D target detection technology is mainly used to identify and locate objects from 3D data (such as point clouds). Pose estimation technology aims to determine the position and orientation of objects in 3D space. Data acquisition and processing technology is the first stage of big data processing, including data acquisition, data preprocessing, data storage, data analysis and mining, and data visualization.
[0003] In the current field of 3D object detection and pose estimation, the process of data collection and annotation is highly dependent on manual work. Traditional pose estimation methods have problems of low accuracy and poor practicality in practical applications, especially when dealing with complex backgrounds and occlusions. It is difficult to achieve high-precision pose estimation. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a three-dimensional point cloud data processing and pose estimation method, device, equipment, medium and program product to solve the problems that the existing three-dimensional target detection technology relies on manual labor and has low accuracy and poor practicality.
[0005] In a first aspect, an embodiment of the present application provides a method for processing three-dimensional point cloud data and estimating a pose, comprising:
[0006] Collect 3D point cloud data and pre-process it;
[0007] Construct a feature fusion model that integrates multi-scale feature extraction networks;
[0008] Extract features from 3D point cloud data based on feature fusion model to convert 3D point cloud data into high-dimensional features;
[0009] Construct a pose estimation framework, perform pose estimation on high-dimensional features based on the pose estimation framework, and obtain pose information;
[0010] According to the posture information, the robot control system is controlled to perform grasping operations.
[0011] In the above implementation process, the embodiment of the present application collects and preprocesses three-dimensional point cloud data; constructs a feature fusion model that integrates a multi-scale feature extraction network; extracts features from the three-dimensional point cloud data based on the feature fusion model to convert the three-dimensional point cloud data into high-dimensional features; constructs a pose estimation framework, and performs pose estimation on the high-dimensional features based on the pose estimation framework to obtain pose information; controls the robot control system to perform grasping operations based on the pose information; improves detection accuracy and robustness through automated data collection and annotation, and through multi-scale and multi-model feature fusion, thereby improving the accuracy and applicability of pose estimation.
[0012] Furthermore, the collecting of three-dimensional point cloud data and preprocessing includes:
[0013] Acquire three-dimensional point cloud data and annotate the three-dimensional point cloud data;
[0014] The farthest point sampling method is used to downsample the 3D point cloud data;
[0015] De-noising of 3D point cloud data.
[0016] In the above implementation process, the data is downsampled and denoised to reduce the number of data points while maintaining the integrity of the object shape features.
[0017] Furthermore, the feature fusion model of building a fused multi-scale feature extraction network includes:
[0018] Construct a feature fusion model that integrates three-dimensional convolutional neural network and point cloud feature learning network.
[0019] In the above implementation process, the model can realize the fusion of multi-scale features, thereby improving the ability to represent three-dimensional objects; this fusion method not only enhances the network's capture of global texture information, but also improves the ability to represent the distance relationship of local point clouds.
[0020] Furthermore, the construction of the pose estimation framework includes:
[0021] Build a 3D target classification and pose estimation database for automatic data annotation.
[0022] In the above implementation process, the breadth and width of the data set are improved through automatic data generation and annotation.
[0023] Furthermore, the pose estimation is performed on the high-dimensional features based on the pose estimation framework to obtain the pose information, including:
[0024] Generate candidate three-dimensional regions based on the global feature model;
[0025] Combined with the local feature model, regression calculation is performed on the three-dimensional point cloud data in the three-dimensional area to obtain the six-degree-of-freedom position information of the target.
[0026] In the above implementation process, the two-stage prediction process not only improves the accuracy of pose estimation, but also makes the algorithm more effective when dealing with objects containing noise.
[0027] Furthermore, the collecting of three-dimensional point cloud data includes:
[0028] Collect 3D point cloud data through a lidar camera or a depth camera.
[0029] In the above implementation process, the accuracy and robustness of data collection are improved.
[0030] In a second aspect, an embodiment of the present application provides a three-dimensional point cloud data processing and pose estimation device, comprising:
[0031] Data acquisition module, used to collect 3D point cloud data and perform preprocessing;
[0032] Model building module, used to build a feature fusion model integrating multi-scale feature extraction network;
[0033] A feature extraction module is used to extract features from three-dimensional point cloud data based on a feature fusion model to convert the three-dimensional point cloud data into high-dimensional features;
[0034] The pose estimation module is used to build a pose estimation framework, perform pose estimation on high-dimensional features based on the pose estimation framework, and obtain pose information;
[0035] The operation control module is used to control the robot control system to perform grasping operations according to the posture information.
[0036] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0037] A processor, a memory and a bus, wherein the processor is connected to the memory via the bus, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement the three-dimensional point cloud data processing and pose estimation method as described above.
[0038] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a server, the three-dimensional point cloud data processing and pose estimation method as described above is implemented.
[0039] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed by a computer, the computer implements the three-dimensional point cloud data processing and pose estimation method as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0041] Figure 1 A flowchart of a three-dimensional point cloud data processing and pose estimation method provided in an embodiment of the present application;
[0042] Figure 2 It is a structural schematic diagram of a three-dimensional point cloud data processing and posture estimation device provided in an embodiment of the present application;
[0043] Figure 3 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0045] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0046] In the current field of three-dimensional target detection and pose estimation, the existing technology has several objective shortcomings, which limit the effectiveness and reliability of the technology in practical applications. First, the process of data collection and annotation is highly dependent on manual labor, which not only consumes a lot of time and human resources, but also significantly increases project costs and limits the expansion capabilities and diversity of data sets. In order to solve this problem, the embodiment of the present application uses a robot operating system environment to automatically complete data collection and annotation. By establishing a target model in a simulated environment, performing pose transformation and annotation information recording, this method can realize the construction of a large-scale three-dimensional point cloud database, significantly reduce labor costs, and improve the efficiency of data collection and annotation.
[0047] Secondly, existing 3D target detection algorithms often lack robustness, efficiency and flexibility when processing 3D target information at different scales and regions. Especially in complex scenes, such as deep-sea environments, these algorithms have difficulty accurately identifying and locating targets. The embodiment of the present application proposes a new 3D target detection framework by combining a deep learning method of global features and local features. This framework can more accurately describe and model feature information of different scales and regions in 3D point cloud data, thereby improving the accuracy and robustness of target classification and posture estimation.
[0048] Third, traditional pose estimation methods have problems of low accuracy and poor practicality in practical applications, especially when dealing with complex backgrounds and occlusions, it is difficult to achieve high-precision pose estimation. The embodiment of the present application proposes a three-dimensional target classification and pose estimation algorithm based on the combination of candidate boxes and pose regression. Through prediction area generation and point cloud regression calculation, this method can achieve more precise target pose estimation, improve the accuracy of pose estimation and its applicability in practical applications.
[0049] Finally, the existing technology often ignores the importance of multi-scale and multi-model feature fusion when processing three-dimensional point cloud data, resulting in performance bottlenecks in object detection and pose estimation at different scales. The embodiment of the present application designs a deep neural network model with multi-model and multi-feature fusion, which realizes the effective combination of global features and local features. This method can improve the network's ability to represent objects and enhance the classification and pose estimation accuracy of subsequent algorithms, especially when dealing with scale changes and complex backgrounds.
[0050] In summary, the embodiments of the present application effectively solve some key shortcomings in the prior art through technical means such as automated data collection and annotation, improved detection accuracy and robustness, improved accuracy and applicability of pose estimation, and realization of multi-scale and multi-model feature fusion, bringing significant progress and innovation to the field of three-dimensional target detection and pose estimation.
[0051] 3D target detection technology mainly relies on point cloud data. It uses 3D point cloud data acquired by radar point cloud cameras or depth cameras, combined with feature extraction and semantic label assignment to achieve target classification and location. In terms of feature modeling methods, it is mainly divided into conversion-based methods, that is, converting 3D point cloud data into 2D images or voxelized matrices, and point processing-based methods, such as PointNet and PointNet++, which directly process point cloud data through multi-layer perceptrons.
[0052] Pose estimation technology involves two stages: rough pose estimation and precise pose estimation. First, a rough pose estimation is performed by generating and screening candidate 3D boxes, and then regression calculations are performed on the point clouds in these areas to obtain precise six-degree-of-freedom (6DoF) pose estimation. In this field, deep learning models such as VoteNet and MV3D are widely used for target classification, positioning, and pose estimation. These models can process complex 3D data and provide relatively accurate results.
[0053] In terms of data collection and processing, in order to reduce labor costs and improve the scalability of data sets, data generation and automatic annotation are simulated in the robot operating system environment. In addition, in order to improve computing efficiency, methods such as farthest point sampling (FPS) are used to downsample the data to reduce the number of data points, thereby reducing the consumption of computing resources while maintaining the shape characteristics of the object. The development and application of these technologies have brought new possibilities and challenges to the field of 3D object detection and pose estimation.
[0054] The embodiment of the present application combines deep learning theory with the robot operating system to achieve automatic collection, processing and pose estimation of three-dimensional point cloud data, thereby improving the accuracy and efficiency of the robot in target grasping tasks. Compared with the prior art, the embodiment of the present application shows significant differences and improvements in many aspects:
[0055] In terms of data collection and automatic labeling, the embodiments of the present application are different from traditional methods that rely on manual labeling. By simulating data generation and automatic labeling in a robot operating system environment, it effectively reduces labor costs and enhances the scalability of the data set. This improvement not only improves the efficiency of data processing, but also provides the possibility of building larger-scale data sets.
[0056] In the design of the feature extraction network, the embodiment of the present application proposes a multi-model fusion network that fuses global features and local features. Compared with the method that relies only on global or local features, it can characterize three-dimensional objects more comprehensively; this multi-model fusion method enables the network to capture the global structure and local details of the object at the same time, thereby improving the ability to extract object features.
[0057] In terms of the pose estimation algorithm, the embodiment of the present application adopts a two-stage prediction process. First, a rough pose estimation is performed by generating candidate three-dimensional regions (3D Bounding Boxes); then, regression calculations are performed on the point cloud data in these regions to obtain more refined six-degree-of-freedom (6DoF) pose information; this staged approach not only improves the accuracy of pose estimation, but also makes the algorithm more robust in practical applications.
[0058] Please see Figure 1 , Figure 1 A flow chart of a method for processing three-dimensional point cloud data and estimating position and posture provided in an embodiment of the present application. Figure 1 The three-dimensional point cloud data processing and pose estimation method comprises:
[0059] 100. Collect 3D point cloud data and perform preprocessing.
[0060] Specifically, three-dimensional point cloud data is collected through a lidar camera or a depth camera, and the data is automatically labeled using a robot operating system.
[0061] Specifically, three-dimensional point cloud data is obtained and data annotation is performed on the three-dimensional point cloud data; the three-dimensional point cloud data is downsampled using the farthest point sampling method; and the three-dimensional point cloud data is denoised.
[0062] It can be understood that in order to improve computing efficiency, the embodiment of the present application adopts the farthest point sampling (FPS) method to downsample the data to reduce the number of data points while maintaining the integrity of the object shape features.
[0063] Exemplarily, a lidar camera assembly for acquiring three-dimensional point cloud data is placed on top of the robot to facilitate data collection.
[0064] Optionally, in addition to using lidar cameras and depth cameras, other types of 3D data acquisition devices such as structured light scanners, time-of-flight (ToF) cameras, or multi-sensor fusion can be used to improve data accuracy and robustness.
[0065] Optionally, in terms of data processing and downsampling techniques, in addition to farthest point sampling, other downsampling methods such as random sampling consistency or Voxel Grid filter can be used, as well as data enhancement techniques such as rotation, scaling, and shearing to increase the diversity of the data set and reduce overfitting.
[0066] 200. Construct a feature fusion model that integrates multi-scale feature extraction networks.
[0067] Specifically, a feature fusion model is constructed by integrating a three-dimensional convolutional neural network and a point cloud feature learning network. This model can realize the fusion of multi-scale features, thereby improving the ability to represent three-dimensional objects. This fusion method not only enhances the network's capture of global texture information, but also improves the ability to represent the distance relationship of local point clouds.
[0068] 300. Feature extraction is performed on three-dimensional point cloud data based on a feature fusion model to convert the three-dimensional point cloud data into high-dimensional features.
[0069] 400. Construct a pose estimation framework, perform pose estimation on high-dimensional features based on the pose estimation framework, and obtain pose information.
[0070] Specifically, a 3D target classification and pose estimation database is constructed for automatic data labeling; by automatically generating data labels, the breadth and width of the data set are improved.
[0071] Optionally, a candidate three-dimensional region is generated based on a global feature model; and in combination with a local feature model, regression calculation is performed on the three-dimensional point cloud data within the three-dimensional region to obtain six-degree-of-freedom pose information of the target.
[0072] It can be understood that in the first stage, candidate three-dimensional regions are generated based on the global feature model; in the second stage, the point clouds in these regions are regressed in combination with the local feature model to obtain accurate pose information; this two-stage prediction process not only improves the accuracy of pose estimation, but also makes the algorithm more effective when dealing with objects containing noise.
[0073] Exemplarily, a posture estimation unit is provided, and a multi-layer neural network is used to calculate the precise posture of the object to achieve the positioning of the grasping point.
[0074] 500. According to the position information, the robot control system is controlled to perform a grasping operation.
[0075] Exemplarily, a robotic arm grasping component is provided, and precise grasping is achieved through a robot control system according to the posture parameters provided by the posture estimation unit.
[0076] As described above, the embodiments of the present application collect and preprocess three-dimensional point cloud data; construct a feature fusion model that integrates a multi-scale feature extraction network; extract features from the three-dimensional point cloud data based on the feature fusion model to convert the three-dimensional point cloud data into high-dimensional features; construct a pose estimation framework, and perform pose estimation on the high-dimensional features based on the pose estimation framework to obtain pose information; according to the pose information, control the robot control system to perform grasping operations; through automated data collection and annotation, and through multi-scale and multi-model feature fusion, improve detection accuracy and robustness, and improve the accuracy and applicability of pose estimation.
[0077] The embodiment of the present application proposes a multi-model fusion network that combines global features with local features to improve the accuracy of three-dimensional object classification and pose estimation; this fusion method makes up for the shortcomings of a single feature extraction method by combining a point cloud feature learning network with a three-dimensional convolutional neural network, enhances the ability to extract detailed features of the target object, and improves the accuracy of object classification and pose estimation.
[0078] In response to the problem that deep learning requires high labor costs for data collection and annotation, the embodiments of the present application utilize a robot operating system to achieve automatic data collection and annotation. This method reduces manual intervention and reduces costs, while improving data consistency and repeatability, which is crucial for training high-quality deep learning models.
[0079] The embodiment of the present application proposes a two-stage pose estimation algorithm framework, which first completes the rough prediction of the target pose by generating and screening candidate three-dimensional boxes, and then performs regression calculation on the point cloud within the three-dimensional box area to obtain a refined target six-degree-of-freedom (6DoF) pose estimation result; this staged method improves the accuracy of the detection method through multi-stage regression calculation while maintaining speed.
[0080] In order to solve the problem of characterizing target scale changes, the embodiment of the present application designs a multi-scale feature extraction network. By fusing features of deep and shallow networks, information integration of different scales is achieved, the loss of detail information in the feature extraction process is reduced, and the model's generalization ability for objects of different scales is improved.
[0081] The embodiment of the present application constructs a three-dimensional target classification and pose estimation database in a robot operating system environment that can automatically complete data labeling. This method improves the breadth and width of the data set by simulating environmental data modeling and automatic data generation and labeling, which is more in line with actual background requirements.
[0082] The embodiment of the present application proposes a classification and posture prediction framework for end-to-end training and prediction, which ensures a balance between accuracy and speed and does not have redundant calculation processes, which is particularly important for real-time requirements in practical applications.
[0083] In summary, the technical solution of the embodiment of the present application achieves objective improvement in three-dimensional target detection and pose estimation through automated data collection and annotation, multi-model feature fusion network, and two-stage pose estimation algorithm, improves efficiency, accuracy, and robustness, and has obvious practical application value. These improvements not only promote the development of technology, but also provide more reliable technical support for the operation of robots in complex environments.
[0084] The above steps are not to be performed in a strict order as described in the numbers, but should be understood as an overall solution.
[0085] In the second aspect, based on the above embodiments, Figure 2 A schematic diagram of the structure of a three-dimensional point cloud data processing and pose estimation device provided in an embodiment of the present application. Figure 2The three-dimensional point cloud data processing and posture estimation device provided in this embodiment specifically includes: a data acquisition module 201, a model construction module 202, a feature extraction module 203, a posture estimation module 204 and an operation control module 205.
[0086] Among them, the data acquisition module 201 is used to collect three-dimensional point cloud data and perform preprocessing; the model construction module 202 is used to construct a feature fusion model that integrates a multi-scale feature extraction network; the feature extraction module 203 is used to extract features from the three-dimensional point cloud data based on the feature fusion model to convert the three-dimensional point cloud data into high-dimensional features; the posture estimation module 204 is used to construct a posture estimation framework, and perform posture estimation on the high-dimensional features based on the posture estimation framework to obtain posture information; the operation control module 205 is used to control the robot control system to perform grasping operations according to the posture information.
[0087] As described above, the embodiments of the present application collect and preprocess three-dimensional point cloud data; construct a feature fusion model that integrates a multi-scale feature extraction network; extract features from the three-dimensional point cloud data based on the feature fusion model to convert the three-dimensional point cloud data into high-dimensional features; construct a pose estimation framework, and perform pose estimation on the high-dimensional features based on the pose estimation framework to obtain pose information; according to the pose information, control the robot control system to perform grasping operations; through automated data collection and annotation, and through multi-scale and multi-model feature fusion, improve detection accuracy and robustness, and improve the accuracy and applicability of pose estimation.
[0088] The three-dimensional point cloud data processing and posture estimation device provided in the embodiment of the present application can be used to execute the three-dimensional point cloud data processing and posture estimation method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0089] In a third aspect, an embodiment of the present application further provides an electronic device that can integrate the three-dimensional point cloud data processing and posture estimation device provided in an embodiment of the present application. Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 3 The electronic device includes: an input device 43, an output device 44, a memory 42 and one or more processors 41; the memory 42 is used to store one or more programs; when the one or more programs are executed by the one or more processors 41, the one or more processors 41 implement the three-dimensional point cloud data processing and pose estimation method provided in the above embodiment. The input device 43, the output device 44, the memory 42 and the processor 41 can be connected by a bus or other means. Figure 3 The example of connecting through bus is taken in the following.
[0090] The processor 41 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 42, that is, realizes the above-mentioned three-dimensional point cloud data processing and pose estimation method.
[0091] The electronic device provided above can be used to execute the three-dimensional point cloud data processing and posture estimation method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0092] In a fourth aspect, an embodiment of the present application also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the three-dimensional point cloud data processing and pose estimation method as described above, and can achieve the same beneficial effects.
[0093] Of course, the storage medium containing computer executable instructions provided in an embodiment of the present application, whose computer executable instructions are not limited to the three-dimensional point cloud data processing and pose estimation method described above, can also execute related operations in the three-dimensional point cloud data processing and pose estimation method provided in any embodiment of the present application.
[0094] In a fifth aspect, the embodiments of the present application also provide a computer program product, and the methods described in the various embodiments of the present application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the processes or functions described in the various embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM (OpenAplicationModel, open application model) or other programmable device.
[0095] The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0096] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0097] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0098] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0099] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0100] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0101] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
Claims
1. A three-dimensional point cloud data processing and pose estimation method, characterized in that: include: Collect 3D point cloud data and pre-process it; Construct a feature fusion model that integrates multi-scale feature extraction networks; Extract features from 3D point cloud data based on feature fusion model to convert 3D point cloud data into high-dimensional features; Construct a pose estimation framework, perform pose estimation on high-dimensional features based on the pose estimation framework, and obtain pose information; According to the posture information, the robot control system is controlled to perform grasping operations.
2. The three-dimensional point cloud data processing and pose estimation method according to claim 1, characterized in that: The collecting and preprocessing of three-dimensional point cloud data includes: Acquire three-dimensional point cloud data and annotate the three-dimensional point cloud data; The farthest point sampling method is used to downsample the 3D point cloud data; De-noising of 3D point cloud data.
3. The three-dimensional point cloud data processing and pose estimation method according to claim 1, characterized in that: The feature fusion model of building a fusion multi-scale feature extraction network includes: Construct a feature fusion model that integrates three-dimensional convolutional neural network and point cloud feature learning network.
4. The three-dimensional point cloud data processing and pose estimation method according to claim 1, characterized in that: The construction of the pose estimation framework includes: Build a 3D target classification and pose estimation database for automatic data annotation.
5. The three-dimensional point cloud data processing and pose estimation method according to claim 1, characterized in that: The method of performing posture estimation on high-dimensional features based on a posture estimation framework to obtain posture information includes: Generate candidate three-dimensional regions based on the global feature model; Combined with the local feature model, regression calculation is performed on the three-dimensional point cloud data in the three-dimensional area to obtain the six-degree-of-freedom position information of the target.
6. The method for processing three-dimensional point cloud data and estimating position and posture according to claim 1, characterized in that: The collecting of three-dimensional point cloud data comprises: Collect 3D point cloud data through a lidar camera or a depth camera.
7. A three-dimensional point cloud data processing and pose estimation device, characterized in that: include: Data acquisition module, used to collect 3D point cloud data and perform preprocessing; Model building module, used to build a feature fusion model integrating multi-scale feature extraction network; A feature extraction module is used to extract features from three-dimensional point cloud data based on a feature fusion model to convert the three-dimensional point cloud data into high-dimensional features; The pose estimation module is used to build a pose estimation framework, perform pose estimation on high-dimensional features based on the pose estimation framework, and obtain pose information; The operation control module is used to control the robot control system to perform grasping operations according to the posture information.
8. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the processor is connected to the memory via the bus, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement the three-dimensional point cloud data processing and pose estimation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a server, implements the three-dimensional point cloud data processing and pose estimation method as described in any one of claims 1-6.
10. A computer program product, characterized in that The computer program product comprises instructions, which, when executed by a computer, enable the computer to implement the three-dimensional point cloud data processing and pose estimation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
6D pose estimation method fusing point cloud local features
CN113221647A
Point cloud target detection method and device, equipment and storage medium
CN115082885A
Human body posture estimation method and system based on multi-modal fusion
CN116453166A
Head posture estimation method, device, equipment and vehicle
CN118351516A
Cited By
Two-stage high-precision three-dimensional point cloud semantic map construction method
CN120451609A