Point cloud recognition method, electronic equipment, storage medium and product
By extracting and fusing the global and local features of the point cloud to generate a third point cloud with the same structure, the problems of low efficiency and accuracy of point cloud recognition are solved, and efficient and accurate point cloud recognition is achieved.
Patent Information
- Application Number
- CN202510058782.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-09-23
AI Technical Summary
In existing technologies, point cloud recognition has low efficiency and accuracy, manual labeling is time-consuming and labor-intensive, and it is difficult to effectively determine the point cloud recognition results.
The global features of the first object and the global features of the second object are extracted, converted into local features and fused to generate a third point cloud with the same structure, and the recognition result of the first object is determined using the point cloud.
It improves the efficiency and accuracy of point cloud recognition, reduces the need for manual labeling, and achieves more efficient and accurate point cloud recognition.
Smart Images

Figure CN120689856A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a point cloud recognition method, electronic equipment, storage medium and product. Background Art
[0002] In the related art, two point clouds can be manually annotated to obtain point cloud recognition results, but the manual annotation method is inefficient, resulting in low efficiency in confirming the point cloud recognition results. In addition, the difference between the two point clouds is large, resulting in low accuracy of the determined point cloud recognition results. Summary of the Invention
[0003] The embodiments of the present application provide a point cloud recognition method, electronic device, storage medium and product, which can improve the efficiency of determining the recognition results of the point cloud and improve the accuracy of the recognition results of the point cloud.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides a method for point cloud recognition, which includes:
[0006] Extracting global features of a first object based on a first point cloud of the first object, and extracting global features of a second object based on a second point cloud of the second object;
[0007] Converting the global features of the first object into local features of the first object, and fusing the local features of the first object with the global features of the second object to obtain a first fused feature;
[0008] Reconstructing a point cloud based on the first fusion feature to obtain a third point cloud, wherein the structure of the third point cloud is the same as the structure of the second object;
[0009] Based on the first point cloud of the first object and the third point cloud, a recognition result of the first point cloud of the first object is determined.
[0010] The present invention provides a point cloud recognition device, comprising:
[0011] an extraction module, configured to extract global features of the first object based on a first point cloud of the first object, and to extract global features of the second object based on a second point cloud of the second object;
[0012] a fusion module, configured to convert the global features of the first object into local features of the first object, and fuse the local features of the first object with the global features of the second object to obtain a first fused feature;
[0013] a reconstruction module, configured to reconstruct a point cloud based on the first fusion feature to obtain a third point cloud, wherein the structure of the third point cloud is the same as that of the second object;
[0014] A determination module is configured to determine a recognition result of the first point cloud of the first object based on the first point cloud of the first object and the third point cloud.
[0015] An embodiment of the present application provides an electronic device, comprising:
[0016] a memory for storing computer-executable instructions or computer programs;
[0017] The processor is used to implement the point cloud recognition method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0018] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the point cloud recognition method provided in the embodiment of the present application when executed by a processor.
[0019] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the point cloud recognition method provided in the embodiment of the present application is implemented.
[0020] The embodiments of the present application have the following beneficial effects:
[0021] In the point cloud recognition method provided in the embodiment of the present application, the global features of the first object can be extracted based on the first point cloud, and the global features of the second object can be extracted based on the second point cloud. The global features of the first object can be converted into local features of the first object, and the local features of the first object and the global features of the second object can be fused to obtain a first fused feature. Based on the first fused feature, a third point cloud with the same structure as the second object can be obtained, and then the recognition result of the first point cloud can be determined based on the first point cloud and the third point cloud.
[0022] In the present application, the third point cloud is obtained based on the local features of the first object and the global features of the second object, which is equivalent to the third point cloud including both the local information of the first point cloud and the overall information of the second point cloud. That is to say, the information included in the third point cloud is more comprehensive, which can improve the accuracy of the recognition results of the first point cloud determined based on the first point cloud and the third point cloud. In addition, the present application does not require manpower to label the first point cloud and the second point cloud. Compared with the manual labeling method, it can improve the efficiency and accuracy of the recognition results of the point cloud. It can be seen that compared with the related technology, the present application can improve the efficiency of determining the recognition results of the point cloud, and can improve the accuracy of the recognition results of the point cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Schematic diagram of the structure of the point cloud recognition system provided in the embodiment of the present application;
[0024] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0025] Figure 3 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 1 ;
[0026] Figure 4 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 2 ;
[0027] Figure 5 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 3 ;
[0028] Figure 6 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 4 ;
[0029] Figure 7 is a structural diagram of the second model provided in an embodiment of the present application;
[0030] Figure 8 Schematic diagram of the second model training process provided in the embodiment of the present application;
[0031] Figure 9 This is a schematic diagram of the process of using the second model provided in the embodiment of the present application;
[0032] Figure 10 Schematic diagram of a point cloud provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0034] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0035] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0036] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0037] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0038] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0039] The embodiments of the present application provide a point cloud recognition method, a point cloud recognition device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the efficiency of determining the recognition results of the point cloud, and can improve the accuracy of determining the recognition results of the point cloud.
[0040] See also Figure 1 , Figure 1 is a schematic diagram of the structure of the point cloud recognition system provided in an embodiment of the present application. Figure 1 The point cloud recognition system 100 shown is used to support a point cloud recognition application. The terminal 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0041] In response to a point cloud recognition instruction triggered by a graphical interface of terminal 400, terminal 400 may send a point cloud recognition request to server 200. Server 200 is configured to receive the point cloud recognition request sent by terminal 400. After receiving the point cloud recognition request, server 200 may extract global features of the first object based on the first point cloud of the first object, and extract global features of the second object based on the second point cloud of the second object.
[0042] Furthermore, server 200 may convert the global features of the first object into local features of the first object, fuse the local features of the first object with the global features of the second object to obtain a first fused feature, and reconstruct a point cloud based on the first fused feature to obtain a third point cloud, where the structure of the third point cloud is the same as that of the second object. Based on the first point cloud and the third point cloud of the first object, a recognition result of the first point cloud of the first object is determined.
[0043] In the present application, the third point cloud is obtained based on the local features of the first object and the global features of the second object, which is equivalent to the third point cloud including both the local information of the first point cloud and the overall information of the second point cloud. That is to say, the information included in the third point cloud is more comprehensive, which can improve the accuracy of the recognition results of the first point cloud determined based on the first point cloud and the third point cloud. In addition, the present application does not require manpower to label the first point cloud and the second point cloud. Compared with the manual labeling method, it can improve the efficiency and accuracy of the recognition results of the point cloud. It can be seen that compared with the related technology, the present application can improve the efficiency of determining the recognition results of the point cloud, and can improve the accuracy of the recognition results of the point cloud.
[0044] The electronic device provided in the embodiment of the present application is described below. The electronic device that implements the point cloud recognition method in the embodiment of the present application can be a terminal, a server, or a combination of the two. The terminal can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and car terminals.
[0045] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server can be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.
[0046] See also Figure 2 , Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 2The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0047] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0048] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0049] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0050] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0051] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0052] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0053] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);
[0054] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0055] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0056] In some embodiments, the point cloud recognition device provided in the embodiments of the present application can be implemented in a software manner. Figure 2 A point cloud recognition device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an extraction module 4551, a fusion module 4552, a reconstruction module 4553, and a determination module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0057] In other embodiments, the point cloud recognition device provided in the embodiments of the present application can be implemented in hardware. As an example, the point cloud recognition device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the point cloud recognition method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0058] The following describes the point cloud recognition method provided by the embodiment of the present application. As mentioned above, the electronic device that implements the point cloud recognition method of the embodiment of the present application can be a terminal, a server, or a combination of the two. Figure 3 , Figure 3 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 1 , combined with Figure 3 The steps shown, and taking the electronic device being a server as an example, illustrate the point cloud recognition method provided in an embodiment of the present application.
[0059] In step 101 , global features of the first object are extracted based on a first point cloud of the first object, and global features of the second object are extracted based on a second point cloud of the second object.
[0060] When the user needs to perform point cloud recognition, in response to the point cloud recognition instruction triggered by the display interface of the server, the server can extract the global features of the first object based on the first point cloud of the first object, and extract the global features of the second object based on the second point cloud of the second object.
[0061] The first object and the second object are described below, where the first object and the second object are only used to distinguish the two as different objects. The following description is combined with the objects. The objects can be physical objects or virtual objects. Physical objects are real objects, people or animals. Virtual objects are various objects or people that can interact in a virtual scene.
[0062] An object has its own shape and volume in a real scene or a virtual scene, and occupies a portion of the space in the scene (corresponding to the real scene or the virtual scene). For example, the first object can be a person, an animal, a plant, a bucket, a wall, a stone, etc. The type of the second object can be the same as that of the first object, or it can be a different type from the first object. The first object and the second object can also be different postures of the same object. For example, the first object can be a person in a first posture, and the second object can be a person in a second posture. The specific settings can be made according to actual usage requirements.
[0063] In some embodiments, see Figure 4 , Figure 4 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 2 , combined with Figure 4 The steps shown are Figure 3 Step 101 of “extracting global features of the first object based on the first point cloud of the first object” is described below.
[0064] In step 1011 , the position information of each point in the first point cloud is determined in a preset three-dimensional coordinate system.
[0065] In some embodiments, the preset three-dimensional coordinate system can be a world coordinate system. For the first point cloud, the first point cloud includes multiple points. For the convenience of distinction, each point in the first point cloud can be called a first point, that is, the first point cloud includes multiple first points. The position information of each first point in the first point cloud can be determined in the preset three-dimensional coordinate system. The position information is the three-dimensional position coordinates of each first point in the first point cloud.
[0066] In step 1012 , for each point in the first point cloud, a first feature and a direction vector of the point are determined based on the position information, and a second feature of the point is determined based on the first feature and the direction vector.
[0067] In some embodiments, the position information is input into a linear layer to perform a linear transformation on the position information, and a linear feature can be obtained. The linear feature is the feature output by the linear layer, and then the linear feature can be input into a nonlinear layer. The nonlinear layer can obtain the first feature and direction vector of the point through an activation function, wherein the first feature of the point can be represented by q, and the direction vector of the point can be represented by k.
[0068] After obtaining the first feature and direction vector of the point, the second feature of the point can be determined based on the first feature and the direction vector. In some embodiments, the angle between the first feature and the direction vector can be determined. If the angle is less than a preset angle threshold, it means that the direction of the first feature is similar to or the same as the direction of the direction vector. Therefore, the first feature can be used as the second feature of the point. The preset angle threshold can be set according to actual usage requirements. For example, the preset angle threshold can be 90 degrees.
[0069] If the angle is greater than or equal to the preset angle threshold, it means that the direction of the first feature is opposite to the direction of the direction vector. Therefore, the projection of the first feature on the direction vector can be determined, and the second feature of the point can be determined based on the projection of the first feature on the direction vector.
[0070] In some embodiments, the square of the modulus of the direction vector can be determined, and then the ratio of the dot product of the first feature and the direction vector to the square of the modulus of the direction vector can be used as a first ratio. The first ratio is multiplied by the direction vector to obtain a first product. The first product is subtracted from the first feature to obtain a second feature. The second feature can be the inverse of the projection of the first feature on the direction vector.
[0071] In some embodiments, the dot product of the first feature and the direction vector can also be determined. When the dot product of the first feature and the direction vector is greater than or equal to 0, the second feature of the first point is determined to be the first feature. When the dot product of the first feature and the direction vector is less than 0, the square of the modulus of the direction vector can be determined, and then the ratio of the dot product of the first feature and the direction vector to the square of the modulus of the direction vector can be used as the first ratio. The first ratio is multiplied by the direction vector to obtain the first product. The first feature is subtracted from the first product to obtain the second feature. For the case where the dot product of the first feature and the direction vector is less than 0, the second feature can be the opposite of the projection of the first feature on the direction vector. Through the above process, when the linear feature is rotated, the direction vector will also rotate, and the nonlinear layer can be used to ensure that the second feature has rotational equivariance.
[0072] Rotational equivariance can be called the rotation group equivariance of three-dimensional space (SO(3)-equivariant). Rotational equivariance means that if a point cloud undergoes a rotation transformation in the SO(3) group, the output of the transformation will also rotate in the same way. When processing point clouds, rotational equivariance can preserve the rotational symmetry of the point cloud.
[0073] In step 1013 , an average value of the second feature of each point in the first point cloud is determined, and the average value is used as the global feature of the first object.
[0074] In some embodiments, after obtaining the second feature of each first point in the first point cloud, the second feature of each first point in the first point cloud can be input into a pooling layer. The pooling layer can determine the average value of the second feature of each first point in the first point cloud, and then the average values can be spliced, and the spliced average values can be used as the global feature of the first object.
[0075] After determining the global features of the first object, to further ensure that the global features of the first object are rotationally equivariant, in some embodiments, a first matrix can be constructed based on the global features of the first object. The first matrix is a Gram matrix, and the first matrix is used as the global features of the first object that are rotationally equivariant. Through steps 1011 to 1013 above, global features of the first object that are rotationally equivariant can be obtained.
[0076] In this application, global features are global feature representations of point clouds. Global features can also be called global feature descriptors. Global features can represent the overall features of point clouds. Local features are local feature representations of point clouds. Local features can also be called local feature descriptors. Local features can represent the characteristics of the geometric properties of a single point in a point cloud or an area composed of multiple adjacent points. Local features are associated with the positions of points in the point cloud, that is, local features can represent the geometric details of the associated positions.
[0077] In some embodiments, the local features of the first object are obtained by converting the global features of the first object by the first model. Before converting the global features of the first object into the local features of the first object, the first model can be trained. The training method of the first model is described below.
[0078] In some embodiments, the first model may include a first feature extraction layer, a first feature conversion layer and a first feature fusion layer, wherein the first feature extraction layer corresponds to the above-mentioned linear layer, the first feature conversion layer corresponds to the above-mentioned nonlinear layer, the first feature fusion layer corresponds to the above-mentioned pooling layer, and the first feature fusion layer may also correspond to the above-mentioned pooling layer and the invariant layer used to construct the first matrix.
[0079] By first extracting the third feature of each sample point in the fourth point cloud of the first object sample, for each sample point, the third feature is feature-converted by the first feature conversion layer of the first model to obtain the fourth feature of the sample point, wherein the method of performing feature conversion on the third feature to obtain the fourth feature of the sample point can refer to the method for obtaining the fourth feature of the sample point by performing feature conversion on the third feature. Figure 4 The description of step 102 is omitted here.
[0080] After obtaining the fourth feature of each sample point, the fourth feature of the sample point can be input into the first feature fusion layer. The first feature fusion layer can determine the average value of the fourth feature of each sample point, and then the average value of the fourth feature of each sample point can be spliced to obtain the fifth feature. Based on the fifth feature and the sample label of the first object sample, the model parameters of the first model are updated.
[0081] In some embodiments, the global features of the second object can be extracted based on the second point cloud of the second object. The second point cloud includes multiple points. For the convenience of distinction, each point in the second point cloud can be called a fourth point, that is, the second point cloud includes multiple fourth points. In a preset three-dimensional coordinate system, the position information of each fourth point in the second point cloud is determined. For each fourth point in the second point cloud, the first feature and direction vector of the fourth point are determined based on the position information, and the second feature of the fourth point is determined based on the first feature of the fourth point and the direction vector of the fourth point. Then, the average value of the second feature of each fourth point in the second point cloud can be determined, and the average value is used as the global feature of the second object. Specifically, Figure 4 The corresponding contents are replaced equivalently and are not described in detail here. In this way, the global features of the second object with rotational equivariance can be obtained.
[0082] Continue to see Figure 3In step 102, the global features of the first object are converted into local features of the first object, and the local features of the first object and the global features of the second object are fused to obtain a first fused feature.
[0083] In some embodiments, the local features of the first object include features of each first point in the first point cloud, and edge features of each first point in the first point cloud are extracted, where the edge features are used to describe local geometric characteristics of each first point in the point cloud.
[0084] After extracting the edge features of each first point in the first point cloud, the edge features of each first point in the first point cloud can be fused with the global features of the first object to obtain the features of each first point in the first point cloud, which is equivalent to converting the global features of the first object into local features of the first object.
[0085] In some embodiments, the global features of the first object have rotational variability, and converting the global features of the first object into local features of the first object is equivalent to converting the global features of the first object with rotational variability into local features. The local features can represent the local semantics and geometric information of the first object. The local features of the first object have rotational invariance, and the rotational invariance is SO(3) invariance. The rotational invariance means that no matter how the point cloud is rotated, the output is the same.
[0086] The following describes a method for fusing edge features and global features. In some embodiments, for each first point edge feature, the edge feature can be spliced with the global feature of the first object to obtain the feature of the first point.
[0087] In some embodiments, for each edge feature of the first point, the edge feature of the first point can be multiplied by the first weight to obtain a first product, and the global feature of the first object can be multiplied by the second weight to obtain a second product. Then, the first product and the second product can be added to obtain the feature of the first point.
[0088] In some embodiments, a multilayer perceptron (MLP) can be used to learn the nonlinear mapping of the features of the first point from the edge features of the first point and the global features of the first object. Of course, other machine learning models can also be used according to actual usage requirements, which is not specifically limited here.
[0089] In some embodiments, the local features of the first object include features of multiple first points in the first point cloud. The multiple first points in the first point cloud can be divided to obtain groups consisting of a preset number of first points. For each group, the distance between any two first points is less than a preset distance threshold.
[0090] For each group, the edge features of each first point included in the group can be extracted, and then the edge features of the first points included in the group can be spliced to obtain the edge features of the group. After extracting the edge features of each group in the first point cloud, the edge features of each group in the first point cloud can be fused with the global features of the first object respectively to obtain the features of each group in the first point cloud, which is equivalent to converting the global features of the first object into local features of the first object.
[0091] The way of fusing the edge features of the group with the global features of the first object can be equivalent to the way of fusing the edge features of the above points with the global features of the first object, which will not be described in detail here.
[0092] In the present application, the local feature can be the feature of a point in the point cloud, or it can be the feature of a preset number of points in the point cloud, and the distance between any two points in the preset number of points is less than the preset distance threshold. In comparison, when the local feature is the feature of a preset number of points in the point cloud, it can better express the local characteristics of the point cloud. The local feature is used to reconstruct the third point cloud, which is equivalent to obtaining a more accurate third point cloud, making the information of the third point cloud more comprehensive, thereby facilitating the subsequent improvement of the accuracy of point cloud recognition.
[0093] When the local feature is the feature of a point in the point cloud, there is no need to fuse the features of multiple points when determining the local feature, which can save computing resources and improve the efficiency of determining the local feature. The local feature is used to reconstruct the third point cloud, which is equivalent to being able to reconstruct the third point cloud more quickly, thereby improving the efficiency of point cloud recognition.
[0094] In some embodiments, after determining the local features of the first object, see Figure 5 , Figure 5 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 3 , combined with Figure 5 The steps shown are Figure 3 The step 102 of “fusing the local features of the first object and the global features of the second object to obtain a first fused feature” is illustrated for explanation.
[0095] In step 1021 , based on the type of each point in the first point cloud, a weight of a feature of the point is determined.
[0096] In some embodiments, the first points in the point cloud can be classified according to the shape of the first point cloud to obtain the type of each first point in the first point cloud. The types may include edge points, plane points, corner points, intersection points, vertices, etc., where edge points are points located at the edge of the object, corner points are points where the surface of the object changes sharply, and the surface of the object changes sharply refers to the position where the curvature of the object surface is greater than the first curvature threshold.
[0097] A plane point is a point located in a relatively flat area of an object. A relatively flat area refers to a position where, for a certain object, the area on the surface of the object is smaller than the second curvature threshold. The first curvature threshold is greater than the second curvature threshold. The first curvature threshold and the second curvature threshold are associated with the object. The first curvature threshold (or second curvature threshold) corresponding to different objects may be the same or different. The first curvature threshold and the second curvature threshold may be set according to actual usage requirements.
[0098] In some embodiments, a correspondence between the candidate types of points and the candidate weights of features can be pre-set. After determining the type of the first point, the weight of the feature of the first point can be determined based on the type of the first point from the correspondence between the candidate types of points and the candidate weights of features.
[0099] In some embodiments, the type of point can also be divided into points with different importance intervals according to actual usage needs. The importance of the point can be set according to actual usage needs. The first point in the first point cloud has a corresponding importance. The importance can be determined in combination with one or more of the curvature difference, normal vector difference, and spatial distance of the point. The curvature difference can be the difference between the curvature of the point and the curvature of other points in the first range corresponding to the point. The normal vector difference can be the difference between the normal vector of the point and the normal vector of other points in the second range corresponding to the point. The spatial distance can be other spatial distances between the point and the point in the third range.
[0100] In some embodiments, the type of a point can be determined based on the interval of the point's importance. The importance of a point is positively correlated with its weight, that is, the higher the importance of a point, the greater its weight, and the lower the importance of a point, the smaller its weight.
[0101] In step 1022 , based on the determined weight, the features of each point in the first point cloud are fused to obtain a second fused feature.
[0102] After determining the weight of each first point, the weight of the first point can be multiplied by the feature of the first point to obtain a second product, and the second product can be spliced to obtain a second fusion feature, which is equivalent to fusing the features of each first point in the first point cloud to obtain the second fusion feature.
[0103] In step 1023, the second fused feature is fused with the global feature of the second object to obtain a first fused feature.
[0104] In some embodiments, the second fused feature can be concatenated with the global feature of the second object, thereby fusing the second fused feature with the global feature of the second object to obtain the first fused feature. In some embodiments, the second fused feature has a corresponding third weight, and the global feature of the second object has a corresponding fourth weight. The product of the second fused feature and the third weight can be used as the third product, the product of the global feature of the second object and the fourth weight can be used as the fourth product, and the sum of the third product and the fourth product can be used as the first fused feature.
[0105] The first fusion feature in this application is obtained by fusing the local features of the first object and the global features of the second object. That is to say, the first fusion feature includes both the local features of the first object and the global features of the second object. The first fusion feature can represent both the local information of the first point cloud and the overall information of the second point cloud.
[0106] In step 103, point cloud reconstruction is performed based on the first fusion feature to obtain a third point cloud.
[0107] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the process of the point cloud recognition method provided in the embodiment of the present application. Figure 4 , combined with Figure 6 The steps shown are Figure 3 Step 103 is shown for explanation.
[0108] In step 1031 , the first fused feature is decoded to obtain a three-dimensional voxel grid corresponding to the second object.
[0109] In some embodiments, the first fusion feature can be decoded by a decoder to obtain a three-dimensional voxel grid corresponding to the second object, wherein the three-dimensional voxel grid can be used to represent data in a three-dimensional space. The three-dimensional voxel grid includes multiple voxels. Voxels are basic units that constitute the three-dimensional voxel grid. Voxels can be used to represent a position in the three-dimensional space.
[0110] The first fusion feature can represent both local information of the first point cloud and overall information of the second point cloud. The three-dimensional voxel grid obtained according to the first fusion feature includes local information of the first point cloud and overall information of the second point cloud.
[0111] In step 1032 , sampling is performed from the voxels included in the three-dimensional voxel grid to obtain sampling points corresponding to a plurality of voxels.
[0112] In some embodiments, sampling may be performed from voxels included in a three-dimensional voxel grid based on a preset sampling strategy, thereby obtaining sampling points corresponding to a plurality of voxels, where one voxel corresponds to one sampling point.
[0113] The sampling strategy can be a centroid sampling strategy. After obtaining the 3D voxel grid, the centroid of each voxel can be determined, and the centroid is used as the sampling point of the voxel. The sampling strategy can also be a proximity sampling strategy. After obtaining the 3D voxel grid, the point with the smallest distance to the voxel center can be determined as the sampling point for each voxel. The sampling strategy can also be a random sampling strategy, that is, a point is randomly selected as the sampling point for each voxel.
[0114] In some embodiments, the sampling points can be determined by at least two of the body centroid sampling strategy, the adjacent point sampling strategy, and the random sampling strategy, which is equivalent to at least two sampling points corresponding to each voxel. For each voxel, the sampling point of the voxel can be determined based on the at least two sampling points corresponding to the voxel.
[0115] In the case of two sampling points corresponding to a voxel, the two sampling points can be connected into a straight line, and then the sampling point corresponding to the voxel can be selected on the straight line. In the case of three sampling points corresponding to a voxel, the three sampling points can be connected in sequence to obtain a corresponding graph, and then the sampling point corresponding to the voxel can be selected in the graph.
[0116] In step 1033 , the sampling points corresponding to the multiple voxels are combined to obtain a third point cloud.
[0117] In some embodiments, after the sampling points corresponding to the voxels are determined, the sampling points corresponding to the multiple voxels may be combined based on the relative positional relationship between the voxels, thereby obtaining a third point cloud, wherein one voxel corresponds to one sampling point.
[0118] When the sampling points corresponding to the voxels are obtained only through the body centroid sampling strategy, the sampling points obtained by the body centroid sampling strategy can better represent the overall shape characteristics of the third point cloud. When the sampling points corresponding to the voxels are obtained only through the adjacent point sampling strategy, the sampling points obtained by the adjacent point sampling strategy can better represent the detailed information of the third point cloud.
[0119] When the sampling points corresponding to the voxels are obtained through the body centroid sampling strategy and the adjacent point sampling strategy, the overall shape characteristics of the third point cloud and the detail information of the third point cloud can be balanced, thereby reconstructing a more accurate third point cloud.
[0120] The third point cloud is obtained by reconstructing the point cloud based on the first fusion feature. The first fusion feature can represent both the local information of the first point cloud and the overall information of the second point cloud. This is equivalent to the third point cloud including both the local information of the first point cloud and the overall information of the second point cloud. In other words, the information included in the third point cloud is more comprehensive.
[0121] The third point cloud includes overall information of the second point cloud, and the structure of the third point cloud is the same as the structure of the second object. The structure of the third point cloud is the same as the structure of the second object, which means that the first similarity between the outline of the object formed by the third point cloud in the three-dimensional space and the outline of the object formed by the second point cloud in the three-dimensional space is greater than the first similarity threshold.
[0122] The structure of the third point cloud is identical to the structure of the second object, and further refers to the second similarity between the relative positional relationship of the third point in the third point cloud in three-dimensional space and the relative positional relationship of the second point in the second point cloud in three-dimensional space being greater than a second similarity threshold. The first and second similarity thresholds can be set based on actual usage requirements.
[0123] Continue to see Figure 3 In step 104 , based on the first point cloud and the third point cloud of the first object, a recognition result of the first point cloud of the first object is determined.
[0124] In some embodiments, the first point cloud includes multiple first points, and the third point cloud includes multiple second points. For each first point, the third point among the multiple second points that is closest to the first point can be determined, which is equivalent to establishing a correspondence between the first point and the third point. Then, based on the correspondence between each first point in the first point cloud and the corresponding third point, the recognition result of the first point cloud of the first object can be determined.
[0125] Among them, the correspondence between the first point and the third point is the correspondence between the first point cloud and the third point cloud. The structure of the third point cloud is the same as the structure of the second object. The correspondence between the first point cloud and the third point cloud is also the correspondence between the first point cloud and the second point cloud.
[0126] In some embodiments, in order to obtain a more accurate recognition result of the first point cloud of the first object, for each first point, the third point closest to the first point among multiple second points can be determined, and then it can be determined whether the distance between the first point and the third point is less than a preset threshold. If the distance between the first point and the third point is less than the preset threshold, it means that the distance between the first point and the third point is small. Therefore, a correspondence between the first point and the third point can be established.
[0127] If the distance between the first point and the third point is greater than a preset threshold, it indicates that the distance between the first point and the third point is large, and a correspondence relationship between the first point and the third point may not be established. Furthermore, a recognition result of the first point cloud of the first object may be determined based on the correspondence relationship between each first point and the corresponding third point in the first point cloud.
[0128] In the present application, the correspondence between the first point and the third point is established only when the distance between the first point and the third point is less than a preset threshold value. This is equivalent to screening the first point and the third point based on the preset threshold value in the process of establishing the correspondence between the first point and the third point, so that a more accurate correspondence between the first point and the third point can be obtained. The correspondence is used to determine the recognition result of the first point cloud of the first object, so that a more accurate recognition result of the first point cloud of the first object can be obtained.
[0129] The recognition result of the first point cloud of the first object is explained below. In some embodiments, the first object and the second object may both be known objects, and the structural information of the second object is known. The structural information is used to describe the structure of the second object. Accordingly, the structural information of the second point cloud of the second object is known, and the structure of the third point cloud is the same as that of the second object, that is, the structural information of the third point cloud is known.
[0130] The structural information includes the correspondence between areas and shapes. After determining the correspondence between the first point and the third point, the third point corresponding to the first area in the third point cloud is determined based on the structural information. The first area corresponds to the first shape. The first shape in the first point cloud can be determined in combination with the correspondence between the first point and the third point.
[0131] For example, if the second object includes a sphere, and accordingly, there is a first region 1 corresponding to the sphere in the third point cloud, the sphere in the first point cloud can be determined based on the correspondence between the first point and the third point. In this application, the shape in the structural information can be a basic geometric shape such as a plane, a cylinder, or a sphere, and the specific shape can be determined based on actual conditions.
[0132] In some embodiments, identifying the first shape in the first point cloud can be used to perform hierarchical analysis on the first point cloud. For example, the first object is a building. Through the point cloud recognition method provided in the embodiment of the present application, the windows, doors, beams and other structures (corresponding geometric shapes) included in the building can be identified.
[0133] In some embodiments, after determining the geometric shape in the first point cloud of the first object, point cloud repair can be performed on the first point cloud. For example, if the first object is an industrial part, after determining the geometric shape in the first point cloud, it can be determined whether there are holes in the area corresponding to the geometric shape. If holes exist, the holes can be repaired.
[0134] In some embodiments, after determining the geometric shape in the first point cloud of the first object, the first object can be optimized. For example, if the first object is a building, after identifying the windows, doors, beams and other structures included in the building, the building can be simplified or the details can be adjusted according to actual usage requirements.
[0135] In some embodiments, the structural information includes the correspondence between regions and topological structures. The topological structures may include connected regions, holes, branches, etc., wherein a connected region refers to a region in a point cloud where any two points can be connected by a series of continuous points.
[0136] Holes are empty spaces in a point cloud—areas within connected regions that are not covered by the point cloud. Branches are secondary structures in a point cloud that extend from a primary structure. For example, in a point cloud of a tree, the trunk is the primary structure, while the branches are the offshoots extending from the trunk.
[0137] After determining the correspondence between the first point and the third point, the third point corresponding to the second area in the third point cloud is determined based on the structural information. The second area corresponds to the first topological structure. The first topological structure in the first point cloud can be determined in combination with the correspondence between the first point and the third point.
[0138] In point cloud segmentation scenarios, point clouds can be segmented into different sub-point clouds by analyzing their connected regions. Each sub-point cloud corresponds to a specific object. Holes can be used to repair the point cloud, further understanding the integrity and internal structure of the original object. Branches can be used for scene analysis.
[0139] In the present application, the third point cloud is obtained based on the local features of the first object and the global features of the second object, which is equivalent to the third point cloud including both the local information of the first point cloud and the overall information of the second point cloud. That is to say, the information included in the third point cloud is more comprehensive, which can improve the accuracy of the recognition results of the first point cloud determined based on the first point cloud and the third point cloud. In addition, the present application does not require manpower to label the first point cloud and the second point cloud. Compared with the manual labeling method, it can improve the efficiency and accuracy of the recognition results of the point cloud. It can be seen that compared with the related technology, the present application can improve the efficiency of determining the recognition results of the point cloud, and can improve the accuracy of the recognition results of the point cloud.
[0140] In this application, the correspondence between the first point and the third point is the correspondence between the first point cloud and the third point cloud. The first point cloud corresponds to the first object. The structure of the third point cloud is the same as the structure of the second object. The correspondence between the first point cloud and the third point cloud is also the correspondence between the first object and the second object.
[0141] After determining the correspondence between the first object and the second object, if the first object is unknown, the first object can be identified based on the correspondence between the first object and the second object to obtain an identification result for the first object. For example, in a robot navigation scenario, the first point cloud of the first object can be an environmental point cloud collected by the robot in real time, which is equivalent to the first object being an object collected by the robot in real time, and the second point cloud of the second object can be an environmental point cloud known to the robot, which is equivalent to the second object being an object known to the robot. After determining the correspondence between the first object and the second object, the identification result for the first object can be determined based on the correspondence between the first object and the second object and a map known to the robot including the second object.
[0142] The recognition result of the first object can be the position of the first object in the map. The robot can determine the current position of the robot based on the position of the first object in the map and the information of the image acquisition device that collects the first point cloud. Then, based on the current position of the robot and the map including the second object, the robot can perform path planning and avoid obstacles.
[0143] In some embodiments, the process of obtaining the third point cloud based on the first point cloud and the second point cloud in the embodiments of the present application can be implemented by a second model. The structure of the second model and the training method of the second model are described below.
[0144] In some embodiments, see Figure 7 , Figure 7 701 is a schematic diagram of the structure of the second model provided in an embodiment of the present application, wherein the second model includes a second feature extraction layer 701, a second feature conversion layer 702, a second feature fusion layer 703, and a feature decoding layer 704. In some embodiments, the second feature extraction layer 701 may be the first model.
[0145] During the model training phase, see Figure 8 , Figure 8 8 is a flow chart of the second model training process provided by an embodiment of the present application. The second feature extraction layer can be used to extract local features 802 of the second sample object from the fifth point cloud 801 of the second sample object. The second feature extraction layer can also be used to extract local features 805 of the third sample object from the sixth point cloud 803 of the third sample object. The second feature extraction layer can also be used to extract global features 804 of the third sample object from the sixth point cloud 803 of the third sample object.
[0146] The second feature extraction layer may include an edge convolution layer, a nonlinear layer, a pooling layer, and an invariant layer. The edge convolution layer corresponds to the above-mentioned linear layer. The second feature extraction layer can be a combination of a Vector Neurons (VN) framework and a Dynamic Graph Convolutional Neural Network (DGCNN). The Vector Neurons (VN) framework can expand one-dimensional scalar neurons to three-dimensional vector neurons, which can improve the model's robustness and generalization ability to point cloud rotation.
[0147] exist Figure 8 In FIG, a local feature extraction layer 810, a local feature fusion layer 820 and a pooling layer 830 composed of multiple edge convolution layers are shown, wherein the nonlinear layer and the invariant layer belong to the vector neuron framework for maintaining rotation equivariance. Figure 8 The fifth point cloud 801 is input to the local feature extraction layer 810 to obtain the local features 802 of the second sample object.
[0148] The sixth point cloud 803 is input to the local feature extraction layer 810 to obtain the local features 805 of the third sample object. The different edge convolution layers in the local feature extraction layer 810 can output features of different scales. The features of different scales can be fused through the local feature fusion layer 820, and the fused features are input to the pooling layer 830. Through the global average pooling operation, the global features 804 of the third sample object can be obtained.
[0149] After obtaining the local features 802 of the second sample object and the local features 805 of the third sample object, in order to enable the second feature conversion layer 840 to learn local feature conversion, the local features 802 of the second sample object can be input to the second feature conversion layer 840, so that the local feature conversion function corresponding to the second feature conversion layer 840 can learn the transformation of each point in the second sample point cloud. In addition, the local features 805 of the third sample object can be input to the second feature conversion layer 840, so that the local feature conversion function corresponding to the second feature conversion layer 840 can learn the transformation of each point in the third sample point cloud.
[0150] For example, the local feature conversion function is {fθ}. For the local features of each point in the second sample point cloud, the local features of each point can be calculated using {fθ} to obtain {fθi} corresponding to the point, where i is the index of the point. After all points in the second sample point cloud are calculated, the {fθi} of multiple points can be combined to obtain the predicted local features of the second sample point cloud.
[0151] In some embodiments, the parameters of the local feature conversion function can be adjusted based on the difference between the predicted local features and the input local features, that is, the parameters of the second feature conversion layer 840 can be adjusted, so as to achieve learning the transformation of each point in the second sample point cloud.
[0152] During the model training process, the second feature conversion layer 840 is trained using local features, which can adapt to the local combination structure and semantic information of the point cloud input to the model. Compared with the use of global features to train the second feature conversion layer 840, it can learn a more accurate expression of local features, thereby obtaining a more accurate representation of the changes and characteristics of the layout area of the point cloud, and capturing the detailed features and semantic information of the local area of the point cloud. The above-mentioned second feature extraction layer and second feature conversion layer can belong to the encoding layer. The encoding layer is used to obtain the global features of a point cloud (i.e., the global features 804 of the third sample object), and the encoding layer is also used to obtain the local features of another point cloud (i.e., the local features 802 of the second sample object). During the training stage, the encoding layer is also used to obtain the local features of a point cloud (i.e., the local features 805 of the third sample object).
[0153] After obtaining the global features 804 of the third sample object and the local features 802 of the second sample object, the global features 804 of the third sample object and the local features 802 of the second sample object may be input into the second feature fusion layer 850 .
[0154] The second feature fusion layer 850 can fuse the global features 804 of the third sample object and the local features 802 of the second sample object to obtain a sample fusion feature, and input the sample fusion feature into the feature decoding layer 860. The feature decoding layer 860 can reconstruct the point cloud based on the sample fusion feature to obtain a fourth sample point cloud.
[0155] The structure of the fourth sample point cloud is the same as that of the third sample object. The fourth sample point cloud includes both the local information of the fifth point cloud 801 of the second sample object and the overall information of the sixth point cloud 803 of the third sample object. That is, the information included in the fourth sample point cloud is more comprehensive.
[0156] After obtaining the global features 804 of the third sample object and the local features 805 of the third sample object, the global features 804 of the third sample object and the local features 805 of the third sample object may be input into the second feature fusion layer 850 .
[0157] The second feature fusion layer 850 can fuse the global features 804 of the third sample object and the local features 805 of the third sample object to obtain a sample fusion feature, and input the sample fusion feature into the feature decoding layer 860. The feature decoding layer 860 can reconstruct the point cloud based on the sample fusion feature to obtain a fifth sample point cloud.
[0158] The structure of the fifth sample point cloud is the same as that of the third sample object. The fifth sample point cloud includes local information of the sixth point cloud 803 of the third sample object and also includes overall information of the sixth point cloud 803 of the third sample object.
[0159] In some embodiments, the structure of the fifth sample point cloud is the same as that of the third sample object, and a self-reconstruction loss can be calculated based on the fifth sample point cloud and the fifth point cloud of the third sample object. The structure of the fourth sample point cloud is the same as that of the third sample object, and a cross-reconstruction loss can be calculated based on the fourth sample point cloud and the sixth point cloud of the third sample object. The parameters of the second model are then adjusted based on the cross-reconstruction loss and the self-reconstruction loss. Training of the second model can be completed when the number of training times exceeds a preset number threshold, or when the sum of the cross-reconstruction loss and the self-reconstruction loss is less than a preset sum threshold.
[0160] Refer to formula (1), which is the method for calculating the self-reconstruction loss. SR is the self-reconstruction loss, λ1 is the first hyperparameter, λ2 is the second hyperparameter, and MES is used to calculate P1 and P 1-1 The mean square error between P1 and P 1-1 The bulldozer distance between P1 and P 1-1 A fifth sample point cloud is reconstructed based on the global features of the third sample object and the local features of the third sample object.
[0161] L SR =λ1MES(P1,P 1-1 )+λ2EMD(P1,P 1-1 ) (1)
[0162] See formula (2), which is the method for calculating the cross reconstruction loss, where L CR is the cross reconstruction loss, λ3 is the third hyperparameter, and CD is used to calculate P1 and P 2-1 The chamfer distance between the third sample point cloud and P 2-1 The fourth sample point cloud is reconstructed based on the global features of the third sample object and the local features of the second sample object.
[0163] L CR =λ3CD(P1,P 2-1 ) (2)
[0164] After the second model is trained, see Figure 9 , Figure 9 This is a flow chart of the second model used in the embodiment of the present application. The second feature extraction layer (corresponding to Figure 9 The dotted box in the figure) can be used to extract the global features 902 of the first object from the first point cloud 901 of the first object, and the second feature extraction layer can also be used to extract the global features 904 of the second object from the second point cloud 903 of the second object.
[0165] After obtaining the global features 902 of the first object, the first point cloud 901 of the first object may be input to the second feature conversion layer 910 , and the second feature conversion layer 910 may convert the global features 902 of the first object into local features of the first object.
[0166] The local features of the first object and the global features 904 of the second object are input to the second feature fusion layer 920. The second feature fusion layer 920 can fuse the local features of the first object and the global features 904 of the second object to obtain a first fused feature, and input the first fused feature to the feature decoding layer 930. The feature decoding layer 930 can reconstruct a point cloud based on the first fused feature to obtain a third point cloud.
[0167] In some embodiments, for the second sample object and the third sample object, the local features of the second sample object and the global features of the second sample object can be extracted, and the local features of the third sample object and the global features of the third sample object can also be extracted, and then the second model is trained in combination with the self-reconstruction loss and the cross-reconstruction loss. For details, please refer to the above description and will not be repeated here.
[0168] For example, see Figure 10 , Figure 10 is a schematic diagram of the point cloud provided in the embodiment of the present application. Figure 10 In the example, for the point cloud 1010 of the second sample object and the point cloud 1020 of the third sample object, a global feature 1001 of the second sample object and a local feature 1003 of the second sample object can be extracted based on the point cloud 1010 of the second sample object. A global feature 1002 of the third sample object and a local feature 1004 of the third sample object can be extracted based on the point cloud 1020 of the third sample object.
[0169] Based on the global features 1001 of the second sample object and the local features 1003 of the second sample object, a point cloud 1030 can be reconstructed. Based on the global features 1001 of the second sample object and the local features 1004 of the third sample object, a point cloud 1040 can be reconstructed. Based on the global features 1002 of the third sample object and the local features 1003 of the second sample object, a point cloud 1050 can be reconstructed. Based on the global features 1002 of the third sample object and the local features 1004 of the third sample object, a point cloud 1060 can be reconstructed. The second model can then be trained by combining the self-reconstruction loss and the cross-reconstruction loss.
[0170] In the present application, the second model includes a second feature conversion layer, which can convert the global features of the point cloud into local features, which is equivalent to being able to realize local feature transformation for the point cloud. By learning dynamic local feature transformation, the global feature descriptor (i.e., global feature) is mapped to the local feature descriptor (i.e., local feature), and a transformation with rotation invariance is formulated for each point.
[0171] The second model mentioned above combines self-reconstruction and cross-reconstruction methods, which can optimize the difference between the input point cloud and the reconstructed point cloud, and can map semantically corresponding points to similar local feature descriptors, thereby facilitating the subsequent establishment of dense point-by-point correspondences between point clouds.
[0172] The following continues to describe the exemplary structure of the point cloud recognition device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the point cloud recognition device 455 of the memory 450 may include:
[0173] An extraction module 4551 is configured to extract global features of a first object based on a first point cloud of the first object, and to extract global features of a second object based on a second point cloud of the second object;
[0174] The fusion module 4552 is configured to convert the global features of the first object into local features of the first object, and fuse the local features of the first object with the global features of the second object to obtain a first fused feature.
[0175] a reconstruction module 4553, configured to reconstruct a point cloud based on the first fusion feature to obtain a third point cloud, wherein the structure of the third point cloud is the same as that of the second object;
[0176] The determination module 4554 is further configured to determine a recognition result of the first point cloud of the first object based on the first point cloud of the first object and the third point cloud.
[0177] In some embodiments, the extraction module 4551 is also used to determine the position information of each point in the first point cloud in a preset three-dimensional coordinate system; for each point in the first point cloud, determine the first feature and direction vector of the point based on the position information, and determine the second feature of the point based on the first feature and the direction vector; determine the average value of the second feature of each point in the first point cloud, and use the average value as the global feature of the first object.
[0178] In some embodiments, the extraction module 4551 is also used to determine the angle between the first feature and the direction vector; if the angle is less than a preset angle threshold, the first feature is used as the second feature of the point; if the angle is greater than or equal to the preset angle threshold, the projection of the first feature on the direction vector is determined, and based on the projection of the first feature on the direction vector, the second feature of the point is determined.
[0179] In some embodiments, the local features of the first object include features of each point in the first point cloud, and the fusion module 4552 is further used to extract edge features of each point in the first point cloud; and fuse the edge features of each point in the first point cloud with the global features of the first object to obtain features of each point in the first point cloud.
[0180] In some embodiments, the fusion module 4552 is further used to determine the weight of the feature of each point in the first point cloud based on the type of the point; based on the determined weight, fuse the feature of each point in the first point cloud to obtain a second fused feature; and fuse the second fused feature with the global feature of the second object to obtain the first fused feature.
[0181] In some embodiments, the reconstruction module 4553 is further used to decode the first fusion feature to obtain a three-dimensional voxel grid corresponding to the second object, where the three-dimensional voxel grid includes multiple voxels; sample from the voxels included in the three-dimensional voxel grid to obtain sampling points corresponding to multiple voxels; and combine the sampling points corresponding to the multiple voxels to obtain the third point cloud.
[0182] In some embodiments, the first point cloud includes multiple first points, and the third point cloud includes multiple second points; the determination module 4554 is also used to determine, for each first point, the third point among the multiple second points that is closest to the first point, and if the distance between the first point and the third point is less than a preset threshold, establish a correspondence between the first point and the third point; based on the correspondence between each first point and the corresponding third point in the first point cloud, determine the recognition result of the first point cloud of the first object.
[0183] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform the point cloud recognition method described in the present invention.
[0184] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the point cloud recognition method provided by the embodiment of the present application, for example, Figure 3 The point cloud recognition method is shown.
[0185] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0186] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0187] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0188] By way of example, computer-executable instructions or computer programs may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0189] In the present application, the third point cloud is obtained based on the local features of the first object and the global features of the second object, which is equivalent to the third point cloud including both the local information of the first point cloud and the overall information of the second point cloud. That is to say, the information included in the third point cloud is more comprehensive, which can improve the accuracy of the recognition results of the first point cloud determined based on the first point cloud and the third point cloud. In addition, the present application does not require manpower to label the first point cloud and the second point cloud. Compared with the manual labeling method, it can improve the efficiency and accuracy of the recognition results of the point cloud. It can be seen that compared with the related technology, the present application can improve the efficiency of determining the recognition results of the point cloud, and can improve the accuracy of the recognition results of the point cloud.
[0190] In the present application, the second model includes a second feature conversion layer, which can convert the global features of the point cloud into local features, which is equivalent to being able to realize local feature transformation for the point cloud. By learning dynamic local feature transformation, the global feature descriptor (i.e., global feature) is mapped to the local feature descriptor (i.e., local feature), and a transformation with rotation invariance is formulated for each point.
[0191] The above-mentioned second model combines self-reconstruction and cross-reconstruction methods, which can optimize the difference between the input point cloud and the reconstructed point cloud, and can realize the mapping of semantic corresponding points to similar local feature descriptors, so as to facilitate the subsequent establishment of dense point-by-point correspondences between point clouds, thereby improving the efficiency of determining the recognition results of the point cloud, and improving the accuracy of the recognition results of the point cloud.
[0192] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A point cloud recognition method, characterized in that: The method comprises: Extracting global features of a first object based on a first point cloud of the first object, and extracting global features of a second object based on a second point cloud of the second object; Converting the global features of the first object into local features of the first object, and fusing the local features of the first object with the global features of the second object to obtain a first fused feature; Reconstructing a point cloud based on the first fusion feature to obtain a third point cloud, wherein the structure of the third point cloud is the same as the structure of the second object; Based on the first point cloud of the first object and the third point cloud, a recognition result of the first point cloud of the first object is determined.
2. The method according to claim 1, characterized in that The extracting the global features of the first object based on the first point cloud of the first object includes: Determining the position information of each point in the first point cloud in a preset three-dimensional coordinate system; For each point in the first point cloud, determining a first feature and a direction vector of the point based on the position information, and determining a second feature of the point based on the first feature and the direction vector; An average value of the second feature of each point in the first point cloud is determined, and the average value is used as a global feature of the first object.
3. The method according to claim 2, characterized in that The determining the second feature of the point based on the first feature and the direction vector includes: determining an angle between the first feature and the direction vector; If the angle is less than a preset angle threshold, the first feature is used as the second feature of the point; If the included angle is greater than or equal to the preset angle threshold, a projection of the first feature on the direction vector is determined, and a second feature of the point is determined based on the projection of the first feature on the direction vector.
4. The method according to claim 1, wherein The local features of the first object include features of each point in the first point cloud, and converting the global features of the first object into the local features of the first object includes: Extracting edge features of each point in the first point cloud; The edge features of each point in the first point cloud are fused with the global features of the first object to obtain the features of each point in the first point cloud.
5. The method according to claim 4, characterized in that The fusing the local features of the first object and the global features of the second object to obtain a first fused feature includes: Determining a weight of a feature of each point based on a type of each point in the first point cloud; Based on the determined weight, the features of each point in the first point cloud are fused to obtain a second fused feature; The second fused feature is fused with the global feature of the second object to obtain the first fused feature.
6. The method according to claim 1, wherein The step of reconstructing the point cloud based on the first fusion feature to obtain a third point cloud includes: decoding the first fused feature to obtain a three-dimensional voxel grid corresponding to the second object, the three-dimensional voxel grid including a plurality of voxels; Sampling from voxels included in the three-dimensional voxel grid to obtain sampling points corresponding to a plurality of voxels; The sampling points corresponding to the plurality of voxels are combined to obtain the third point cloud.
7. The method according to claim 1, characterized in that The first point cloud includes a plurality of first points, and the third point cloud includes a plurality of second points; The determining, based on the first point cloud of the first object and the third point cloud, a recognition result of the first point cloud of the first object includes: For each of the first points, determining a third point that is closest to the first point among the plurality of second points, and establishing a correspondence between the first point and the third point if the distance between the first point and the third point is less than a preset threshold; Based on the correspondence between each first point and a corresponding third point in the first point cloud, a recognition result of the first point cloud of the first object is determined.
8. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; A processor, configured to implement the point cloud recognition method according to any one of claims 1 to 7 when executing the computer-executable instructions or computer program stored in the memory.
9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the point cloud recognition method according to any one of claims 1 to 7 is implemented.
10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the point cloud recognition method according to any one of claims 1 to 7 is implemented.