3D Object Detection
By extracting point features in the three-dimensional object point cloud data and generating candidate object feature representations, the correlation is determined using the attention module, and the problems of poor detection effect of the existing three-dimensional object detection methods and relying on manual rules are solved, and high-accuracy three-dimensional object detection is achieved.
Patent Information
- Application Number
- CN202110212544.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-02-25
AI Technical Summary
The existing three-dimensional object detection method is difficult to directly apply to three-dimensional scenes, and the point aggregation-based method relies on manual rules, resulting in poor detection results and the inability to fully utilize the information in point cloud data.
By extracting the feature representations of multiple points from the point cloud data of the three-dimensional object, including position information and appearance characteristics, the initial feature representation of the candidate three-dimensional object is generated, and the autocorrelation between candidate objects and cross-correlation between points and objects is determined through the attention module, and the detection results of the three-dimensional object are generated.
It realizes that locating and identifying three-dimensional objects in a three-dimensional scene with high accuracy without point aggregation, improving detection accuracy and saving computing resources.
Smart Images

Figure CN114973231B_ABST
Abstract
Description
Background Art
[0001] 3D object detection is used to locate and identify 3D objects contained in a 3D scene, such as pedestrians, vehicles, objects, etc. Currently, 3D object detection plays an important role in applications such as autonomous driving, robot control, and augmented reality. In conventional 3D object detection methods, irregular and sparse point clouds are usually used to describe the 3D scene and the 3D objects in the scene. Therefore, it is difficult to directly apply 2D object detection methods based on regular grids to 3D object detection. Based on this, there is a need for methods that can perform object detection for 3D scenes. Summary of the Invention
[0002] According to an implementation of the present disclosure, a solution for 3D object detection is proposed. In this solution, a feature representation of multiple points is extracted from the point cloud data of the 3D object, and the feature representation of each point includes the position information and appearance features of the point. Based on the feature representations of multiple points, an initial feature representation of a set of candidate 3D objects is determined. The initial feature representation of each candidate 3D object includes the position feature and appearance feature of the candidate 3D object. Based on the feature representations of multiple points and the initial feature representations of a set of candidate 3D objects, by determining the self-correlation between a set of candidate 3D objects and the cross-correlation between multiple points and a set of candidate 3D objects, the detection result of the 3D object is generated. In this way, this solution can locate and identify 3D objects in a 3D scene only based on the correlation between each point in the point cloud and the candidate 3D objects and the correlation between the candidate 3D objects without aggregating points into candidate objects.
[0003] The Summary of the Invention section is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings
[0004] Figure 1 A block diagram of a computing device capable of implementing multiple implementations of the present disclosure is shown;
[0005] Figure 2 A system architecture diagram for 3D object detection according to an implementation of the present disclosure is shown;
[0006] Figure 3 A schematic diagram showing the process of generating a first set of candidate detection results using a first attention module according to an implementation of the present disclosure is shown;
[0007] Figure 4 A schematic diagram showing the process of determining a detection result from at least one set of candidate detection results according to an implementation of the present disclosure; and
[0008] Figure 5 A flowchart of a method for three-dimensional object detection according to an implementation of the present disclosure is shown;
[0009] In these figures, the same or similar reference signs are used to denote the same or similar elements. Detailed implementation manners
[0010] The present disclosure will now be described with reference to several example implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, rather than implying any limitation on the scope of the present disclosure.
[0011] As used herein, the term "comprising" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The terms "one implementation" and "an implementation" are to be construed as "at least one implementation". The term "another implementation" is to be construed as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included hereinafter.
[0012] As used herein, a "neural network" is capable of processing inputs and providing corresponding outputs, and generally includes an input layer and an output layer as well as one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications generally include many hidden layers, thus extending the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also referred to as processing nodes or neurons), and each node processes the input from the previous layer. A CNN is a type of neural network that includes one or more convolutional layers for performing convolution operations on respective inputs. A CNN can be used in various scenarios and is particularly suitable for processing image or video data. In this article, the terms "neural network", "network" and "neural network model" may be used interchangeably.
[0013] As described above, since three-dimensional scenes and three-dimensional objects are usually described using irregular and sparse point clouds, it is difficult to directly apply two-dimensional object detection methods based on regular grids to three-dimensional object detection. Based on this, some methods for three-dimensional object detection have been proposed. In conventional three-dimensional object detection methods, a point grouping step is usually required to group specific points in the point cloud into corresponding candidate objects. Then, the features of the corresponding objects can be calculated based on the points belonging to each candidate object, so as to locate and identify three-dimensional objects in a three-dimensional scene. However, point grouping usually requires manually set rules to achieve. Although these manual rules can describe the relationship between points and objects to a certain extent, they are not very accurate. Therefore, the detection effect of three-dimensional object detection methods based on manually set rules needs to be further improved. In addition, the three-dimensional object detection method based on point grouping cannot fully utilize the information contained in the point cloud data.
[0014] Some of the problems existing in conventional three-dimensional object detection schemes are discussed above. According to the implementation of the present disclosure, a scheme for three-dimensional object detection is proposed, aiming to solve one or more of the above problems and other potential problems. In this scheme, a feature representation of multiple points is extracted from the point cloud data of a three-dimensional object, and the feature representation of each point includes the position information and appearance features of the point. Based on the feature representations of multiple points, an initial feature representation of a set of candidate three-dimensional objects is determined. The initial feature representation of each candidate three-dimensional object includes the position feature and appearance feature of the candidate three-dimensional object. Based on the feature representations of multiple points and the initial feature representations of a set of candidate three-dimensional objects, the detection result of the three-dimensional object is generated by determining the self-correlation between a set of candidate three-dimensional objects and the cross-correlation between multiple points and a set of candidate three-dimensional objects. Various example implementations of this scheme are further described in detail below with reference to the accompanying drawings.
[0015] Figure 1 A block diagram of a computing device 100 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 1 The computing device 100 shown is merely exemplary and should not constitute any limitation to the functions and scope of the implementations described in the present disclosure. As Figure 1 shown, the computing device 100 includes a computing device 100 in the form of a general-purpose computing device. The components of the computing device 100 may include, but are not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.
[0016] In some implementations, computing device 100 may be implemented as various user terminals or service terminals with computing capabilities. The service terminal may be a server, a large computing device, etc. provided by various service providers. The user terminal may be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that computing device 100 can support any type of user interface (such as "wearable" circuits, etc.).
[0017] Processing unit 110 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in memory 120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of computing device 100. Processing unit 110 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0018] Computing device 100 generally includes multiple computer storage media. Such media may be any available media accessible to computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 120 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Memory 120 may include a three-dimensional object detection module 122, and these program modules are configured to perform the functions of various implementations described herein. The three-dimensional object detection module 122 may be accessed and run by processing unit 110 to implement corresponding functions.
[0019] Storage device 130 may be removable or non-removable media and may include machine-readable media that can be used to store information and / or data and can be accessed within computing device 100. Computing device 100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 1As shown, a disk drive for reading from or writing to a removable, non-volatile disk and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data medium interfaces.
[0020] The communication unit 140 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 100 can be implemented in a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Thus, the computing device 100 can operate in a networked environment using a logical connection with one or more other servers, personal computers (PCs), or another general network node.
[0021] The input device 150 can be one or more of various input devices such as a mouse, keyboard, trackball, voice input device, etc. The output device 160 can be one or more output devices such as a display, speaker, printer, etc. The computing device 100 can also communicate with one or more external devices (not shown) as needed via the communication unit 140, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 100, or communicate with any device that enables the computing device 100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0022] In some implementations, in addition to being integrated on a single device, some or all of the various components of the computing device 100 can also be arranged in the form of a cloud computing architecture. In a cloud computing architecture, these components can be remotely located and can work together to implement the functions described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services that do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network such as the Internet. For example, a cloud computing provider provides applications over a wide area network, and they can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated at a remote data center location or they can be distributed. The cloud computing infrastructure can provide services through a shared data center, even though they appear as a single access point for the user. Thus, the components and functions described herein can be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they can also be provided from a conventional server, or they can be directly or otherwise installed on a client device.
[0023] The computing device 100 can perform three-dimensional object detection according to various implementations of the present disclosure. As Figure 1 shown, the computing device 100 can receive point cloud data 170 about a three-dimensional object through the input device 150. The input device 150 can transmit the point cloud data 170 to the three-dimensional object detection module 122. The three-dimensional object detection module 122 generates a detection result 190 of the three-dimensional object from the point cloud data 170 about the three-dimensional object. For example, the detection result 190 of the three-dimensional object can indicate the position characteristics (e.g., position coordinates and geometric dimensions) and / or other information (e.g., color, shape, category, etc.) of the three-dimensional object, and thus can be used to locate and identify the three-dimensional object in a three-dimensional scene. In Figure 1 the example shown, the three-dimensional object detection module 122 generates a detection result 190 of the three-dimensional object in the living room scene from the point cloud data 170 describing the living room scene. For example, the detection result 190 can include a detection result 190-1 indicating a coffee table and a detection result 190-2 indicating a sofa (collectively referred to as the detection result 190). It should be noted that only the position characteristics in the detection result 190 are shown here, and the position characteristics are represented by a bounding box surrounding the three-dimensional object.
[0024] Figure 2 The architecture diagram of a system 200 for three-dimensional object detection according to an implementation of the present disclosure is shown. The system 200 can be implemented in Figure 1 the computing device 100. The system 200 can be an end-to-end neural network model. As Figure 2 shown, the system 200 can include a point feature extraction module 210, a candidate three-dimensional object generation module 220, an attention module 230, and a selection module 240.
[0025] The point feature extraction module 210 extracts the feature representations 201 of multiple points from the received point cloud data 170. The point cloud data 170 can be a set of vectors used to represent N points in a three-dimensional coordinate system. For example, the vector representing each point can include the three-dimensional coordinates of the point (e.g., x, y, z), color information (e.g., values in the RGB color space), or reflection intensity information. Multiple points can be sampled from the point cloud data 170 and the corresponding feature representations 201 of these points can be extracted through various methods. For example, the PointNet++ architecture can be adopted to sample M points from the point cloud data 170 regarding N points and extract the feature representations 201 of the M points (N≥M). The feature representations 201 of the M points are a set of vectors used to represent the M points, including the position information and appearance features of each point. The position information of each point can be the three-dimensional coordinates of the point in the three-dimensional coordinate system. The appearance feature of each point can be generated based on the feature transformation of the three-dimensional coordinates, color information, and reflection intensity information of the point as described above. The appearance feature of each point can also include information about multiple adjacent points thereto. The appearance feature of the point can indicate the color, shape of the region including the point, and / or the type of the object described by at least one point within the region. Therefore, by sampling multiple points from the point cloud data 170 and extracting the feature representations 201 of the multiple points, the amount of data input to the subsequent modules can be reduced with minimal information loss, thereby saving computing resources. The scope of the present disclosure is not limited in terms of the method for point feature extraction.
[0026] The candidate three-dimensional object generation module 220 determines an initial feature representation 202 of a set of candidate three-dimensional objects based on the feature representation 201 of multiple points. The candidate three-dimensional object generation module 220 may include, for example, a generation unit 221 and a sampling unit 222. The generation unit 221 may generate an initial detection result of multiple candidate three-dimensional objects corresponding to the multiple points based on the extracted feature representation 201 of the multiple points. In some implementations, an initial detection result of M candidate three-dimensional objects corresponding to M points may be generated by using a fully connected layer based on the feature representation 201 of the M points. The initial detection result of each candidate three-dimensional object may indicate the position feature (e.g., position coordinates and geometric dimensions) and / or other information (e.g., color, shape, category, etc.) of the candidate three-dimensional object. The position feature of the candidate three-dimensional object may be represented by a bounding box surrounding the candidate three-dimensional object. The bounding box may be a cuboid, a cube, an ellipsoid, etc. When using a cuboid-type bounding box to represent the position feature of the candidate three-dimensional object, the position feature may be a 6×1-dimensional vector including the length (l), width (w), height (h), and center point coordinates (x, y, z) of the cuboid. The position feature may indicate the positioning of the candidate three-dimensional object in the three-dimensional scene and the geometric dimensions of the candidate three-dimensional object. For each candidate three-dimensional object, an initial feature representation of the candidate three-dimensional object may be generated based on the appearance feature of the corresponding point and the position feature (e.g., position coordinates and geometric dimensions) in the initial detection result. That is, the initial feature representation of each candidate three-dimensional object may include the appearance feature and the position feature of the candidate three-dimensional object. As described above, since the appearance feature of the point may indicate the color, shape, and / or the category of the object described by at least one point within the region including the point, the appearance feature of the candidate three-dimensional object may indicate the color, shape, category, etc. of the region including the point (i.e., the candidate three-dimensional object).
[0027] The sampling unit 222 in the candidate 3D object generation module 220 may determine an initial feature representation 202 of a set of candidate 3D objects based on the initial feature representations of multiple candidate 3D objects. The number of the determined initial feature representations 202 of the set of candidate 3D objects is preset and can be adjusted according to the actual application. A set (e.g., K, K≤M) of initial feature representations 202 of candidate 3D objects can be sampled from the initial feature representations of multiple (e.g., M) candidate 3D objects by using a variety of sampling methods. Examples of sampling methods include Furthest Point Sampling (FPS), k-Closest Points Sampling (KPS), and KPS with non-maximum suppression, etc. In some implementations, the non-maximum suppression algorithm can also be directly used to sample a set of initial feature representations 202 of candidate 3D objects from the initial feature representations of multiple candidate 3D objects. In some implementations, other information in the initial detection result may include information indicating the type of the candidate 3D object. For example, the probability or score that the candidate 3D object is classified into a certain specific type. A set of candidate 3D objects can be selected from multiple candidate 3D objects corresponding to multiple points based on the score indicating the type in the initial detection result. For example, candidate 3D objects with higher scores can be included in a set of candidate 3D objects, so as to determine an initial feature representation 202 of the set of candidate 3D objects. In some implementations, through sampling, object detection results with higher confidence are retained, and only one object detection result is retained among multiple very similar object detection results (e.g., bounding boxes that overlap each other). In this way, the amount of data input to the subsequent module can be reduced, thus saving computing resources.
[0028] The attention module 230 may determine the self-correlation between a set of candidate 3D objects and the cross-correlation between multiple points and a set of candidate 3D objects based on the initial feature representation 202 of the set of candidate 3D objects and the feature representation 201 of multiple points, so as to generate a detection result 190 of the 3D object. As described above, the detection result 190 may indicate the position features (e.g., position coordinates and geometric dimensions) of the detected 3D object and / or other information (e.g., color, shape, and type, etc.). The position features of the 3D object can be represented by a bounding box that encloses the 3D object. The bounding box can be a cuboid, a cube, an ellipsoid, etc. When using a cuboid type of bounding box to represent the position features, the position features can be a 6×1-dimensional vector including the length (l), width (w), height (h), and center point coordinates (x, y, z) of the cuboid. In some implementations, the detection result 190 may also include information indicating the type of the 3D object. For example, the probability or score that the 3D object is classified into a certain specific type.
[0029] Since the correlation among a set of candidate three-dimensional objects and the correlation between multiple points and the candidate three-dimensional objects are considered, the three-dimensional object indicated by the detection result 190 can be closer to the actual three-dimensional object. Therefore, the three-dimensional object detection method implemented according to the present disclosure can achieve high detection accuracy by using the attention module 230 without using the point aggregation operation commonly used in conventional object detection methods.
[0030] In some implementations, the attention module 230 may include a plurality of stacked attention modules 230-1, 230-2... 230-N (collectively or individually referred to as the attention module 230, where N≥1). For example, the attention module 230-1 (hereinafter also referred to as the "first attention module") may be used to generate a first set of candidate detection results 203-1, the attention module 230-2 (hereinafter also referred to as the "second attention module") may be used to generate a second set of candidate detection results 203-2... the attention module 230-N may be used to generate an Nth set of candidate detection results 203-N. Details regarding using the attention module 230 to generate at least one set of candidate detection results 203 will be described below with reference to Figure 3 and Figure 4 to describe.
[0031] The selection module 240 may select the detection result 190 from the at least one set of candidate detection results 203 generated. In Figure 1 the example shown, the detection result 190 includes the detection result 190-1 indicating a coffee table and the detection result 190-2 indicating a sofa. Multiple methods may be used to select the detection result 190 from the at least one set of candidate detection results 203. Examples of multiple methods include the non-maximum suppression algorithm described above. Alternatively or additionally, the detection result 190 may be determined according to the scores for indicating the object types in the candidate detection results. The number of the selected detection results 190 may also be adjusted according to the actual application. It should be understood that the structure and function of the system 200 are described only for exemplary purposes and do not imply any limitation on the scope of the subject matter described herein. The subject matter described herein may be embodied in different structures and / or functions.
[0032] Figure 3 FIG. shows a schematic diagram of using the first attention module 230-1 in the attention module 230 to generate a first set of candidate detection results 203-1 according to an implementation of the present disclosure. As shown in the figure, the first attention module 230-1 may include a self-attention module 310 and a cross-attention module 320.
[0033] In some implementations, to generate the first set of candidate detection results 203-1, the first attention module 230-1 may receive an initial feature representation 202 of a set of candidate three-dimensional objects. As described above, the initial feature representation 202 of a set of candidate three-dimensional objects may include the position features and appearance features of the candidate three-dimensional objects. In some implementations, the position features of the candidate three-dimensional objects may be subjected to a feature transformation, and then the transformed position features may be fused with the appearance features of the candidate three-dimensional objects to generate a combined feature representation of the candidate three-dimensional objects. For example, the position features of the candidate three-dimensional objects may be encoded as vectors of the same dimension as the appearance features, and then the position features and the appearance features may be added vectorially to generate a combined feature representation of the candidate three-dimensional objects.
[0034] In some implementations, the self-attention module 310 in the first attention module 230-1 may determine the self-correlation between a set of candidate three-dimensional objects based on the initial feature representation 202 of the set of candidate three-dimensional objects. For example, the self-attention vectors of a set of candidate three-dimensional objects may be calculated based on the self-attention algorithm to represent the self-correlation between the set of candidate three-dimensional objects. In some implementations, the query matrix (Q), the key matrix (K), and the value matrix (V) may be multiplied by the feature representation o of the candidate three-dimensional objects respectively, so as to map the feature representation o of the candidate three-dimensional objects to the query vector, the key vector, and the value vector respectively. Then, the similarity between the query vector of each candidate three-dimensional object in the set of candidate three-dimensional objects and the key vectors of all candidate three-dimensional objects may be determined, and the self-attention weights of all candidate three-dimensional objects for a given candidate three-dimensional object may be determined based on the similarity. Finally, the self-attention weights of all candidate three-dimensional objects may be multiplied by their corresponding value vectors and summed to obtain the self-attention vector of the given candidate three-dimensional object. In this way, the self-attention vectors of each candidate three-dimensional object in the set of candidate three-dimensional objects may be obtained, and the self-attention vector is used to measure the correlation between the candidate three-dimensional object and the set of candidate three-dimensional objects. In addition, it should be understood that the multi-head attention algorithm may be used to calculate the self-attention vector. In the multi-head attention algorithm, each head uses the corresponding query matrix (Q), key matrix (K), and value matrix (V), and generates the corresponding self-attention vector. The weighted sum of the self-attention vectors generated by each head may be used to obtain the final self-attention vector. Specifically, the self-attention vector Self-Att of candidate three-dimensional object j may be calculated with reference to formula (1).
[0035]
[0036] where l represents the index l of the attention module used to calculate the self-attention vector, j represents the index of the candidate three-dimensional object in a set of candidate three-dimensional objects with the number K, and h represents the index of the head in the multi-head (with the number H) attention algorithm. represents the feature representation of candidate 3D object j, and {o (l)} represents the feature representations of a set of candidate 3D objects. represents the self-attention weight between candidate 3D object j and candidate 3D object k in the h-th head self-attention calculation of the l-th attention module. represents the value vector of the k-th candidate 3D object. and are the multi-head weighted matrix, query matrix, key matrix, and value matrix respectively. These matrices can be obtained through the training process of the neural network.
[0037] In some implementations, in the above formula (1), the feature representations {o (l)} of a set of candidate 3D objects can be the appearance features in the initial feature representations 202 of a set of candidate 3D objects. Alternatively, the feature representations {o (l)} of a set of candidate 3D objects can be the combined feature representations of a set of candidate 3D objects. As described above, the combined feature representations can be generated by re-encoding the position features of the candidate 3D objects and adding them to the appearance feature vectors.
[0038] In some implementations, the self-attention module 310 can update the initial feature representations 202 of a set of candidate 3D objects to the first set of intermediate feature representations 350 of a set of candidate 3D objects based on the determined self-correlation. For example, the calculated self-attention vector Self-Att can be added to the appearance features in the initial feature representations 202 of a set of candidate 3D objects, thereby generating the first set of intermediate feature representations 350 of a set of candidate 3D objects. Additionally, the combined feature representations generated based on the initial feature representations 202 can also be added to the first set of intermediate feature representations for updating.
[0039] In some implementations, the cross-attention module 320 in the first attention module 230-1 can receive the first set of intermediate feature representations 350 and the feature representations 201 of multiple points. As described above, the feature representations 201 of multiple points can include the position information of the points and the appearance features of the points. Similarly, the position information of the points can be subjected to feature transformation, and then the transformed position features can be fused with the appearance features of the points. For example, the position information of the points can be encoded into a vector with the same dimension as the appearance features, and then the vectors of the position features and the appearance features can be added to generate the combined feature representations of the points.
[0040] In some implementations, the cross-attention module 320 can determine the cross-correlation between a set of candidate 3D objects and multiple points based on the first set of intermediate feature representations 350 and the feature representations 201 of multiple points. Similarly, the cross-attention vectors of each candidate 3D object in a set of candidate 3D objects with respect to multiple points can be calculated based on the cross-attention algorithm, and the cross-attention vectors are used to represent the cross-correlation between the set of candidate 3D objects and multiple points. In some implementations, the query matrix (Q) can be multiplied by the feature representation o of the candidate 3D object to map the feature representation o of the candidate 3D object to a query vector. By multiplying the key matrix (K) and the value matrix (V) with the feature representation z of the points respectively, the feature representation z of the points can be mapped to a key vector and a value vector respectively. Then, the similarity between the query vector of a given candidate 3D object in the set of candidate 3D objects and the key vector of each point can be determined, and the cross-attention weight of each point with respect to the given candidate 3D object can be determined based on the similarity. Finally, the cross-attention weight of each point is multiplied by its corresponding value vector and summed to obtain the cross-attention vector of the given candidate 3D object with respect to multiple points. In this way, the cross-attention vectors of each candidate 3D object in the set of candidate 3D objects can be obtained, and the cross-attention vector is used to measure the correlation between the candidate 3D object and multiple points. Similarly, the multi-head attention algorithm can be used to calculate the cross-attention vector. Specifically, the cross-attention vector Cross-Att of candidate 3D object j can be calculated with reference to formula (2).
[0041]
[0042] where m represents the index of the point among the multiple points with the number M, {z (l)} represents the feature representations of M points, represents the cross-attention weight between candidate 3D object j and point i in the h-th head self-attention calculation of the l-th attention module, represents the value vector of point i.
[0043] In some implementations, in the above formula (2), the feature representations {z (l)} of the points can be the appearance features in the feature representations 201 of the points. Alternatively, the feature representations {z (l)} of the points can be the combined feature representations of the points as described above. As described above, the combined feature representation can be generated by re-encoding the position information of the points and adding it to the appearance feature vector.
[0044] In some implementations, the mutual attention module 320 may update the first set of intermediate feature representations 350 to a first set of candidate feature representations 360-1 of a set of candidate 3D objects based on the determined mutual correlation. For example, the calculated mutual attention vector may be added to the first set of intermediate feature representations 350 of a set of candidate 3D objects to generate a first set of candidate feature representations 360-1 of a set of candidate 3D objects. Additionally, the combined feature representation generated based on the initial feature representation 202 of a set of candidate 3D objects may be added to the first set of candidate feature representations 360-1 for updating. Additionally, the first set of candidate feature representations 360-1 may be normalized. Similar to the initial feature representation 202, the first set of candidate feature representations 360-1 includes the position features and appearance features of a set of candidate 3D objects. The position features may be represented using bounding boxes enclosing the 3D objects, and the appearance features may indicate the color, shape, category, etc. of the candidate 3D objects.
[0045] In some implementations, the first attention module 230-1 may generate a first set of candidate detection results 203-1 based on the first set of candidate feature representations 360-1. For example, the first set of candidate feature representations 360-1 may be input into multiple fully connected layers 330 for feature transformation to obtain a first set of candidate detection results 203-1. Additionally, before inputting the first set of candidate feature representations 360-1 into multiple fully connected layers 330, the first set of candidate feature representations 360-1 may be input into another multiple fully connected layers for feature transformation to generate a transformed first set of candidate feature representations, and the combined feature representation generated based on the initial feature representation 202 of a set of candidate 3D objects may be added to it for updating (not shown in Figure 3 ). Since the self-correlation between a set of candidate 3D objects and the mutual correlation between a set of candidate 3D objects and multiple points are considered, compared with the initial detection results of a set of candidate 3D objects, the set of candidate 3D objects indicated by the first set of candidate detection results 203-1 may be closer to the actual 3D objects. Using the attention algorithm, points with strong correlations can be identified as belonging to the same object without manual rules. Therefore, the 3D object detection method according to the implementation of the present disclosure can achieve 3D object detection without a point aggregation operation.
[0046] The above references Figure 3 describe the process of generating the first set of candidate detection results 203-1 using the first attention module 230-1 in the attention module 230. Based on the above discussion, the process of using multiple stacked attention modules 230 to generate at least one set of candidate detection results 203 and determining the detection results 190 therefrom will be described in detail below with reference to Figure 4 As Figure 4As shown, the first attention module 230-1 can receive the feature representations 201 of multiple points and the initial feature representations 202 of a set of candidate 3D objects. As described above with reference to Figure 3 As described, the first attention module 230-1 can generate a first set of candidate feature representations 360-1 based on the feature representations 201 of multiple points and the initial feature representations 202 of a set of candidate 3D objects. The first set of candidate feature representations 360-1 can be provided to the second attention module 230-2.
[0047] In some implementations, the self-attention module (not shown here) in the second attention module 230-2 can use the method described with reference to Figure 3 As described, based on the first set of candidate feature representations 360-1, determine the self-correlation between a set of candidate 3D objects. Similarly, based on the determined self-correlation, the self-attention module in the second attention module 230-2 can update the first set of candidate feature representations 360-1 to a second set of intermediate feature representations of a set of candidate 3D objects. Then, the cross-attention module in the second attention module 230-2 can determine the cross-correlation between a set of candidate 3D objects and multiple points based on the second set of intermediate feature representations and the feature representations 201 of multiple points. Based on the determined cross-correlation, the cross-attention module in the second attention module 230-2 can update the second set of intermediate feature representations to a second set of candidate feature representations 360-2 of a set of candidate 3D objects. The second attention module 230-2 can generate a second set of candidate detection results 203-2 based on the second set of candidate feature representations 360-2.
[0048] Since the process of generating the second set of candidate detection results 203-2 is similar to the process of generating the first set of candidate detection results 203-1, the details of the process will not be elaborated here. Compared with the process of generating the first set of candidate detection results 203-1, when generating the second set of candidate detection results 203-2, the second attention module receives the first set of candidate feature representations 360-1 as input, rather than the initial feature representations 202 of a set of candidate 3D objects. Since the first set of candidate feature representations 360-1 is generated based on the initial feature representations 202 and the correlation between a set of candidate 3D objects and multiple points, inputting the first set of candidate feature representations 360-1 into the second attention module 230-2 can generate a second set of candidate feature representations 360-2 with more correlation information.
[0049] Therefore, when there are N attention modules in the attention module 230, each attention module can generate a corresponding set of candidate detection results. Different from the first attention module 230-1 receiving the initial feature representation 202 of a set of candidate three-dimensional objects as input, each subsequent attention module receives a set of candidate feature representations generated by the previous attention module as input. In this way, the correlation (including self-correlation and cross-correlation) between a set of candidate three-dimensional objects and multiple points can be accumulated to improve the accuracy of the three-dimensional objects indicated by the generated set of candidate feature representations.
[0050] It should be noted that although it is generally considered that subsequent attention modules can generate a more accurate set of candidate detection results, in the implementation of the present disclosure, a set of candidate detection results generated by each attention module can be input into the selection module 240 for determining the detection result 190. In some implementations, a union operation can be performed on at least one set of candidate detection results 203 generated by at least one attention module 230, and the detection result 190 can be selected from the union. In some implementations, the detection result 190 can also be selected from the candidate detection results generated by some attention modules. In some implementations, multiple methods can be used to select the detection result 190. For example, the non-maximum suppression algorithm can be used to select the detection result 190 from at least one set of candidate detection results 203. In other implementations, the detection result 190 can be determined according to the scores indicating the types of three-dimensional objects in the candidate detection results.
[0051] The above references Figures 1-4 The working principle of the method for three-dimensional object detection according to the implementation of the present disclosure has been described in detail above. The training process of the end-to-end neural network model used in this method will be described below.
[0052] In some implementations, the neural network is trained in a supervised learning manner using a training data set of manually annotated three-dimensional scenes. The training data set can include the manually annotated object types, bounding box center positions, and bounding box geometric dimensions. The core idea of training the neural network is to maximize the similarity between the three-dimensional objects detected by the neural network and the actual three-dimensional objects, that is, to minimize the differences between the object types, bounding box center positions, and bounding box geometric dimensions detected by the neural network and the corresponding manually annotated items. The difference between the detected three-dimensional objects and the actual three-dimensional objects can be described based on a loss function. By adjusting the architecture and parameters of the neural network to minimize the loss function, an optimized network architecture and parameters can be obtained. By applying the neural network for three-dimensional object detection using the optimized network architecture and parameters, three-dimensional object detection results with high accuracy can be obtained. Specifically, for the end-to-end neural network model for three-dimensional object detection according to the implementation of the present disclosure, the loss function can be expressed as:
[0053] wherein represents the overall loss function, represents the loss function for the attention module 230, and represents the loss function for the candidate 3D object generation module 220.
[0054] As described above, the attention module 230 can be a stack of multiple attention modules. Thus, as shown in Equation (4), can be the average of the loss functions for each attention module.
[0055]
[0056] The loss function for each attention module can be calculated using Equation (5).
[0057]
[0058] wherein, represents the loss function for the object category, represents the loss function for the bounding box classification, represents the loss function for the bounding box center offset, the loss function for the coarse-grained bounding box size, represents the loss function for the fine-grained bounding box size, and β1 to β5 represent the corresponding coefficients. In an exemplary implementation, the values of β can default to 0.5, 0.1, 1.0, 0.1, and 0.3, respectively.
[0059] Similarly, the loss function for the candidate 3D object generation module 220 can also be calculated using Equation (5). As described above, the initial detection results of a set of candidate 3D objects determined in the candidate 3D object generation module 220 can include position features in the form of a bounding box and information indicating the category of the candidate 3D object. Thus, the loss function for the candidate 3D object generation module 220 can also be and a weighted sum of. Alternatively, in the candidate 3D object generation module 220, only the position features of the object can be determined. In this implementation, the loss function for the candidate 3D object generation module 220 can be only and a weighted sum of.
[0060] Figure 5FIG. 0 shows a flowchart of a method 500 for three-dimensional object detection according to some implementations of the present disclosure. The method 500 may be implemented by a computing device 100, for example, at a three-dimensional object detection module 122 that may be implemented in a memory 120 of the computing device 100.
[0061] As Figure 5 shown, at block 510, the computing device 100 extracts a feature representation of a plurality of points from point cloud data regarding a three-dimensional object, where the feature representation of each point includes the position information and appearance features of the point. At block 520, the computing device 100 determines an initial feature representation of a set of candidate three-dimensional objects based on the feature representations of the plurality of points, where the initial feature representation of each candidate three-dimensional object includes the position features and appearance features of the candidate three-dimensional object. At block 530, the computing device 100 generates a detection result of the three-dimensional object based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects by determining the self-correlation between the set of candidate three-dimensional objects and the cross-correlation between the plurality of points and the set of candidate three-dimensional objects.
[0062] In some implementations, generating the detection result of the three-dimensional object includes: generating at least one set of candidate detection results of the set of candidate three-dimensional objects by using at least one attention module; and determining the detection result of the three-dimensional object from the at least one set of candidate detection results.
[0063] In some implementations, the at least one attention module includes a first attention module, and generating at least one set of candidate detection results by using the at least one attention module includes: generating a first set of candidate detection results in the at least one set of candidate detection results by using the first attention module based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects.
[0064] In some implementations, generating the first set of candidate detection results includes: determining the self-correlation between the set of candidate three-dimensional objects by using a self-attention module in the first attention module based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects; updating the initial feature representations of the set of candidate three-dimensional objects to a first set of intermediate feature representations of the set of candidate three-dimensional objects based on the determined self-correlation; determining the cross-correlation between the set of candidate three-dimensional objects and the plurality of points by using a cross-attention module in the first attention module based on the first set of intermediate feature representations and the feature representations of the plurality of points; updating the first set of intermediate feature representations to a first set of candidate feature representations of the set of candidate three-dimensional objects based on the determined cross-correlation; and generating the first set of candidate detection results based on the first set of candidate feature representations.
[0065] In some implementations, the at least one attention module further includes a second attention module, and generating at least one set of candidate detection results by using the at least one attention module further includes: generating a second set of candidate detection results in the at least one set of candidate detection results by using the second attention module based on the feature representations of multiple points and a first set of candidate feature representations.
[0066] In some implementations, generating the second set of candidate detection results includes: determining the self-correlation among a set of candidate three-dimensional objects by using the self-attention module in the second attention module based on the feature representations of multiple points and the first set of candidate feature representations; updating the first set of candidate feature representations to a second set of intermediate feature representations of a set of candidate three-dimensional objects based on the determined self-correlation; determining the cross-correlation between a set of candidate three-dimensional objects and multiple points by using the cross-attention module in the second attention module based on the second set of intermediate feature representations and the feature representations of multiple points; updating the second set of intermediate feature representations to a second set of candidate feature representations of a set of candidate three-dimensional objects based on the determined cross-correlation; and generating the second set of candidate detection results based on the second set of candidate feature representations.
[0067] In some implementations, determining the initial feature representations of a set of candidate three-dimensional objects includes: generating the initial feature representations of multiple candidate three-dimensional objects corresponding to multiple points based on the feature representations of multiple points; and selecting the initial feature representations of a set of candidate three-dimensional objects from the initial feature representations of multiple candidate three-dimensional objects.
[0068] In some implementations, the detection result indicates at least one of the position coordinates, geometric dimensions, color, shape, and type of the three-dimensional object.
[0069] It can be seen from the above description that the three-dimensional object detection solution according to the implementation of the present disclosure can determine the final feature representations of three-dimensional objects only based on the correlations between each point in the point cloud and candidate three-dimensional objects and the correlations between candidate three-dimensional objects without aggregating points into candidate objects, so as to locate and identify three-dimensional objects in a three-dimensional scene.
[0070] Some example implementation manners of the present disclosure are listed below.
[0071] In one aspect, the present disclosure provides a computer-implemented method. The method includes: extracting a feature representation of a plurality of points from point cloud data of a three-dimensional object, where the feature representation of each point includes the position information and appearance features of the point; determining an initial feature representation of a set of candidate three-dimensional objects based on the feature representations of the plurality of points, where the initial feature representation of each candidate three-dimensional object includes the position features and appearance features of the candidate three-dimensional object; and generating a detection result of the three-dimensional object based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects by determining the self-correlation between the set of candidate three-dimensional objects and the cross-correlation between the plurality of points and the set of candidate three-dimensional objects.
[0072] In some implementations, generating a detection result of a three-dimensional object includes: using at least one attention module to generate at least one set of candidate detection results of a set of candidate three-dimensional objects; and determining a detection result of the three-dimensional object from the at least one set of candidate detection results.
[0073] In some implementations, the at least one attention module includes a first attention module, and using the at least one attention module to generate at least one set of candidate detection results includes: generating a first set of candidate detection results in the at least one set of candidate detection results using the first attention module based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects.
[0074] In some implementations, generating the first set of candidate detection results includes: determining the self-correlation between a set of candidate three-dimensional objects using a self-attention module in the first attention module based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects; updating the initial feature representations of the set of candidate three-dimensional objects to a first set of intermediate feature representations of the set of candidate three-dimensional objects based on the determined self-correlation; determining the cross-correlation between the set of candidate three-dimensional objects and the plurality of points using a cross-attention module in the first attention module based on the first set of intermediate feature representations and the feature representations of the plurality of points; updating the first set of intermediate feature representations to a first set of candidate feature representations of the set of candidate three-dimensional objects based on the determined cross-correlation; and generating the first set of candidate detection results based on the first set of candidate feature representations.
[0075] In some implementations, the at least one attention module further includes a second attention module, and using the at least one attention module to generate at least one set of candidate detection results further includes: generating a second set of candidate detection results in the at least one set of candidate detection results using the second attention module based on the feature representations of the plurality of points and the first set of candidate feature representations.
[0076] In some implementations, generating a second set of candidate detection results includes: determining the self-correlation between a set of candidate three-dimensional objects by using the self-attention module in the second attention module based on the feature representations of multiple points and the first set of candidate feature representations; updating the first set of candidate feature representations to a second set of intermediate feature representations of a set of candidate three-dimensional objects based on the determined self-correlation; determining the cross-correlation between a set of candidate three-dimensional objects and multiple points by using the cross-attention module in the second attention module based on the second set of intermediate feature representations and the feature representations of multiple points; updating the second set of intermediate feature representations to a second set of candidate feature representations of a set of candidate three-dimensional objects based on the determined cross-correlation; and generating a second set of candidate detection results based on the second set of candidate feature representations.
[0077] In some implementations, determining an initial feature representation of a set of candidate three-dimensional objects includes: generating initial feature representations of multiple candidate three-dimensional objects corresponding to multiple points based on the feature representations of multiple points; and selecting an initial feature representation of a set of candidate three-dimensional objects from the initial feature representations of multiple candidate three-dimensional objects.
[0078] In some implementations, the detection result indicates at least one of the position coordinates, geometric dimensions, color, shape, and type of the three-dimensional object.
[0079] In another aspect, the present disclosure provides an electronic device. The electronic device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, which when executed by the processing unit cause the device to perform actions including: extracting feature representations of multiple points from point cloud data regarding a three-dimensional object, where the feature representation of each point includes the position information and appearance features of the point; determining an initial feature representation of a set of candidate three-dimensional objects based on the feature representations of multiple points, where the initial feature representation of each candidate three-dimensional object includes the position feature and appearance feature of the candidate three-dimensional object; and generating a detection result of the three-dimensional object by determining the self-correlation between the set of candidate three-dimensional objects and the cross-correlation between the multiple points and the set of candidate three-dimensional objects based on the feature representations of multiple points and the initial feature representation of the set of candidate three-dimensional objects.
[0080] In some implementations, generating a detection result of a three-dimensional object includes: generating at least one set of candidate detection results of a set of candidate three-dimensional objects by using at least one attention module; and determining a detection result of the three-dimensional object from the at least one set of candidate detection results.
[0081] In some implementations, at least one attention module includes a first attention module, and generating at least one set of candidate detection results by using the at least one attention module includes: generating a first set of candidate detection results in the at least one set of candidate detection results by using the first attention module based on the feature representations of multiple points and the initial feature representations of a set of candidate three-dimensional objects.
[0082] In some implementations, generating the first set of candidate detection results includes: determining the self-correlation between a set of candidate three-dimensional objects by using the self-attention module in the first attention module based on the feature representations of multiple points and the initial feature representations of a set of candidate three-dimensional objects; updating the initial feature representations of a set of candidate three-dimensional objects to a first set of intermediate feature representations of a set of candidate three-dimensional objects based on the determined self-correlation; determining the cross-correlation between a set of candidate three-dimensional objects and multiple points by using the cross-attention module in the first attention module based on the first set of intermediate feature representations and the feature representations of multiple points; updating the first set of intermediate feature representations to a first set of candidate feature representations of a set of candidate three-dimensional objects based on the determined cross-correlation; and generating the first set of candidate detection results based on the first set of candidate feature representations.
[0083] In some implementations, the at least one attention module further includes a second attention module, and generating at least one set of candidate detection results by using the at least one attention module further includes: generating a second set of candidate detection results in the at least one set of candidate detection results by using the second attention module based on the feature representations of multiple points and the first set of candidate feature representations.
[0084] In some implementations, generating the second set of candidate detection results includes: determining the self-correlation between a set of candidate three-dimensional objects by using the self-attention module in the second attention module based on the feature representations of multiple points and the first set of candidate feature representations; updating the first set of candidate feature representations to a second set of intermediate feature representations of a set of candidate three-dimensional objects based on the determined self-correlation; determining the cross-correlation between a set of candidate three-dimensional objects and multiple points by using the cross-attention module in the second attention module based on the second set of intermediate feature representations and the feature representations of multiple points; updating the second set of intermediate feature representations to a second set of candidate feature representations of a set of candidate three-dimensional objects based on the determined cross-correlation; and generating the second set of candidate detection results based on the second set of candidate feature representations.
[0085] In some implementations, determining the initial feature representations of a set of candidate three-dimensional objects includes: generating the initial feature representations of multiple candidate three-dimensional objects corresponding to multiple points based on the feature representations of multiple points; and selecting the initial feature representations of a set of candidate three-dimensional objects from the initial feature representations of multiple candidate three-dimensional objects.
[0086] In some implementations, the detection result indicates at least one of the position coordinates, geometric dimensions, color, shape, and type of the three-dimensional object.
[0087] In another aspect, the present disclosure provides a computer program product tangibly stored in a non-transitory computer storage medium and including machine-executable instructions that, when executed by a device, cause the device to perform the methods of the above aspects.
[0088] In another aspect, the present disclosure provides a computer program product including machine-executable instructions that, when executed by a device, cause the device to perform the methods of the above aspects.
[0089] In another aspect, the present disclosure provides a computer-readable medium having machine-executable instructions stored thereon that, when executed by a device, cause the device to perform the methods of the above aspects.
[0090] The functions described above herein can be performed at least in part by one or more hardware logic components. By way of example and not limitation, the types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0091] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0092] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0093] Moreover, although the operations are depicted in a particular order, this should be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, the various features that are described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.
[0094] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A computer-implemented method, comprising: extracting a feature representation of a plurality of points from point cloud data of a three-dimensional object, wherein the feature representation of each point includes the position information and appearance features of the point; determining an initial feature representation of a set of candidate three-dimensional objects based on the feature representations of the plurality of points, wherein the initial feature representation of each candidate three-dimensional object includes the position features and appearance features of the candidate three-dimensional object; determining the self-correlation between the set of candidate three-dimensional objects by calculating the self-attention vectors of the set of candidate three-dimensional objects based on a self-attention algorithm based on the initial feature representations of the set of candidate three-dimensional objects, and updating the initial feature representations of the set of candidate three-dimensional objects to a set of intermediate feature representations of the set of candidate three-dimensional objects based on the self-correlation; determining the cross-correlation between the set of candidate three-dimensional objects and the plurality of points by calculating the cross-attention vectors of each candidate three-dimensional object in the set of candidate three-dimensional objects with respect to the plurality of points based on a cross-attention algorithm based on the set of intermediate feature representations and the feature representations of the plurality of points, and updating the set of intermediate feature representations to a set of candidate feature representations of the set of candidate three-dimensional objects based on the cross-correlation; and generating a detection result of the three-dimensional object based on the set of candidate feature representations.
2. The method according to claim 1, wherein generating the detection result of the three-dimensional object includes: generating at least one set of candidate detection results of the set of candidate three-dimensional objects by using at least one attention module; and determining the detection result of the three-dimensional object from the at least one set of candidate detection results.
3. The method according to claim 2, wherein the at least one attention module includes a first attention module, and generating the at least one set of candidate detection results by using the at least one attention module includes: generating a first set of candidate detection results in the at least one set of candidate detection results by using the first attention module based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects.
4. The method according to claim 3, wherein the at least one attention module further includes a second attention module, and generating the at least one set of candidate detection results by using the at least one attention module further includes: generating a second set of candidate detection results in the at least one set of candidate detection results by using the second attention module based on the feature representations of the plurality of points and the first set of candidate feature representations, wherein the first set of candidate feature representations replaces the initial feature representations of the set of candidate three-dimensional objects.
5. The method according to claim 1, wherein determining the initial feature representation of a set of candidate three-dimensional objects includes: generating the initial feature representations of a plurality of candidate three-dimensional objects corresponding to the plurality of points based on the feature representations of the plurality of points; and selecting the initial feature representation of the set of candidate three-dimensional objects from the initial feature representations of the plurality of candidate three-dimensional objects.
6. The method according to claim 1, wherein the detection result indicates at least one of position coordinates, geometric dimensions, color, shape, and type of the three-dimensional object.
7. An electronic device, comprising: a processing unit; and a memory, coupled to the processing unit and containing instructions stored thereon, which when executed by the processing unit cause the electronic device to perform actions, the actions including: extracting a feature representation of a plurality of points from point cloud data of a three-dimensional object, the feature representation of each point including position information and appearance features of the point; determining an initial feature representation of a set of candidate three-dimensional objects based on the feature representations of the plurality of points, the initial feature representation of each candidate three-dimensional object including position features and appearance features of the candidate three-dimensional object; determining the self-correlation between the set of candidate three-dimensional objects by calculating a self-attention vector of the set of candidate three-dimensional objects based on a self-attention algorithm based on the initial feature representation of the set of candidate three-dimensional objects, and updating the initial feature representation of the set of candidate three-dimensional objects to a set of intermediate feature representations of the set of candidate three-dimensional objects based on the self-correlation; determining the cross-correlation between the set of candidate three-dimensional objects and the plurality of points by calculating a cross-attention vector of each candidate three-dimensional object in the set of candidate three-dimensional objects with respect to the plurality of points based on a cross-attention algorithm based on the set of intermediate feature representations and the feature representations of the plurality of points, and updating the set of intermediate feature representations to a set of candidate feature representations of the set of candidate three-dimensional objects based on the cross-correlation; and generating a detection result of the three-dimensional object based on the set of candidate feature representations.
8. The electronic device according to claim 7, wherein generating the detection result of the three-dimensional object includes: generating at least one set of candidate detection results of the set of candidate three-dimensional objects by using at least one attention module; and determining the detection result of the three-dimensional object from the at least one set of candidate detection results.
9. The electronic device according to claim 8, wherein the at least one attention module includes a first attention module, and generating the at least one set of candidate detection results by using the at least one attention module includes: generating a first set of candidate detection results in the at least one set of candidate detection results by using the first attention module based on the feature representations of the plurality of points and the initial feature representation of the set of candidate three-dimensional objects.
10. The electronic device according to claim 9, wherein the at least one attention module further includes a second attention module, and generating the at least one set of candidate detection results by using the at least one attention module further includes: generating a second set of candidate detection results in the at least one set of candidate detection results by using the second attention module based on the feature representations of the plurality of points and the first set of candidate feature representations, wherein the first set of candidate feature representations replaces the initial feature representation of the set of candidate three-dimensional objects.
11. The electronic device according to claim 7, wherein determining the initial feature representation of a set of candidate three-dimensional objects includes: Generate an initial feature representation of a plurality of candidate three-dimensional objects corresponding to the plurality of points based on the feature representations of the plurality of points; and Select the initial feature representations of the set of candidate three-dimensional objects from the initial feature representations of the plurality of candidate three-dimensional objects.
12. The electronic device according to claim 7, wherein the detection result indicates at least one of the position coordinates, geometric dimensions, color, shape, and type of the three-dimensional object.
13. A computer program product, the computer program product comprising machine-executable instructions that, when executed by a device, cause the device to perform operations, the operations including: Extract feature representations of a plurality of points from point cloud data of a three-dimensional object, where the feature representation of each point includes the position information and appearance features of the point; Based on the feature representations of the plurality of points, determine an initial feature representation of a set of candidate three-dimensional objects, where the initial feature representation of each candidate three-dimensional object includes the position features and appearance features of the candidate three-dimensional object; Based on the initial feature representations of the set of candidate three-dimensional objects, determine the self-correlation between the set of candidate three-dimensional objects by calculating the self-attention vectors of the set of candidate three-dimensional objects based on the self-attention algorithm, and update the initial feature representations of the set of candidate three-dimensional objects to a set of intermediate feature representations of the set of candidate three-dimensional objects based on the self-correlation; Based on the set of intermediate feature representations and the feature representations of the plurality of points, determine the cross-correlation between the set of candidate three-dimensional objects and the plurality of points by calculating the cross-attention vectors of each candidate three-dimensional object in the set of candidate three-dimensional objects with respect to the plurality of points based on the cross-attention algorithm, and update the set of intermediate feature representations to a set of candidate feature representations of the set of candidate three-dimensional objects based on the cross-correlation; and Generate a detection result of the three-dimensional object based on the set of candidate feature representations.
14. The computer program product according to claim 13, wherein generating the detection result of the three-dimensional object includes: Using at least one attention module to generate at least one set of candidate detection results of the set of candidate three-dimensional objects; and Determine the detection result of the three-dimensional object from the at least one set of candidate detection results.
15. The computer program product according to claim 14, wherein the at least one attention module includes a first attention module, and using the at least one attention module to generate the at least one set of candidate detection results includes: Based on the feature representations of the plurality of points and the initial feature representations of the set of candidate three-dimensional objects, use the first attention module to generate a first set of candidate detection results in the at least one set of candidate detection results.
16. The computer program product according to claim 15, wherein the at least one attention module further includes a second attention module, and using the at least one attention module to generate the at least one set of candidate detection results further includes: Based on the feature representation of the plurality of points and the first set of candidate feature representations, wherein the first set of candidate feature representations replaces the initial feature representation of the set of candidate three-dimensional objects, the second attention module is used to generate a second set of candidate detection results in the at least one set of candidate detection results.
17. The computer program product according to claim 13, wherein determining an initial feature representation of a set of candidate three-dimensional objects comprises: generating, based on the feature representation of the plurality of points, initial feature representations of a plurality of candidate three-dimensional objects corresponding to the plurality of points; and selecting the initial feature representation of the set of candidate three-dimensional objects from the initial feature representations of the plurality of candidate three-dimensional objects.
18. The computer program product according to claim 13, wherein the detection result indicates at least one of the position coordinates, geometric dimensions, color, shape, and type of the three-dimensional object.
Citation Information
Patent Citations
Brain disease classification system based on self-attention mechanism
CN109165667A
3D point cloud data semantic segmentation method based on deep learning and self-attention
CN110245709A