Task Recognition Point Cloud Downsampling
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2026-08-14
Smart Images

Figure 0007905332000006 
Figure 0007905332000007 
Figure 0007905332000008
Abstract
Description
Technical Field
[0001] This principle generally relates to the field of point cloud processing. This document is also understood in the context of the analysis, interpolation, representation, and understanding of point cloud signals.
Background Art
[0002] This section intends to introduce readers to various technical aspects that may be related to the various aspects of the present principle described and / or claimed below. This discussion is thought to be helpful in providing readers with background information to facilitate a better understanding of the various aspects of the present principle. Therefore, it should be understood that these descriptions should be read from this perspective and should not be read as an admission of prior art.
[0003] Point clouds are a data format used across several business fields, including autonomous driving, robotics, AR / VR, civil engineering, computer graphics, and the animation / movie industry. 3D LIDAR sensors are deployed in autonomous vehicles, and affordable LIDAR sensors are included, for example, in the Apple iPad Pro 2020 and the Intel RealSense LIDAR Camera L515. With the advancement of sensing technology, three-dimensional (3D) point cloud data has become more practical and is expected to be a valuable enabling factor in the applications mentioned.
[0004] At the same time, point cloud data may consume, for example, network traffic between connected cars over a 5G network and most of the traffic in immersive communications (virtual reality or augmented reality (VR / AR)). The understanding and communication of point clouds essentially lead to an efficient representation format. In particular, raw point cloud data needs to be properly organized and processed for world modeling and sensing purposes.
[0005] Furthermore, point clouds can represent a continuous representation of the same scene containing multiple moving objects. These are called dynamic point clouds, in contrast to static point clouds, which are captured from static scenes or static objects. Dynamic point clouds may be composed of different frames, each captured at a different time.
[0006] 3D point cloud data is essentially a collection of distinct samples on the surface of an object or scene. In reality, a vast number of points are needed to fully represent the real world with point samples. For example, a typical VR immersive scene contains millions of points, while a point cloud map typically contains hundreds of millions. Therefore, processing such large point clouds is computationally expensive, especially for consumer devices with limited computing power, such as smartphones, tablets, and car navigation systems.
[0007] Point cloud data is crucial for various applications such as autonomous driving, VR / AR, geomorphology, and cartography, but consuming large point clouds directly leads to significant computational costs. Therefore, adaptively downsampling input point clouds is important to facilitate subsequent tasks. Such downsampling processes are useful for scene flow estimation, point cloud compression, and other common computer vision tasks. [Overview of the project]
[0008] The following is a simplified overview of the Principle to provide a basic understanding of some aspects of it. This overview is not a comprehensive overview of the Principle. It is not intended to identify any important or significant elements of the Principle. The following overview merely presents some aspects of the Principle in a simplified form as a prelude to the more detailed explanation provided below.
[0009] This principle describes a method for generating point-level feature vectors for each point in a point cloud, and set-level feature vectors for the point cloud, using a neural network. Based on the point-level and set-level feature vectors, representative locations are generated. The representative locations and set-level feature vectors are output as set descriptors.
[0010] In another embodiment, a method for extracting a point cloud from a data stream involves obtaining a downsampled point cloud and a residual point cloud from the data stream. The downsampled point cloud is fed to a predictor builder module to obtain a predicted point cloud. The point cloud is extracted by adding the predicted point cloud to the residual point cloud.
[0011] The principle also relates to a device comprising at least one processor associated with at least one memory configured to implement embodiments corresponding to the method described above. [Brief explanation of the drawing]
[0012] This disclosure will be better understood by reading the following description, which will reveal other specific features and advantages, and this specification will refer to the attached drawings. [Figure 1] A non-limiting embodiment of the present principle is shown, and a method 10 for downsampling an input point cloud X at n points for a subsequent mechanical task is presented. [Figure 2] A schematic representation of the SD function based on a non-limiting embodiment of this principle is shown below. [Figure 3] Here is an example of how point A is selected as the representative point because it has the greatest weight. [Figure 4] A fifth embodiment of downsampling the input point cloud using this principle is shown. [Figure 5] This paper outlines how the task recognition point cloud downsampling method of this principle can be integrated with subsequent machine tasks. [Figure 6] A seventh embodiment of an integrated task-recognition point cloud downsampling method is shown. [Figure 7] This document demonstrates a point cloud compression method using an embodiment of the task-recognition point cloud downsampling method based on this principle. [Figure 8] An embodiment of the decoder based on this principle is shown. [Figure 9] An exemplary architecture of device 30, which may be configured to implement the method described in relation to Figure 1, is shown. [Modes for carrying out the invention]
[0013] The principle is fully described below with reference to the accompanying drawings, which illustrate examples of the principle. However, the principle can be embodied in many alternative forms and should not be construed as being limited to the embodiments described herein. Thus, the principle is open to various modifications and alternative forms, specific examples of which are shown as examples in the drawings and described in detail herein. However, it should be understood that there is no intention to limit the principle to any particular form disclosed, on the contrary, this disclosure covers all modifications, equivalents, and alternatives that fall within the spirit and scope of the principle as defined by the claims.
[0014] The terms used herein are for the purpose of illustrating only specific embodiments and are not intended to limit the principles herein. Where used herein, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context otherwise explicitly indicates. Where used herein, the terms “comprises,” “comprising,” “includes,” and / or “including” specify the presence of the described features, integers, steps, actions, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof. Furthermore, where an element is referred to as “responding to” or “connecting to” another element, it may directly respond to or be able to connect to the other element, or an intervening element may exist. In contrast, where an element is referred to as “directly responding to” or “directly connecting to” another element, there is no intervening element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated enumerated items, and may be abbreviated as " / ".
[0015] In this specification, terms such as "first," "second," etc., may be used to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, the first element may be called the second element, and similarly, the second element may be called the first element without deviating from the teachings of this principle.
[0016] Some parts of the diagram include arrows on the communication path to indicate the main direction of communication, but please understand that communication may occur in the opposite direction to the depicted arrows.
[0017] Some examples are described with respect to block diagrams and flowcharts of portions of circuit elements, modules, or code, where each block contains one or more executable instructions for implementing a specified logical function. Note also that in other implementations, the functions described in the blocks may occur in the order described. For example, two blocks shown in sequence may actually be executed substantially simultaneously, or the blocks may be executed in the reverse order, depending upon the functions involved.
[0018] As used herein, "according to an example" or "in an example" means that a particular feature, structure, or characteristic described in connection with the example may be included in at least one implementation of the principle. The appearances of the phrases "according to an example" or "in an example" in various places in this specification are not necessarily all referring to the same example, and in separate or alternative examples, they are not necessarily mutually exclusive of other examples.
[0019] The reference numbers appearing in the claims are for illustrative purposes only and shall not have a limiting effect on the scope of the claims. Although not explicitly stated, the examples and variations may be used in any combination or partial combination.
[0020] The automotive industry and self-driving vehicles are areas where point clouds can be used. Self-driving vehicles should be able to "explore" their environment and make good driving decisions based on the reality immediately around them. Representative sensors such as LIDAR produce (dynamic) point clouds that are used by the decision engine. These point clouds are not intended to be viewed by the human eye; they are typically sparse, not necessarily color-coded, and dynamic at a high capture frequency. They may have other attributes such as the reflectivity provided by LIDAR, as this attribute can indicate the material of the object being sensed and can help in making that determination.
[0021] Virtual reality and immersive worlds have recently been widely considered and predicted by many as the future of 2D flat video. The basic idea is to immerse the viewer in the environment around them, in contrast to a standard TV where the viewer only sees the virtual world in front of them. Depending on the degree of freedom of the viewer within the environment, there are several levels of immersion. Point clouds are good candidate formats for delivering virtual reality worlds. They can be static or dynamic and typically have an average size, for example, less than millions of points per instance.
[0022] Point clouds can also be used for various purposes such as cultural heritage / buildings, where objects like statues or buildings there are scanned in 3D to share the spatial configuration of the objects without sending or physically visiting the object. This also provides a way to preserve information and data about the object, for example, in cases where a temple can be destroyed by an earthquake. Such point clouds are usually static, colored, and large in quantity.
[0023] [[ID=...]] Another use case is in topography and cartography where 3D representations and maps are not limited to a flat plane and can include elevation. Google Maps is an example of a |D map that uses a mesh instead of point clouds. Nevertheless, point clouds can be a suitable data format for 3D maps, and such point clouds are usually static, colored, and relatively large in quantity.
[0024] World modeling and perception via point clouds is a technology that enables machines to gain knowledge about the 3D world around them and is useful for the applications discussed above.
[0025] 3D point cloud data is essentially discrete samples on the surface of an object or scene. In reality, a large number of points are required to fully represent the real world with point samples. Therefore, the processing of such large-scale point clouds is computationally costly, especially for consumer devices with limited computing power, such as smartphones, tablets, and automotive navigation systems.
[0026] To process the input point cloud at a reasonable computational cost, one solution is to first downsample it, and the downsampled point cloud summarizes the geometry of the input point cloud while having significantly fewer points. The downsampled point cloud is then fed into subsequent machine tasks for further consumption. However, point cloud data can be utilized for various tasks such as scene flow estimation, classification, detection, segmentation, and compression. Different tasks focus on different aspects of the point cloud. For example, classification depends on geometric salient points, object segmentation needs to distinguish points on one object from others, and scene flow estimation counts the dynamics of the point cloud. Therefore, a task-aware adaptive point cloud downsampling algorithm is useful. Thus, when facing different tasks, the same point cloud can be downsampled differently to facilitate subsequent tasks.
[0027] FIG. 1 shows a method 10 for downsampling an input point cloud X at n points for subsequent instrument tasks according to the present principle. In step 11, a first downsampled point cloud having m points (m < n) is selected. Using any applicable method, a set of m points such as point 110 of the input point cloud is selected. In step 12, for the points 110 (referred to herein as "anchor points") within the first downsampled point cloud, the neighboring points are aggregated from the point cloud X to obtain a local point set 120. In this way, each anchor point within the first downsampled point cloud is associated with a local point set from the point cloud X. In step 13, each point set is fed into a module called the Set Distillation (SD) function herein, resulting in a representative point 130 and its corresponding set-level feature.
[0028] According to this principle, given a set of points (and other auxiliary information, if available), the SD function first computes a point-level feature vector for each point in the set and a set-level feature vector describing the entire set. This step is achieved, for example, using a structured neural network module according to this principle (referred to herein as P-Net). Representative positions are computed by taking the point-level and set-level feature vectors for each point as input. This step is achieved through either a deterministic approach or another neural network module. The SD function then outputs the representative positions along with the set-level features to represent the geometry of the set of points. By using the SD function, the representative positions obtained are not limited to points within the set of points.
[0029] In step 14, the m representative points are aggregated as an updated, downsampled point cloud, which is then fed into subsequent tasks for further processing. Optionally, m set-level features are also output and fed into subsequent tasks.
[0030] The downsampling method 10 is integrated with subsequent tasks and trained end-to-end, enabling the downsampling method 10 to be task-recognized, i.e., adaptive to machine tasks. Meanwhile, through end-to-end training, the downsampled point clouds obtained by method 10 can capture the underlying geometry for a particular machine task, regardless of how the original input point cloud is sampled from the scene. Specifically, given two different point clouds sampling the same surface, i.e., one point cloud being a resampled version of the other, method 10 yields two downsampled point clouds that are very similar to each other for the same subsequent machine task.
[0031] Figure 2 schematically shows an example of an SD function. Given a set of points 20, the SD function supplies the set of points to the PointNet architecture, as described, for example, in "PointNet: Deep learning on point sets for 3D classification and segmentation" by CRQi, H.Su, K.Mo, and LJGuibas, in proc. IEEE Conference on Computer Vision and Pattern Recognition, pp. 652-660, 2017. The PointNet module 21 computes point-level feature vectors 22 for each point using shared multi-layer perception (MLP). These point-level feature vectors 22 are then aggregated in a maximum pooling operation 23 to yield a set-level feature vector 24 that describes the entire set of points.
[0032] According to this principle, a set of weights 26 is calculated for each point in the entire point set. To do this, affinity values (e.g., weight estimates) between each point-level feature vector 22 and the set-level feature vector 24 are provided by module 25, which calculates the dot product between them. This affinity value describes the extent to which the associated point represents the entire point set. Module 27 generates representative locations for the point set by performing a weighted average 28 of the points using the calculated weights 26. The affinity values are converted into a set of weights using the Softmax(·) function, so that all weight values are greater than 0 and summed to 1. A weighted average of the x-coordinates of all points in the point set is performed using the obtained weights to obtain the x-coordinate of the generated representative point. Similarly, the y-coordinate and z-coordinate of the representative point are calculated using the weights. The generated x, y, and z coordinates form the location of the representative point. The SD function outputs the representative point and the set-level features generated by PointNet.
[0033] In a first embodiment of this principle, downsampling of a given point cloud X containing n points is performed using the presented SD function. In the first step, the initial downsampled point cloud with m points is generated using the Farthest point sampling (FPS) method, and the resulting points are called "anchor points". Farthest point sampling is a known point cloud downsampling technique, described, for example, in "The Farthest point strategy for progressive image sampling", IEEE Trans. on Image Processing, vol.6, no.9, pp.1306-1315, 1997. FPS is based on iteratively selecting the next sample point within a minimum search area. Given a point cloud X and its sampled subset, the FPS algorithm selects the point furthest from the remaining points in X to the subset using some distance measure. This farthest point is then added to the subset, where the subset is initialized by randomly picking points from X. The FPS algorithm repeats this point selection process until certain conditions are met, for example, until the number of points in a subset reaches a predefined threshold. This classical sampling method is deterministic and does not consider downstream tasks.
[0034] In the second step of the first embodiment, for each anchor point, its neighbors are collected through a ball query procedure. That is, all points in X that are within a predefined distance r from the anchor point are identified and collected to form a local point set for that anchor point. In the third step, all local point sets (total m) are individually fed into the SD function to obtain an updated downsampled point cloud (having m points) along with m set-level features. In the fourth step, the m downsampled points (and optionally, set-level features) are fed into a subsequent task. This downsampling method is trained end-to-end with the subsequent machine task to enable the neural network layer in the SD function to recognize the task, i.e., to adapt to the subsequent task.
[0035] In the second embodiment, the calculation of point-by-point weights in the SD function is different. Specifically, in the SD function of this embodiment, a distance is calculated for each point in the set of points, which is the Euclidean distance between its point-level feature vector and the set-level feature vector. In this specification, this distance value is given by d for i. i This is shown by d. i The values are plugged into a Gaussian kernel to calculate the weights, that is, it has a constant σ.
[0036]
number
[0037] In the third embodiment, each downsampled point is obtained by selecting a critical point within a local point set. The difference between this embodiment and the first embodiment is the SD function, which selects a representative point from the input point set. As in the first embodiment, given a point set, the SD function calculates a set of weights for each point in the entire set. The SD function then returns the point with the highest weight directly as the representative point, as well as the set-level features generated by PointNet. Figure 3 shows an example where point A has the highest weight and is therefore selected as the representative point. In a variation of the third embodiment, the method in the second embodiment for calculating the weights of points in the point set, where the weights are obtained through a Gaussian kernel, may also be used.
[0038] In the fourth embodiment, the SD function takes as input not only a local set of points from the point cloud X, but also one-hot vectors indicating which points are anchor points in the set. As a result, the SD function in this embodiment can utilize knowledge of anchor point locations to generate representative points. In particular, the SD function expands the position vector of each point in the set of points by appending the anchor point position vector (and, if available, the anchor feature vector, which is another input to the SD function) before the PointNet calculation. With the anchor position information appended, the expanded set of points is processed by PointNet to obtain point-level and set-level feature vectors for each point.
[0039] Figure 4 shows a fifth embodiment of downsampling the input point cloud according to the present principle. Instead of generating representative points by weighted averaging, this embodiment directly modifies the position of the anchor point 41 and then returns the modified position 42 as the representative point. Specifically, as in the fourth embodiment, the SD function in this embodiment also takes as input a local point set 20 and a one-hot vector indicating which point is the anchor point 41 of the point set. Once the point-level feature vector 22 and set-level features 24 are obtained, they are fed into another neural network 43, which is referred to herein as "M-Net". Specifically, the M-Net outputs a modification vector 44 for the anchor point position. This may be implemented using a PointNet architecture. The representative point position 42 is obtained by adding the modification vector 44 and the anchor position 41. Finally, the SD function still returns the representative point and set-level feature vectors. The fifth embodiment can be combined with the fourth embodiment in which the points fed into the SD function are first expanded with anchor position information.
[0040] Figure 5 schematically illustrates how the task-recognition point cloud downsampling method of this principle is integrated with subsequent machine tasks. As an example, without loss of generality, the task of scene flow estimation for a 3D point cloud is considered in order to illustrate this sixth embodiment. This task takes two consecutive 3D point cloud frames in a point cloud sequence, e.g., a first point cloud frame 51 and a second point cloud frame 52, as input and aims to estimate the scene flow from the first point cloud frame to the second point cloud frame, i.e., the movement of each 3D point from the first point cloud frame to the second point cloud frame. The difficulty is that the index of the points is lost from one frame to consecutive frames. In this scenario, the output scene flow 53 includes a set of 3D vectors, each 3D vector associated with a point in the first point cloud frame. The 3D vectors represent how the points from the first point cloud frame physically move to the surface of the second point cloud frame. In other words, the scene flow between two point cloud frames represents the point cloud dynamics, which are essential for many practical applications, such as autonomous driving, AR / VR, and robotics.
[0041] In this sixth embodiment, the downsampling method presented in the previous embodiment is applied multiple times. The overall neural network architecture of this embodiment takes the form of an hourglass structure with skip connections. The method of this sixth embodiment each includes a first step of generating first and second downsampled point clouds from first and second point cloud frames. This is achieved using two task-aware downsampling modules 54a and 54b (based on any one of the previous embodiments) for both inputs. Two consecutive point clouds in a point cloud sequence may be considered as a single point cloud carrying temporal information indicating whether a point belongs to the first point cloud or the second point cloud. In fact, two point clouds in a point cloud sequence share the same reference frame, and their points can be merged into a single point cloud. In the second step, each set of points in the first downsampled point cloud is aggregated by searching for its nearest neighbors from the second downsampled point cloud. This method calculates a first interframe feature for each point in the first downsampled point cloud, using the point information (its position and point-level features) and its associated nearest neighbor set to fuse information from both point cloud frames. This second stage is achieved using a neural network module 55, referred to herein as "F1-Net". In the third stage, the first downsampled point cloud is further downsampled using a task recognition downsampling module 54c based on this principle, taking the points and associated interframe features as input. In the fourth stage, a second interframe feature is calculated for each point in the first point cloud frame using an upsampled neural network module 56, referred to herein as "F2-Net". This F2-Net module corresponds to stacking setup transformation layers, for example, as presented in "FlowNet3D: Learning scene flow in 3D point clouds", in proc. IEEE Conference on Computer Vision and Pattern Recognition, pp. 529-537, 2020.Such neural network modules hierarchically interpolate point-by-point features. In the fifth stage, a scene flow vector is calculated for each point in the first point cloud frame using a feature-to-flow transform neural network module 57, referred to herein as "F3-Net". F3-Net is implemented using point-by-point MLP layers. According to this principle, a skip connection between the task recognition downsampling module and F2-Net is added to merge information from the initial layers.
[0042] The entire neural network architecture shown in Figure 5, including the task-recognition downsampling modules 54a-c and other neural network modules, is trained end-to-end using an end-point error (EPE) loss function as described in "FlowNet3D". After the training process, the proposed task-recognition downsampling modules integrate well with the other neural network modules, which facilitates accurate scene flow vector estimation.
[0043] In Figure 5, the integration of the task-recognition point cloud downsampling method is explained in relation to the scene flow estimation task. The same principle may be applied to any other task dealing with point clouds as described above without loss of generality.
[0044] Figure 6 shows a seventh embodiment of the integrated task-recognition point cloud downsampling method. In this seventh embodiment, a scene flow 53 is estimated based on two input point cloud frames 51 and 52. The method iteratively updates the estimated scene flow to fine-tune its accuracy through a flow interpolation module 61. An initial point-by-point scene flow for the first point cloud frame 51 is estimated using the method described above with respect to Figure 5. Based on this initial scene flow, a point-by-point scene flow is generated for the first downsampled point cloud. This is achieved based on a scene flow interpolation neural network module 61, referred to herein as "I-Net". I-Net is implemented in the same manner as the setup transformation layer, as described in "FlowNet3D". The shifted downsampled point cloud is generated by shifting each point in the first downsampled point cloud by its associated scene flow vector. Then, for each point in the shifted downsampled point cloud, the set of points is aggregated by searching for its nearest neighbor from the second downsampled point cloud. In this way, each point in the first downsampled point cloud is associated with its shifted version, as well as an updated set of nearest neighbors that uses the shifted version as a query point. These updated sets of nearest neighbors are more accurate / useful for scene flow estimation.
[0045] The second point-by-point scene flow for the first point cloud frame is obtained by running F1-Net 55, the second task-aware downsampling module 54c, F2-Net 56, and F3-Net 57 again using the first downsampled point cloud and all updated nearest neighbor sets. An alternative to this step is to run F1-Net, the second task-aware downsampling module, F2-Net, and F3-Net again based on the shifted downsampled point cloud and updated nearest neighbor sets to obtain the point-by-point scene flow of the residuals. The second point-by-point scene flow for the first point cloud frame can be obtained by adding the point-by-point scene flow of the residuals to the initial point-by-point scene flow. Finally, the second point-by-point scene flow is output as the result. This recursive scene flow estimation scheme can be run iteratively over more than two iterations until certain conditions are met, for example, until the number of iterations reaches a predetermined threshold.
[0046] To this seventh embodiment, the integration of a task-recognition point cloud downsampling method is described in relation to scene flow estimation. The same iterative principle may be applied to any other task dealing with point clouds as described above without loss of generality.
[0047] Figure 7 shows a method of point cloud compression using an embodiment of the task recognition point cloud downsampling method according to the present principle. In this embodiment of the encoder, the downsampled point cloud is used to construct a predicted point cloud for a predictive coding task. Given an input point cloud to be encoded X, the downsampled point cloud is generated using the task recognition point cloud downsampling method 71, as described in relation to one of the embodiments described above. The downsampled point cloud, and optionally the generated set-level feature vector, are encoded by a first entropy encoder 72 to obtain a first bitstream BS1. On the other hand, the downsampled point cloud, and optionally the generated set-level feature vector, are used to construct a predicted point cloud X close to X. PThe predictor construction module 73 attempts to generate the residual point group X. Finally, the second entropy encoder 74 supplies the residual point group X to the predictor construction module 73. R =XX P The first bitstream is encoded to obtain a second bitstream BS2. The two bitstreams are sent together to the decoder. The entropy encoders 72 and 74 can be either reversible or irreversible.
[0048] Figure 8 shows an embodiment of the decoder of this principle. The downsampled point cloud (and feature vectors, if available) is decoded from the first bitstream BS1 by the decoder module 81 and supplied to the predictor construction module 82 to predict the point cloud
[0049]
number
[0050]
number
[0051]
number
[0052]
number
[0053] This decoder embodiment can be used for either inter-frame predictive coding or intra-frame predictive coding. This differs from conventional scalable coding in two ways. On the one hand, the decoder does not restrict the downsampled point cloud to being a subset of the input point cloud. On the other hand, a predicted point cloud can also be generated using feature vectors produced by the task-aware downsampling module of the encoder associated with Figure 7, separately from the downsampled point cloud, thereby providing further flexibility to this predictive coding scheme.
[0054] Figure 9 shows an exemplary architecture of device 30 that may be configured to implement the methods described in relation to Figures 1, 5, 6, 7, and 8. Different embodiments of encoders and decoders according to the present principle may implement this architecture. Alternatively, the encoder and / or decoder circuits according to the present disclosure may be connected together, for example, via their bus 31 and / or via the I / O interface 36, in a device according to the architecture of Figure 9.
[0055] Device 30 consists of the following elements, which are linked together by the data and address bus 31: For example, a microprocessor 32 (or CPU) which is a DSP (Digital Signal Processor), ROM (Read Only Memory) 33 and, • RAM (Random Access Memory) 34 and • Storage interface 35, • An I / O interface 36 for receiving data to be sent from the application, It includes a power source, such as a battery (not shown).
[0056] For example, the power supply is external to the device. In each of the memories mentioned herein, the term “register” as used herein may refer to a small area of capacity (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). ROM33 contains at least a program and parameters. ROM33 can store algorithms and instructions for performing the technology according to this principle. When switched on, CPU32 uploads the program in RAM and executes the corresponding instructions.
[0057] RAM34 contains, within its registers, a program executed by CPU32 and uploaded after device30 is switched on, input data within the registers, intermediate data for different states of methods within the registers, and other variables used for executing methods within the registers.
[0058] The implementations described herein may be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even if considered only in the context of a single implementation (e.g., considered only as a method or device), the implementations of the considered features may also be implemented in other forms (e.g., programs). Apparatus may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented in apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as computers, mobile phones, and portable / mobile information terminals ("Personal Digital Assistants, PDAs"), which facilitate the communication of information between end users.
[0059] According to the examples of this disclosure, device 30 is • Mobile devices and, • Communication devices and, • Gaming devices and, • A tablet (or tablet computer) and • Laptop and, For example, a still camera or video camera equipped with a depth sensor, • A rig for a still camera or video camera, • Encoding chip, Belongs to a set that includes a server (e.g., a broadcast server, video-on-demand server, or web server).
[0060] The various processes and features described herein may be embodied in a variety of different devices or applications, particularly in devices or applications associated with other processing of images and associated texture information and / or depth information, for example, data encoding, data decoding, view generation, texture processing, and related texture information and / or depth information. Examples of such devices include encoders, decoders, post-processors that process the output from decoders, pre-processors that provide input to encoders, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be clear, the devices may be mobile and may be installed in mobile vehicles.
[0061] In addition, the method may be implemented by instructions executed by the processor, and such instructions (and / or data values produced by the implementation) may be stored on a processor-readable medium such as an integrated circuit, software carrier or other storage device, e.g., a hard disk, a compact diskette ("CD"), an optical disc (e.g., a DVD, often referred to as a digital multipurpose disc or digital video disc), random access memory ("RAM") or read-only memory ("ROM"). The instructions may form an application program explicitly embodied on the processor-readable medium. The instructions may be, for example, hardware, firmware, software, or a combination of the two. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. Thus, a processor may be characterized as both, for example, a device configured to execute a process and a device containing a processor-readable medium (such as a storage device) having instructions for executing the process. Furthermore, the processor-readable medium may store data values produced by the implementation in addition to, or instead of, instructions.
[0062] As will be apparent to those skilled in the art, the implementations can produce a variety of signals formatted to carry, for example, information that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry, as data, rules for writing or reading the syntax of a described embodiment, or actual syntax values written by a described embodiment. Such signals may be formatted, for example, as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as baseband signals. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The signals carried by the signals may be, for example, analog or digital information. Signals may be transmitted by a variety of different wired or wireless links, as is known. Signals may be stored in a processor-readable medium.
[0063] Many implementations are described. Nevertheless, it will be understood that various modifications are possible. For example, elements of different implementations can be combined, supplemented, modified, or deleted to create other implementations. In addition, those skilled in the art will understand that other structures and processes can be substituted for those disclosed, and that the resulting implementations will achieve at least substantially the same results as the disclosed implementations, performing at least substantially the same functions in at least substantially the same manner. Accordingly, these and other implementations are contemplated in this application.
Claims
1. A method, for point clouds, For each point in the point cloud, a point-level feature vector is generated by inputting the position vector of the point into a neural network, and a set-level feature vector of the point cloud is generated by using the neural network. Using a set distillation (SD) function, the affinity value between each point-level feature vector and the set-level feature vector is calculated, The SD function determines the representative position of the point cloud as a weighted combination of points in the point cloud using the affinity value, A method comprising outputting the representative position and the set-level feature vector as a set descriptor of the point cloud.
2. To generate the aforementioned representative position, The weight coefficient of each point in the point cloud is calculated by calculating the dot product based on the similarity measure between the point-level feature vector and the set-level feature vector, The method according to claim 1, comprising generating the representative position as a weighted average of all points using the weight coefficients.
3. To generate the point-level feature vector and the set-level feature vector, Accessing the anchor points of the point cloud by performing furthest point sampling, For each point in the point cloud, an extended point is generated by adding the position vector of the anchor point to the position vector of the point. The method according to claim 1, comprising using the extension points as inputs to the neural network.
4. To generate the aforementioned representative position, By performing sampling at the farthest point, the anchor position of the point cloud is accessed, The process involves generating a correction vector for the anchor position by an extended neural subnetwork that is trained end-to-end with subsequent machine tasks and adjusts the anchor position in a task-specific manner, The method according to claim 1, comprising generating the representative position by adding the correction vector to the anchor position.
5. The method according to claim 1, wherein the point cloud is first downsampled by a task-recognition downsampling function that includes the SD function trained in conjunction with a subsequent neural network for a predetermined downstream task, thereby making the downsampled point cloud adaptive to the downstream task.
6. The method according to claim 5, wherein the task recognition downsampling function is integrated with a subsequent machine task, and the integration is achieved by jointly training the task recognition downsampling function and the subsequent machine task in an end-to-end manner.
7. The downstream task is a predictive coding or compression task, and the downsampled point cloud is Encoded by the first entropy coding method, It is supplied to the predictor building module to obtain the predicted point cloud, The method according to claim 6, further comprising encoding a residual point group, which is the difference between the point group and the predicted point group, by a second entropy coding method.
8. The method according to claim 1, wherein the representative position is generated without establishing an explicit correspondence between the points.
9. A method for extracting point clouds from a data stream, From the aforementioned data stream, the downsampled point cloud, residual point cloud, and set-level feature vectors are obtained. The downsampled point cloud and the set-level feature vector are supplied to a predictor building module trained with a task recognition downsampling function to obtain a predicted point cloud. A method comprising: extracting the point cloud by adding the predicted point cloud to the residual point cloud.
10. A device comprising a processor associated with memory, wherein the processor performs the following actions on a point cloud: For each point in the point cloud, a point-level feature vector is generated by inputting the position vector of the point into a neural network, and a set-level feature vector of the point cloud is generated by using the neural network. The affinity value between each point-level feature vector and the set-level feature vector is calculated using the set distillation (SD) function. The SD function determines the representative position of the point cloud as a weighted combination of points in the point cloud using the affinity value, A device configured to output the aforementioned representative position and the aforementioned set-level feature vector as a set descriptor of the point cloud.
11. The aforementioned processor, The weight coefficient of each point in the point cloud is calculated by calculating the dot product based on the similarity measure between the point-level feature vector and the set-level feature vector, The device according to claim 10, configured to generate the representative position by generating the representative position as a weighted average of all points using the weight coefficients.
12. The aforementioned processor, Accessing the anchor points of the point cloud by performing furthest point sampling, For each point in the point cloud, an extended point is generated by adding the position vector of the anchor point to the position vector of the point. The device according to claim 10, configured to generate the point-level feature vector and the set-level feature vector by using the extension points as input to the neural network.
13. The aforementioned processor, By performing sampling at the farthest point, the anchor position of the point cloud is accessed, The process involves generating a correction vector for the anchor position by an extended neural subnetwork that is trained end-to-end with subsequent machine tasks and adjusts the anchor position in a task-specific manner, The device according to claim 10, configured to generate the representative position by adding the correction vector to the anchor position, and by performing the following:
14. The device according to claim 10, wherein the processor is further configured to first downsample the point cloud by using a task-aware downsampling function which includes the SD function trained in conjunction with a subsequent neural network for a predetermined downstream task, thereby making the downsampled point cloud adaptive to the downstream task.
15. The device according to claim 14, wherein the task recognition downsampling function is integrated with a subsequent machine task, and the integration is achieved by jointly training the task recognition downsampling function and the subsequent machine task in an end-to-end manner.
16. The downstream task is a predictive coding or compression task, and the downsampled point cloud is Encoded by the first entropy coding method, It is supplied to the predictor building module to obtain the predicted point cloud, The device according to claim 15, wherein the processor is further configured to encode a residual point cloud, which is the difference between the point cloud and the predicted point cloud, by a second entropy coding method.
17. The device according to claim 10, wherein the representative position is generated without establishing an explicit correspondence between points.
18. A device for extracting point clouds from a data stream, wherein the device comprises a processor associated with memory, and the processor is From the aforementioned data stream, the downsampled point cloud, residual point cloud, and set-level feature vectors are obtained. The downsampled point cloud and the set-level feature vector are supplied to a predictor building module trained with a task recognition downsampling function to obtain a predicted point cloud. A device configured to extract the point cloud by adding the predicted point cloud to the residual point cloud.
Citation Information
Patent Citations
Method for encoding a point cloud representing a scene, encoder system, and non-transitory computer-readable recording medium having a program stored thereon
JP2019521417A