Image processing apparatus and image processing method
By dividing local areas in three-dimensional space and using representative values for filtering, the problem of increasing point cloud data processing time is solved, and efficient filtering is achieved.
Patent Information
- Application Number
- CN202510109041.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-11
- Filing Date
- 2019-06-27
- Publication Date
- 2025-05-13
AI Technical Summary
When filtering point cloud data is performed in the prior art, the processing time is easily increased, especially since the point cloud contains a large number of points, the processing burden of nearest neighbor search is heavy.
By dividing local areas of three-dimensional space, point cloud data is filtered using representative values of each local area, and two-dimensional planar images of point cloud data subjected to the filtering process are encoded to generate a bit stream.
High-efficiency filtering of point cloud data is realized, reducing the increase in processing time and improving processing speed.
Smart Images

Figure CN119996679A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application number 201980045106.4, application date June 27, 2019, and invention name “Image processing device and image processing method”. Technical Field
[0002] The present disclosure relates to an image processing device and an image processing method, and more particularly to an image processing device and an image processing method capable of suppressing an increase in processing time of a filter process for point cloud data. Background Art
[0003] Conventionally, as a method for encoding 3D data (eg, point cloud) representing a three-dimensional structure, encoding has been performed using voxels (eg, octree) (eg, see Non-Patent Literature 1).
[0004] In recent years, as another encoding method, for example, a method has been proposed in which position and color information about a point cloud is projected onto a two-dimensional plane for each small area, respectively, and encoded by an encoding method for two-dimensional images (hereinafter, also referred to as a video-based method) (for example, see non-patent documents 2 to 4).
[0005] In such encoding, in order to suppress degradation of subjective image quality when imaging a point cloud restored from a decoded two-dimensional image, a method of acquiring peripheral points by nearest neighbor search and applying a three-dimensional smoothing filter has been considered.
[0006] Citation List
[0007] Non-patent literature
[0008] Non-patent document 1: R. Mekuria, Student Member IEEE, K. Blom, P. Cesar., Member, IEEE, "Design, Implementation and Evaluation of a Point Cloud Codec for Tele-Immersive Video", tcsvt_paper_submitted_february.pdf
[0009] Non-patent document 2: Tim Golla and Reinhard Klein, “Real-time Point Cloud Compression,” IEEE, 2015
[0010] Non-patent document 3: K. Mammou, “Video-based and Hierarchical Approaches Point Cloud Compression”, MPEG m41649, October 2017
[0011] Non-patent document 4: K. Mammou, “PCC Test Model Category 2v0”, N17248 MPEG output document, October 2017 Summary of the invention
[0012] Problems to be solved by the present invention
[0013] However, in general, a point cloud contains a large number of points, and the processing load of the nearest neighbor search becomes very heavy. For this reason, there is a possibility that this method will increase the processing time.
[0014] The present disclosure has been made in view of such circumstances, and an object of the present disclosure is to enable filter processing to be performed on point cloud data at a higher speed than conventional methods and to suppress an increase in processing time.
[0015] Solution to the problem
[0016] An image processing device according to one aspect of the present technology is an image processing device including: a filtering processing unit that performs filtering processing on point cloud data using a representative value of the point cloud data of each local area obtained by dividing a three-dimensional space; and an encoding unit that encodes a two-dimensional plane image on which the point cloud data subjected to filtering processing by the filtering processing unit is projected and generates a bit stream.
[0017] An image processing method according to one aspect of the present technology is an image processing method comprising: performing filtering processing on point cloud data using a representative value of the point cloud data of each local area obtained by dividing a three-dimensional space; and encoding a two-dimensional plane image on which the point cloud data subjected to filtering processing is projected and generating a bit stream.
[0018] An image processing device on the other hand of the present technology is an image processing device, which includes: a decoding unit that decodes a bit stream and generates encoded data of a two-dimensional plane image on which point cloud data is projected; and a filtering processing unit that uses a representative value of the point cloud data of each local area obtained by dividing the three-dimensional space to perform filtering processing on the point cloud data restored from the two-dimensional plane image generated by the decoding unit.
[0019] The image processing method on the other hand of the present technology is an image processing method, comprising: decoding a bit stream and generating encoded data of a two-dimensional plane image on which point cloud data is projected; and using a representative value of the point cloud data of each local area obtained by dividing the three-dimensional space, performing filtering processing on the point cloud data recovered from the generated two-dimensional plane image.
[0020] An image processing device on the other hand of the present technology is an image processing device, which includes: a filtering processing unit that performs filtering processing on some points in point cloud data; and an encoding unit that encodes a two-dimensional plane image on which the point cloud data subjected to filtering processing by the filtering processing unit is projected and generates a bit stream.
[0021] An image processing method according to another aspect of the present technology is an image processing method including: performing filtering processing on some points in point cloud data; and encoding a two-dimensional plane image on which the point cloud data subjected to filtering processing is projected and generating a bit stream.
[0022] An image processing device according to another aspect of the present technology is an image processing device including: a decoding unit that decodes a bit stream and generates encoded data of a two-dimensional plane image on which point cloud data is projected; and a filtering processing unit that performs filtering processing on some points in the point cloud data restored from the two-dimensional plane image generated by the decoding unit.
[0023] An image processing method according to another aspect of the present technology is an image processing method, which includes: decoding a bit stream and generating encoded data of a two-dimensional plane image on which point cloud data is projected; and performing filtering processing on some points in the point cloud data recovered from the generated two-dimensional plane image.
[0024] In an image processing device and an image processing method of one aspect of the present technology, filtering processing is performed on the point cloud data using a representative value of the point cloud data for each local area obtained by dividing the three-dimensional space, and a two-dimensional plane image projected by the point cloud data subjected to the filtering processing is encoded, and a bit stream is generated.
[0025] In an image processing device and an image processing method on another aspect of the present technology, a bit stream is decoded and encoded data of a two-dimensional plane image on which point cloud data is projected is generated, and filtering processing is performed on the point cloud data recovered from the generated two-dimensional plane image using representative values of the point cloud data for each local area obtained by dividing the three-dimensional space.
[0026] In an image processing device and an image processing method of another aspect of the present technology, filtering processing is performed on some points in point cloud data, and a two-dimensional plane image on which the point cloud data subjected to filtering processing is projected is encoded, and a bit stream is generated.
[0027] In an image processing device and an image processing method of another aspect of the present technology, a bit stream is decoded, and encoded data of a two-dimensional plane image on which point cloud data is projected is generated, and filtering processing is performed on some points in the point cloud data restored from the generated two-dimensional plane image.
[0028] Beneficial effects of the present invention
[0029] According to the present disclosure, it is possible to process an image and, in particular, suppress an increase in processing time for filter processing on point cloud data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a diagram for explaining an example of smoothing processing.
[0031] Figure 2 is a diagram summarizing the main features associated with the present technology.
[0032] Figure 3 is a graph illustrating a nearest neighbor search.
[0033] Figure 4 This is a diagram illustrating an example of the outline of the filtering process using the present technique.
[0034] Figure 5 It is a diagram for explaining comparison of processing time.
[0035] Figure 6 is a diagram illustrating an example of a local area division technique.
[0036] Figure 7 It is a diagram illustrating parameters related to a local area.
[0037] Figure 8 This is a diagram for explaining the transmission of information.
[0038] Fig. 9 It is a diagram for explaining the target of filtering processing.
[0039] Fig.10 This is a diagram for explaining a method for obtaining a representative value.
[0040] Fig.11 It is a diagram for explaining arithmetic operations of filtering processing.
[0041] Fig.12 It is a diagram for explaining the target range of the filtering process.
[0042] Fig.13 It is a diagram for explaining a case where filtering processing using a nearest neighbor search is applied.
[0043] Fig.14 It is a diagram for explaining a case where filtering processing using a representative value for each local area is applied.
[0044] Fig.15 It is a diagram for explaining comparison of processing time.
[0045] Fig.16 is a block diagram showing a main configuration example of an encoding device.
[0046] Fig.17 is a diagram illustrating a main configuration example of a patch decomposition unit.
[0047] Fig.18 : is a diagram illustrating a main configuration example of a three-dimensional position information smoothing processing unit.
[0048] Fig.19 is a flowchart illustrating an example of the flow of encoding processing.
[0049] Fig. 20 : is a flowchart illustrating an example of the flow of patch decomposition processing.
[0050] Fig.21 is a flowchart illustrating an example of the flow of the smoothing process.
[0051] Fig. 22 is a flowchart illustrating an example of the flow of the smoothing range setting process.
[0052] Fig.23 is a block diagram showing a main configuration example of a decoding device.
[0053] Fig.24 is a diagram illustrating a main configuration example of a 3D reconstruction unit.
[0054] Fig.25 : is a diagram illustrating a main configuration example of a three-dimensional position information smoothing processing unit.
[0055] Fig.26 is a flowchart for explaining an example of the flow of decoding processing.
[0056] Fig. 27 is a flowchart illustrating an example of the flow of point cloud reconstruction processing.
[0057] Fig.28 is a flowchart illustrating an example of the flow of the smoothing process.
[0058] Fig.29 is a block diagram showing a main configuration example of a computer. DETAILED DESCRIPTION
[0059] Modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described below. Note that the description is given in the following order.
[0060] 1. Accelerate filtering processing
[0061] 2. First Embodiment (Encoding Device)
[0062] 3. Second Embodiment (Decoding Device)
[0063] 4. Variant
[0064] 5. Additional Notes
[0065] <1. Accelerated filtering processing>
[0066] <Documents supporting technical content and terminology, etc.>
[0067] The scope disclosed in the present technology includes not only the contents described in the embodiments but also the contents described in the following non-patent documents known at the time of filing.
[0068] Non-patent document 1: (as described above)
[0069] Non-patent document 2: (as described above)
[0070] Non-patent document 3: (as described above)
[0071] Non-patent document 4: (as described above)
[0072] Non-patent document 5: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union), "Advanced video coding for generic audiovisual services", H.264, 04 / 2017
[0073] Non-patent document 6: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union), "High efficiency video coding", H.265, 12 / 2016
[0074] Non-Patent Literature 7: Jianle Chen, Elena Alshina, Gary J. Sullivan, Jens-Rainer, Jill Boyce, “Algorithm Description of Joint Exploration Test Model 4”, JVET-G1001_v1, Joint Video Exploration Team (JVET) of ITU-TSG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 Seventh Meeting: Turin, Italy, July 13 to July 21, 2017
[0075] In other words, the contents described in the above-mentioned non-patent documents are also the basis for explaining the support requirements. For example, even if the quadtree block structure described in non-patent document 6 and the quadtree plus binary tree (QTBT) block structure described in non-patent document 7 are not directly described in the embodiments, these technologies are interpreted as being within the scope of the disclosure of the present technology and satisfying the support requirements of the claims. In addition, similarly, for example, technical terms such as parsing, grammar, and semantics are interpreted as being within the scope of the disclosure of the present technology and satisfying the support requirements of the claims even if they are not directly described in the embodiments.
[0076] <Point Cloud>
[0077] Conventionally, there are data such as a point cloud representing a three-dimensional structure by point cloud position information or attribute information, and a mesh consisting of vertices, edges, and faces and defining a three-dimensional shape using polygonal representation.
[0078] For example, in the case of a point cloud, a spatial structure is expressed as a collection of a large number of points (point cloud). In other words, the data of a point cloud consists of position information and attribute information (e.g., color) about each point in the point cloud. Therefore, the data structure is relatively simple, and any three-dimensional structure can be represented with sufficient accuracy by using a sufficiently large number of points.
[0079] <Overview of Video-Based Methods>
[0080] A video-based method has been proposed in which position information and color information about such a point cloud are projected onto a two-dimensional plane for each small area, respectively, and encoded by a two-dimensional image encoding method.
[0081] In this video-based method, the input point cloud is divided into multiple segments (also called regions), and each region is projected onto a two-dimensional plane. Note that, as described above, the data of the point cloud for each position (i.e., the data of each point) consists of position information (geometry (also called depth)) and attribute information (texture), and the position information and attribute information are projected onto a two-dimensional plane for each region, respectively.
[0082] Each of these segments (also referred to as patches) projected onto the two-dimensional plane is then arranged on the two-dimensional plane image and encoded, for example, by a coding technique for two-dimensional plane images (e.g., Advanced Video Coding (AVC) or High Efficiency Video Coding (HEVC)).
[0083] <Occupied Map>
[0084] When 3D data is projected onto a two-dimensional plane using a video-based method, an occupancy map is generated in addition to a two-dimensional plane image (also called a geometric image) projected with position information and a two-dimensional plane image (also called a texture image) projected with attribute information as described above. The occupancy map is map information indicating the presence or absence of position information and attribute information at each position on the two-dimensional plane. More specifically, in the occupancy map, the presence or absence of position information and attribute information is indicated for each area called precision.
[0085] Since the point cloud (each point of the point cloud) is restored in units of blocks defined by the precision of the occupancy map, the larger the size of the block, the coarser the resolution of the point. Therefore, due to the large size of the precision, there is a possibility that the subjective image quality when imaging the point cloud encoded and decoded by the video-based method will be reduced.
[0086] For example, in the case where a point cloud encoded and decoded by a video-based method is imaged, when the size of the precision is large, a fine notch like a sawtooth is formed at the boundary between the white part and the black part, such as Figure 1 As shown in A, there is a possibility that the subjective image quality will be reduced.
[0087] Therefore, a method has been considered in which points around a point to be processed are acquired by nearest neighbor search (also referred to as nearest neighbor (NN)), and a three-dimensional smoothing filter is applied to the point to be processed using the acquired points. By applying such a three-dimensional smoothing filter, as Figure 1 As shown in B, notches at the boundary between the white portion and the black portion are suppressed and a smooth linear shape is obtained, making it possible to suppress a decrease in subjective image quality.
[0088] However, in general, a point cloud contains a large number of points, and the processing load of the nearest neighbor search becomes very heavy. For this reason, there is a possibility that this method will increase the processing time.
[0089] Due to such an increase in processing time, for example, it is difficult to immediately (in real time) perform the video-based method as described above (for example, to encode a moving image at 60 frames per second).
[0090] As general schemes for accelerating NN, a method of searching by approximation (approximate NN), a method using hardware capable of high-speed processing, and the like are considered, but even with these methods, immediate processing is actually difficult.
[0091] <Accelerated 3D filtering processing>
[0092] <#1. Use representative values for each local area for acceleration>
[0093] Therefore, the three-dimensional smoothing filter processing is accelerated. Figure 2 As shown in part #1 in , the three-dimensional space is divided into local areas, a representative value of the point cloud is calculated for each local area, and the representative value for each local area is used as a reference value in the filtering process.
[0094] For example, when the point Figure 3 When the three-dimensional smoothing filter is applied to the black point (curPoint) at the center, the smoothing is performed by referring to the data of the gray points (nearPoint) around the black point (used as a reference value).
[0095] exist Figure 3 The pseudo code of the conventional method is shown in B. In the conventional case, the nearest neighbor search (NN) is used to resolve the peripheral points (nearPoint) of the processing target point (curPoint) (nearPoint=NN(curPoint)), and when all peripheral points do not belong to the same patch (if (!all same patch(nearPoints))), that is, when the processing target point is located at the end of the patch, the average value of the data of the peripheral points (curPoint=average(nearPoints)) is used to smooth the processing target point.
[0096] In contrast, Figure 4 As indicated by the quadrilateral in A of , the three-dimensional space is divided into local areas, a representative value (x) of the point cloud is obtained for each local area, and the processing target point (black point) is smoothed using the obtained representative value. The pseudo code of this process is Figure 4In this case, first, the average value (Average Point) of the points in the local area is obtained as the representative value of each local area (grid). Then, the peripheral grids (nearby grids) located around the grid (processing target grid) to which the processing target point belongs are specified.
[0097] As the peripheral grid, a grid having a predetermined positional relationship pre-established with respect to the processing target grid is selected. For example, a grid adjacent to the processing target grid may be used as the peripheral grid. Figure 4 In case of A, when the square at the center is assumed as the processing target grid, eight grids surrounding the processing target grid are adopted as peripheral grids.
[0098] Then, when all peripheral points do not belong to the same patch (if (!all same patch(nearPoints))), that is, when the processing target point is located at the end of the patch, three-dimensional smoothing filtering is performed on the processing target point (curPoint=trilinear(averagePoints)) by using trilinear filtering of a set of representative values of these peripheral grids (averagePoints=AveragePoint(near grid)).
[0099] By performing the processing in this way, filtering processing (three-dimensional smoothing filtering processing) can be achieved without performing the nearest neighbor search (NN) which carries a large load. Therefore, a smoothing effect equivalent to that of conventional three-dimensional smoothing filtering can be achieved, while the processing time of the filtering processing can be significantly reduced. Figure 5 An example of comparison between the processing time of a three-dimensional smoothing filter (NN) when using the nearest neighbor search and the processing time of a three-dimensional smoothing filter (trilinear) to which the present technique is applied is shown. This shows that by applying the present technique, Figure 5 The processing time required as shown in the figure on the left can be Figure 5 is shortened as shown in the figure on the right.
[0100] In the following, reference will be made to Figures 6 to 15 describe Figure 2 Each part of the.
[0101] <#1-1. Local area division technology>
[0102] The way of dividing the three-dimensional space (partitioning technique of local regions) is optional. For example, the three-dimensional space can be evenly divided into N×N×N cubic regions, such as Figure 6By dividing the three-dimensional space in this way, the three-dimensional space can be easily divided into local areas, so that an increase in the processing time of the filtering process can be suppressed (the filtering process can be accelerated).
[0103] In addition, for example, the three-dimensional space can be evenly divided into M×N×L rectangular regions, as in Figure 6 By dividing the three-dimensional space in this way, the three-dimensional space can be easily divided into local areas, so that the increase in the processing time of the filtering process can be suppressed (the filtering process can be accelerated). In addition, since the degree of freedom of the shape of the local area is increased compared to the case where the three-dimensional space is divided into cubic areas, the processing load can be further smoothed between the local areas (the load imbalance can be suppressed).
[0104] Furthermore, for example, the three-dimensional space can be partitioned so that the number of points in each local area is constant, as Figure 6 By dividing the three-dimensional space in this way, the processing load and resource usage can be smoothed among the local areas (load imbalance can be suppressed), compared to the case where the three-dimensional space is divided into cubic areas or rectangular areas.
[0105] In addition, for example, a local area with any shape and size can be set at any position in three-dimensional space, such as Figure 6 As in the row with "4" in the ID column of the table. By setting the local area in this way, even for an object with a complex three-dimensional shape, smoothing processing more suitable for a specific shape can be performed, and more smoothing can be achieved compared to each of the above methods.
[0106] In addition, for example, selection from among the above-mentioned methods having IDs "1" to "4" can be realized, as in Figure 6 By implementing the selection in this way, more appropriate smoothing processing can be performed in different situations, and more smoothing can be implemented. Note that how to make this selection (based on the selection content) is optional. In addition, information indicating which method has been selected (a signal of method selection information) can be sent from the encoding side to the decoding side.
[0107] <#1-2. Local area parameter settings>
[0108] In addition, the method and content of setting parameters of such local areas are optional. For example, dividing the local area of the three-dimensional space (for example, Figure 6 The shape and size of L, M, N in can have fixed values, such as Figure 7 In the row with "1" in the ID column of the table in . For example, these values can be set in advance according to a standard, etc. By setting the values in this way, setting the shape and size of the local area can be omitted, so that the filtering process can be further accelerated.
[0109] In addition, for example, the shape and size of the local area can be set according to the point cloud and the situation, such as in Figure 7 In other words, the parameters of the local area can be made variable. By adopting a variable parameter in this way, a more appropriate local area can be formed according to the situation, so that the filtering process can be performed more appropriately. For example, the process can be further accelerated, the imbalance in the process can be suppressed, and more smoothing can be achieved.
[0110] For example, the size of the local area (e.g., Figure 6 The L, M, N) in Figure 7 In addition, for example, the number of points included in the local area can be made variable, as in the row with "2-2" in the ID column. In addition, for example, the shape and position of the local area can be made variable, as in the row with "2-3" in the ID column. In addition, for example, the user can be allowed to select the setting method of the local area, as in the row with "2-4" in the ID column. For example, the user can be allowed to decide from Figure 6 Which method is selected from the methods with IDs "1" to "4" in the table.
[0111] <#1-3.Signal>
[0112] In addition, information about the filtering process may be sent from the encoding side to the decoding side or may not be sent from the encoding side to the decoding side. Figure 8 As in the row with "1" in the ID column of the table of , all parameters related to the filtering process can be preset by the standard or the like so that information about the filtering process is not transmitted. By presetting all parameters in this way, since the amount of information to be transmitted is reduced, the encoding efficiency can be improved. In addition, since the derivation of parameters is unnecessary, the load of the filtering process can be reduced, and the filtering process can be further accelerated.
[0113] In addition, for example, Figure 8As in the row with "2" in the ID column of the table of , it is possible to derive the optimal values of all parameters related to the filtering process from other internal parameters (for example, the accuracy of the occupancy map) so that information about the filtering process is not transmitted. By achieving the derivation of the optimal value in this way, the encoding efficiency can be improved because the amount of information to be transmitted is reduced. In addition, a local area more suitable for the situation can be set.
[0114] Furthermore, for example, information about the filtering process can be sent in the header of the bitstream, such as in Figure 8 In that case, the parameter has a fixed value in the bit stream. By transmitting information in the header of the bit stream in this way, the amount of information to be transmitted can be relatively small, so that a decrease in encoding efficiency can be suppressed. In addition, since the parameter has a fixed value in the bit stream, an increase in the load of the filtering process can be suppressed.
[0115] In addition, for example, information about the filtering process can be sent in the header of the frame, such as in Figure 8 In that case, the parameter can be made variable for each frame. Therefore, a local area more suitable for the case can be set.
[0116] <#1-4. Filtering Processing Target>
[0117] The target of the filtering process is optional. For example, the position information of the point cloud can be used as the target, as in Fig. 9 In other words, three-dimensional smoothing filter processing is performed on the position information about the processing target point. By performing smoothing filter processing in this way, smoothing of the positions between the respective points of the point cloud can be achieved.
[0118] In addition, for example, attribute information (color, etc.) related to the point cloud can be used as a target, for example, as in Fig. 9 In other words, three-dimensional smoothing filter processing is performed on the position information about the processing target point. By performing smoothing filter processing in this way, smoothing of colors and the like between individual points in the point cloud can be achieved.
[0119] <#1-5. Method of deriving representative values>
[0120] The method of obtaining a representative value for each local area is optional. Fig.10The average value of the data of the points in the local area (contained in the local area) can be used as the representative value, as in the row with "1" in the ID column of the table of FIG. Since the average value can be calculated by simple arithmetic operations, the representative value can be calculated at a higher speed by using the average value as the representative value in this way. That is, the filtering process can be further accelerated.
[0121] In addition, for example, Fig.10 The median of the data of the points in the local area (included in the local area) can be used as a representative value, as in the row with "2" in the ID column of the table of FIG. Since the median is less susceptible to the influence of unique data, a more stable result can be obtained even in the presence of noise. That is, a more stable filtering processing result can be obtained.
[0122] Of course, the method for deriving the representative value may be different from these examples. In addition, for example, the representative value may be derived by a plurality of methods so that a more favorable value can be selected. In addition, for example, different derivation methods may be allowed for each local area. For example, the derivation method may be selected based on the characteristics of the three-dimensional structure represented by the point cloud. For example, the representative value may be derived by the median of a part with a small shape including a lot of noise (e.g., hair), while the representative value may be derived by the average value of a part with a clear boundary (e.g., clothes).
[0123] <#1-6. Filter processing arithmetic operation>
[0124] The arithmetic operation of the filtering process (three-dimensional smoothing filter) is optional. For example, Fig.11 Trilinear interpolation can be used as in the row with "1" in the ID column of the table of . Trilinear interpolation has a good balance between processing speed and quality of processing results. Alternatively, for example, ternary cubic interpolation can be used, as in Fig.11 The ternary cubic interpolation can obtain a processing result of higher quality than the processing result of the trilinear interpolation. In addition, for example, the nearest neighbor search (NN) can be used, as in Fig.11 This method can obtain processing results at a higher speed than trilinear interpolation. Of course, three-dimensional smoothing filtering can be achieved by any arithmetic operation other than these methods.
[0125] <#2. Simplification of 3D filtering processing>
[0126] In addition, if Figure 2 As shown in part #2 in FIG, filtering processing can be performed only in part of the area. Fig.12 is a diagram showing an example of an occupancy map. Fig.12In the illustrated occupancy map 51, the white portion indicates an area (precision) where there is data in a geometric image where position information related to the point cloud is projected on a two-dimensional plane and where there is data in a texture image where attribute information related to the point cloud is projected on a two-dimensional plane, and the black portion indicates an area where there is no data in the geometric image or the texture image. In other words, the white portion indicates an area where a patch of the point cloud is projected, and the black portion indicates an area where a patch of the point cloud is not projected.
[0127] like Fig.12 As indicated by arrow 52 in Figure 1 The gaps shown in A appear at the boundaries between patches. Figure 2 As shown in part #2-1 in FIG. 1 , the three-dimensional smoothing filter process can be performed only on the points corresponding to such boundary portions between patches (the ends of the patches in the occupancy map). In other words, the ends of the patches in the occupancy map can be used as the partial area on which the three-dimensional smoothing filter process is performed.
[0128] By using the end of the patch as a partial area in this way, the three-dimensional smoothing filter processing can be performed only on some areas. In other words, since the area where the three-dimensional smoothing filter processing is performed can be reduced, the three-dimensional smoothing filter processing can be further accelerated.
[0129] This method can be used with Fig.13 In other words, as in Fig.13 As in the pseudo code shown in B of FIG. 1 , the three-dimensional smoothing filtering process including the nearest neighbor search (k nearest neighbors) may be performed only when the position of the processing target point corresponds to the end of the patch (if (is_Boundary (curPos))).
[0130] In addition, if Fig.14 As shown in A of FIG. 1 , the filtering process described above in #1 can be used in combination with the application of the present technology. In other words, as in Fig.14 As in the pseudo code shown in B of , three-dimensional smoothing filtering processing by trilinear interpolation using representative values of the local area can be performed only when the position of the processing target point corresponds to the end of the patch (if (is_Boundary (curPos))).
[0131] Fig.15An example of comparison of processing time between the various methods is shown. The first figure from the left shows the processing time of smoothing filter processing using conventional nearest neighbor search. The second figure from the left shows the processing time of three-dimensional smoothing filter processing performed by trilinear interpolation using representative values of local areas. The third figure from the left shows the processing time when smoothing filter processing using conventional nearest neighbor search is performed only on points in the occupancy map corresponding to the ends of the patch. The fourth figure from the left shows the processing time when three-dimensional smoothing filter processing by trilinear interpolation using representative values of local areas is performed only on points in the occupancy map corresponding to the ends of the patch. In this way, by performing three-dimensional smoothing filtering only on some areas, the processing time can be reduced regardless of the method of the filtering process.
[0132] <2. First Embodiment>
[0133] <Encoding device>
[0134] Next, a configuration for realizing each scheme as mentioned above will be described. Fig.16 : is a block diagram showing an example of the configuration of an encoding device as an exemplary form of an image processing device to which the present technology is applied. Fig.16 The illustrated encoding device 100 is a device that projects 3D data such as a point cloud onto a two-dimensional plane and encodes the projected 3D data by an encoding method for two-dimensional images (an encoding device to which a video-based method is applied).
[0135] Notice, Fig.16 shows the main parts of the processing units, data flow, etc., and Fig.16 In other words, in the encoding device 100, there may be Fig.16 The processing units are not shown as blocks, or may exist in Fig.16 100. Processing or data flow as arrows, etc. is not shown in FIG. Similarly, this also applies to other drawings illustrating processing units, etc. in the encoding device 100.
[0136] like Fig.16 As shown, the encoding device 100 includes a patch decomposition unit 111, a packing unit 112, an OMap generation unit 113, an auxiliary patch information compression unit 114, a video encoding unit 115, a video encoding unit 116, an OMap encoding unit 117 and a multiplexer 118.
[0137] The patch decomposition unit 111 performs processing related to decomposition of 3D data. For example, the patch decomposition unit 111 acquires 3D data (e.g., point cloud) representing a three-dimensional structure, which has been input to the encoding device 100. In addition, the patch decomposition unit 111 decomposes the acquired 3D data into a plurality of segments to project the 3D data onto a two-dimensional plane of each segment, and generates patches of position information and patches of attribute information.
[0138] The patch decomposition unit 111 supplies information on each generated patch to the packing unit 112. Furthermore, the patch decomposition unit 111 supplies auxiliary patch information as information on decomposition to the auxiliary patch information compression unit 114.
[0139] The packing unit 112 performs processing related to data packing. For example, the packing unit 112 acquires data (patches) of a two-dimensional plane on which 3D data is projected for each region, which is provided from the patch decomposition unit 111. In addition, the packing unit 112 arranges each acquired patch on a two-dimensional image, and packs the obtained two-dimensional image into a video frame. For example, the packing unit 112 packs patches of position information (geometry) indicating the position of a point and patches of attribute information (texture) such as color information added to the position information, respectively, as video frames.
[0140] The packetizing unit 112 supplies the generated video frame to the OMap generating unit 113. Furthermore, the packetizing unit 112 supplies the multiplexer 118 with control information regarding the packetizing.
[0141] The OMap generation unit 113 performs processing related to the generation of an occupancy map. For example, the OMap generation unit 113 acquires data provided from the packing unit 112. In addition, the OMap generation unit 113 generates an occupancy map corresponding to the position information and the attribute information. The OMap generation unit 113 provides the generated occupancy map and various information acquired from the packing unit 112 to a subsequent processing unit. For example, the OMap generation unit 113 provides a video frame of the position information (geometry) to the video encoding unit 115. In addition, for example, the OMap generation unit 113 provides a video frame of the attribute information (texture) to the video encoding unit 116. In addition, for example, the OMap generation unit 113 provides the occupancy map to the OMap encoding unit 117.
[0142] The auxiliary patch information compression unit 114 performs processing related to compression of the auxiliary patch information. For example, the auxiliary patch information compression unit 114 acquires the data provided from the patch decomposition unit 111. The auxiliary patch information compression unit 114 encodes (compresses) the auxiliary patch information included in the acquired data. The auxiliary patch information compression unit 114 provides the obtained encoded data of the auxiliary patch information to the multiplexer 118.
[0143] The video encoding unit 115 performs processing related to encoding of the video frame of the position information (geometry). For example, the video encoding unit 115 obtains the video frame of the position information (geometry) provided from the OMap generation unit 113. In addition, for example, the video encoding unit 115 encodes the obtained video frame of the position information (geometry) by any encoding method (for example, AVC or HEVC) for two-dimensional images. The video encoding unit 115 provides the encoded data obtained by encoding (the encoded data of the video frame of the position information (geometry)) to the multiplexer 118.
[0144] The video encoding unit 116 performs processing related to encoding of the video frame of the attribute information (texture). For example, the video encoding unit 116 acquires the video frame of the attribute information (texture) provided from the OMap generation unit 113. In addition, for example, the video encoding unit 116 encodes the acquired video frame of the attribute information (texture) by any encoding method (for example, AVC or HEVC) for two-dimensional images. The video encoding unit 116 provides the encoded data (encoded data of the video frame of the attribute information (texture)) obtained by encoding to the multiplexer 118.
[0145] The OMap encoding unit 117 performs processing related to encoding of the occupancy map. For example, the OMap encoding unit 117 acquires the occupancy map provided from the OMap generation unit 113. In addition, the OMap encoding unit 117 encodes the acquired occupancy map by any encoding method such as arithmetic coding, for example. The OMap encoding unit 117 supplies the encoded data obtained by the encoding (encoded data of the occupancy map) to the multiplexer 118.
[0146] The multiplexer 118 performs processing related to multiplexing. For example, the multiplexer 118 obtains the encoded data of the auxiliary patch information provided from the auxiliary patch information compression unit 114. In addition, the multiplexer 118 obtains the control information about the packing provided from the packing unit 112. In addition, the multiplexer 118 obtains the encoded data of the video frame of the position information (geometry) provided from the video encoding unit 115. In addition, the multiplexer 118 obtains the encoded data of the video frame of the attribute information (texture) provided from the video encoding unit 116. In addition, the multiplexer 118 obtains the encoded data of the occupancy map provided from the OMap encoding unit 117.
[0147] The multiplexer 118 multiplexes the acquired information to generate a bit stream. The multiplexer 118 outputs the generated bit stream to the outside of the encoding device 100.
[0148] In such an encoding device 100, the patch decomposition unit 111 acquires the occupancy map generated by the OMap generation unit 113 from the OMap generation unit 113. In addition, the patch decomposition unit 111 acquires the encoded data (also referred to as a geometry image) of the video frame of the position information (geometry) generated by the video encoding unit 115 from the video encoding unit 115.
[0149] Then, the patch decomposition unit 111 uses these data to perform three-dimensional smoothing filter processing on the point cloud. In other words, the patch decomposition unit 111 projects the 3D data subjected to the three-dimensional smoothing filter processing onto a two-dimensional plane and generates patches of position information and patches of attribute information.
[0150] <Patch decomposition unit>
[0151] Fig.17 It is shown Fig.16 A block diagram of a main configuration example of the patch decomposition unit 111 in FIG. Fig.17 As shown, the patch decomposition unit 111 includes a patch decomposition processing unit 131 , a geometry decoding unit 132 , a three-dimensional position information smoothing processing unit 133 and a texture correction unit 134 .
[0152] The patch decomposition processing unit 131 acquires a point cloud to decompose the acquired point cloud into a plurality of segments, and projects the point cloud onto a two-dimensional plane for each segment to generate a patch of position information (geometry patch) and a patch of attribute information (texture patch). The patch decomposition processing unit 131 supplies the generated geometry patch to the packing unit 112. In addition, the patch decomposition processing unit 131 supplies the generated texture patch to the texture correction unit 134.
[0153] The geometry decoding unit 132 acquires the encoded data of the geometry image (geometry encoded data). The encoded data of the geometry image is obtained by packing the geometry patches generated by the patch decomposition processing unit 131 into a video frame in the packing unit 112 and encoding the video frame in the video encoding unit 115. The geometry decoding unit 132 decodes the geometry encoded data by a decoding technique corresponding to the encoding technique of the video encoding unit 115. In addition, the geometry decoding unit 132 reconstructs a point cloud (position information about the point cloud) from the geometry image obtained by decoding the geometry encoded data. The geometry decoding unit 132 provides the obtained position information about the point cloud (geometry point cloud) to the three-dimensional position information smoothing processing unit 133.
[0154] The three-dimensional position information smoothing processing unit 133 acquires the position information on the point cloud supplied from the geometry decoding unit 132. In addition, the three-dimensional position information smoothing processing unit 133 acquires an occupancy map. The occupancy map has been generated by the OMap generating unit 113.
[0155] The three-dimensional position information smoothing processing unit 133 performs three-dimensional smoothing filter processing on the position information about the point cloud (geometric point cloud). At this time, as described above, the three-dimensional position information smoothing processing unit 133 performs three-dimensional smoothing filter processing using a representative value for each local area obtained by dividing the three-dimensional space. In addition, the three-dimensional position information smoothing processing unit 133 uses the acquired occupancy map to perform three-dimensional smoothing filter processing only on points in a partial area corresponding to the end of a patch in the acquired occupancy map. By performing the three-dimensional smoothing filter processing in this way, the three-dimensional position information smoothing processing unit 133 can perform filtering processing at a higher speed.
[0156] The three-dimensional position information smoothing processing unit 133 provides the geometric point cloud subjected to filtering processing (also referred to as a smoothed geometric point cloud) to the patch decomposition processing unit 131. The patch decomposition processing unit 131 decomposes the provided smoothed geometric point cloud into a plurality of segments to project the point cloud onto a two-dimensional plane for each segment, and generates patches of position information (smoothed geometric patches) to provide the generated patches to the packing unit 112.
[0157] Furthermore, the three-dimensional position information smoothing processing unit 133 also supplies the smoothed geometric point cloud to the texture correction unit 134 .
[0158] The texture correction unit 134 acquires the texture patch provided from the patch decomposition processing unit 131. In addition, the texture correction unit 134 acquires the smoothed geometric point cloud provided from the three-dimensional position information smoothing processing unit 133. The texture correction unit 134 uses the acquired smoothed geometric point cloud to correct the texture patch. When the position information on the point cloud is changed due to three-dimensional smoothing, the shape of the patch projected on the two-dimensional plane may also change. In other words, the texture correction unit 134 reflects the change in the position information on the point cloud caused by three-dimensional smoothing in the patch of attribute information (texture patch).
[0159] The texture correction unit 134 provides the corrected texture patch to the packing unit 112 .
[0160] The packing unit 112 packs the smoothed geometry patches and the corrected texture patches provided from the patch decomposition unit 111 into video frames, respectively, and generates a video frame of position information and a video frame of attribute information.
[0161] <Three-dimensional position information smoothing processing unit>
[0162] Fig.18 It is shown Fig.17 133 in the three-dimensional position information smoothing processing unit 133. Fig.18As shown, the three-dimensional position information smoothing processing unit 133 includes a region dividing unit 141, a representative value deriving unit 142 within the region, a processing target region setting unit 143, a smoothing processing unit 144 and a sending information generating unit 145.
[0163] The region division unit 141 obtains the position information about the point cloud (geometry point cloud) provided from the geometry decoding unit 132. The region division unit 141 divides the region of the three-dimensional space including the obtained geometry point cloud, and sets the local region (grid). At this time, the region division unit 141 divides the three-dimensional space and sets the local region by the method described above in <#1. Acceleration using a representative value for each local region>.
[0164] The region division unit 141 provides information about the set local region (for example, information about the shape and size of the local region) and information about the geometric point cloud to the intra-region representative value derivation unit 142. In addition, when the information about the local region is transmitted to the decoding side, the region division unit 141 provides the information about the local region to the transmission information generation unit 145.
[0165] The intra-region representative value deriving unit 142 acquires information on the local region and the geometric point cloud provided from the region dividing unit 141. Based on this information, the intra-region representative value deriving unit 142 derives a representative value of the geometric point cloud in each local region set by the region dividing unit 141. At this time, the intra-region representative value deriving unit 142 derives the representative value by the method described above in <#1. Acceleration using the representative value for each local region>.
[0166] The intra-region representative value deriving unit 142 provides information about the local region, the geometric point cloud, and the representative value derived for each local region to the smoothing processing unit 144. In addition, when the representative value derived for each local region is to be sent to the decoding side, information indicating the representative value of each local region is provided to the transmission information generating unit 145.
[0167] The processing target region setting unit 143 acquires an occupancy map. The processing target region setting unit 143 sets a region to which the filtering process is to be applied based on the acquired occupancy map. At this time, the processing target region setting unit 143 sets the region by the method described above in <#2. Simplification of three-dimensional filtering process>. In other words, the processing target region setting unit 143 sets a partial region corresponding to the end of a patch in the occupancy map as a processing target region for filtering process.
[0168] The processing target region setting unit 143 supplies information indicating the set processing target region to the smoothing processing unit 144. Furthermore, when information indicating the processing target region is to be transmitted to the decoding side, the processing target region setting unit 143 supplies the transmission information generating unit 145 with the information indicating the processing target region.
[0169] The smoothing processing unit 144 acquires information on the local area, the geometric point cloud, and the representative value for each local area supplied from the in-area representative value derivation unit 142. Furthermore, the smoothing processing unit 144 acquires information indicating the processing target area that has been supplied from the processing target area setting unit 143.
[0170] The smoothing processing unit 144 performs three-dimensional smoothing filter processing based on this information. In other words, as described above in <Accelerated Three-Dimensional Filter Processing>, the smoothing processing unit 144 uses the representative value of each local area as a reference value to perform three-dimensional smoothing filter processing on the points of the geometric point cloud in the processing target area. Therefore, the smoothing processing unit 144 can perform three-dimensional smoothing filter processing at a higher speed.
[0171] The smoothing processing unit 144 supplies the geometric point cloud subjected to the three-dimensional smoothing filter process (smoothed geometric point cloud) to the patch decomposition processing unit 131 and the texture correction unit 134 .
[0172] The transmission information generation unit 145 acquires the information about the local area provided from the area division unit 141, the information indicating the representative value for each local area provided from the intra-area representative value deriving unit 142, and the information indicating the processing target area provided from the processing target area setting unit 143. The transmission information generation unit 145 generates transmission information including these information. The transmission information generation unit 145 provides the generated transmission information to, for example, the auxiliary patch information compression unit 114, and causes the auxiliary patch information compression unit 114 to transmit the provided transmission information to the decoding side as auxiliary patch information.
[0173] <Encoding Process Flow>
[0174] Will refer to Fig.19 An example of the flow of the encoding process performed by the encoding device 100 is described with reference to the flowchart in FIG.
[0175] Once the encoding process starts, in step S101 , the patch decomposition unit 111 of the encoding device 100 projects a point cloud onto a two-dimensional plane and decomposes the projected point cloud into patches.
[0176] In step S102 , the auxiliary patch information compressing unit 114 compresses the auxiliary patch information generated in step S101 .
[0177] In step S103, the packing unit 112 packs each patch of the position information and attribute information generated in step S101 into a video frame. In addition, the OMap generation unit 113 generates an occupancy map corresponding to the video frame of the position information and attribute information.
[0178] In step S104 , the video encoding unit 115 encodes the geometry video frame, which is the video frame of the position information generated in step D103 , by an encoding method for a two-dimensional image.
[0179] In step S105 , the video encoding unit 116 encodes the color video frame, which is the video frame of the attribute information generated in step S103 , by the encoding method for a two-dimensional image.
[0180] In step S106 , the OMap encoding unit 117 encodes the occupancy map generated in step S103 by a predetermined encoding method.
[0181] In step S107 , the multiplexer 118 multiplexes the various information generated as described above, and generates a bit stream including the information.
[0182] In step S108 , the multiplexer 118 outputs the bit stream generated in step S107 to the outside of the encoding device 100 .
[0183] Once the processing in step S108 ends, the encoding process ends.
[0184] <Patch decomposition process>
[0185] Next, we will refer to Fig. 20 The flowchart in Fig.19 An example of the flow of the patch decomposition processing performed in step S101 in FIG.
[0186] Once the patch decomposition process is started, the patch decomposition processing unit 131 decomposes the point cloud into patches in step S121 , and generates a geometry patch and a texture patch.
[0187] In step S122 , the geometry decoding unit 132 decodes geometry encoding data obtained by packing the geometry patches generated in step S121 into video frames and encoding the video frames, and reconstructs a point cloud to generate a point cloud of geometry.
[0188] In step S123 , the three-dimensional position information smoothing processing unit 133 performs smoothing processing, and performs three-dimensional smoothing filter processing on the geometric point cloud generated in step S122 .
[0189] In step S124 , the texture correction unit 134 corrects the texture patch generated in step S121 using the smoothed geometric point cloud obtained by the process in step S123 .
[0190] In step S125 , the patch decomposition processing unit 131 decomposes the smoothed geometric point cloud obtained by the process in step S123 into patches, and generates smoothed geometric patches.
[0191] Once the processing in step S136 is completed, the patch decomposition processing is completed, and the processing returns to Fig.19 .
[0192] <Flow of Smoothing Process>
[0193] Next, we will refer to Fig.21 The flowchart in Fig. 20 An example of the flow of the smoothing process performed in step 123 of FIG.
[0194] Once the smoothing process starts, in step S141, the region division unit 141 divides the three-dimensional space including the point cloud into local regions. The region division unit 141 divides the three-dimensional space and sets local regions by the method described above in <#1. Acceleration using a representative value for each local region>.
[0195] In step S142, the in-region representative value deriving unit 142 derives a representative value of the point cloud for each local region set in step S141. The in-region representative value deriving unit 142 derives the representative value by the method described above in <#1. Acceleration using representative value for each local region>.
[0196] In step S143, the processing target region setting unit 143 sets a range for performing smoothing processing. The processing target region setting unit 143 sets the region by the method described above in <#2. Simplification of three-dimensional filtering processing>. In other words, the processing target region setting unit 143 sets a partial region corresponding to the end of the patch in the occupancy map as a processing target region for filtering processing.
[0197] In step S144, the smoothing processing unit 144 performs smoothing processing on the processing target range set in step S143 by referring to the representative value of each area. As described above in <Accelerated Three-Dimensional Filter Processing>, the smoothing processing unit 144 uses the representative value of each local area as a reference value to perform three-dimensional smoothing filtering processing on the points of the geometric point cloud in the processing target area. Therefore, the smoothing processing unit 144 can perform three-dimensional smoothing filtering processing at a higher speed.
[0198] In step S145 , the transmission information generation unit 145 generates transmission information about smoothing to supply the generated transmission information to, for example, the auxiliary patch information compression unit 114 , and causes the auxiliary patch information compression unit 114 to transmit the supplied transmission information as auxiliary patch information.
[0199] Once the processing in step S145 ends, the smoothing process ends, and the process returns to Fig. 20 .
[0200] <Flow of Smoothing Range Setting Process>
[0201] Next, we will refer to Fig. 22 The flowchart in the description is in Fig.21 An example of the flow of the smoothing range setting process performed in step S143 of FIG.
[0202] Once the smoothing range setting process is started, the processing target area setting unit 143 determines in step S161 whether the current position (x, y) in the occupancy map (processing target block) is located at the end of the occupancy map. For example, when the horizontal width of the occupancy map is assumed to be the width and the vertical width is assumed to be the height, the following determination is made.
[0203] x!=0&y!=0&x!=width-1&y!=height-1
[0204] When it is determined that the determination is true, that is, the current position is not located at the end of the occupancy map, the process proceeds to step S162.
[0205] In step S162, the processing target area setting unit 143 determines whether all values of the peripheral parts of the current position in the occupancy map have 1. When it is determined that all values of the peripheral parts of the current position in the occupancy map have 1, that is, all peripheral parts have position information and attribute information and are not located near the boundary between a part having position information and attribute information and a part having no position information or attribute information, the process proceeds to step S163.
[0206] In step S163, the processing target area setting unit 143 determines whether all patches to which the peripheral part of the current position belongs are consistent with the patch to which the current position belongs. When the patches are placed side by side, the parts where the value of the occupancy map has 1 are continuous. Therefore, even in the case where it is determined in step S162 that all peripheral parts of the current position have data, parts where multiple patches are adjacent to each other may be cases where there is data, and the current position may still be located at the end of the patch. Then, since the image is basically discontinuous between different patches, even in parts where multiple patches are adjacent to each other, a large size such as that shown in FIG. 1 may be formed due to the large size of the accuracy of the occupancy map. Figure 1Therefore, as described above, it is determined whether all patches to which the peripheral portion of the current position belongs are consistent with the patch to which the current position belongs.
[0207] When it is determined that all peripheral portions and the current position belong to the same patch as each other, that is, the current position is not located in a portion where a plurality of patches are adjacent to each other and is not located at an end of a patch, the process proceeds to step S164.
[0208] In step S164, the processing target area setting unit 143 determines the three-dimensional point (the point of the point cloud corresponding to the processing target block) restored from the current position (x, y) as a point that will not be subjected to the smoothing filter processing. In other words, the current position is excluded from the smoothing processing target range. Once the processing in step S164 is completed, the processing proceeds to step S166.
[0209] Furthermore, when it is determined in step S161 that the above determination is false, that is, the current position is located at the end of the occupancy map, the process proceeds to step S165.
[0210] Furthermore, when it is determined in step S162 that there is a peripheral portion where the value of the occupancy map does not have 1, that is, there is a peripheral portion having no position information or attribute information, and the current position is at the end of the patch, the process proceeds to step S165.
[0211] Furthermore, when it is determined in step S163 that there is a peripheral portion belonging to a patch different from the patch to which the current position belongs, that is, the current position is located in a portion where a plurality of patches are adjacent to each other, the process proceeds to step S165 .
[0212] In step S165, the processing target area setting unit 143 determines the three-dimensional point (the point of the point cloud corresponding to the processing target block) restored from the current position (x, y) as the point to be subjected to the smoothing filter processing. In other words, the current position is set as the smoothing processing target range. Once the processing in step S165 is completed, the processing proceeds to step S166.
[0213] In step S166, the processing target area setting unit 143 determines whether all positions (blocks) in the occupied map have been processed. When it is determined that there are unprocessed positions (blocks), the process returns to step S161, and the subsequent process is repeated for the unprocessed blocks assigned as processing target blocks. In other words, the processes in steps S161 to S166 are repeated for each block.
[0214] Then, when it is determined in step S166 that all positions (blocks) in the occupation map have been processed, the smoothing range setting process ends, and the process returns to Fig.21 .
[0215] By performing each process as described above, an increase in the processing time of the filter process on the point cloud data can be suppressed (the filter process can be performed at a higher speed).
[0216] <3. Second Embodiment>
[0217] <Decoding device>
[0218] Next, a configuration for realizing each scheme as mentioned above will be described. Fig.23 : is a block diagram showing an example of the configuration of a decoding device as an exemplary form of an image processing device to which the present technology is applied. Fig.23 The decoding device 200 shown in FIG. 1 is a device that decodes coded data obtained by projecting 3D data such as a point cloud onto a two-dimensional plane and encoding the projected 3D data by a decoding method for a two-dimensional image, and projects the decoded data into a three-dimensional space (a decoding device to which a video-based method is applied). For example, the decoding device 200 decodes coded data obtained by projecting 3D data such as a point cloud onto a two-dimensional plane and encoding the projected 3D data by a decoding method for a two-dimensional image, and projects the decoded data into a three-dimensional space (a decoding device to which a video-based method is applied). Fig.16 ) to decode the bit stream generated and reconstruct the point cloud.
[0219] Notice, Fig.23 shows the main parts of the processing units, data flow, etc., and Fig.23 Not all of them are necessarily shown. In other words, in the decoding device 200, there may be Fig.23 The processing units are not shown as blocks, or may exist in Fig.23 200. Processing or data flows as arrows, etc. are not shown in FIG. Similarly, this also applies to other figures illustrating processing units, etc. in the decoding device 200.
[0220] like Fig.23 As shown, the decoding apparatus 200 includes a demultiplexer 211 , an auxiliary patch information decoding unit 212 , a video decoding unit 213 , a video decoding unit 214 , an OMap decoding unit 215 , a depacketizing unit 216 and a 3D reconstruction unit 217 .
[0221] The demultiplexer 211 performs processing related to data demultiplexing. For example, the demultiplexer 211 acquires a bit stream input to the decoding device 200. The bit stream is provided from, for example, the encoding device 100. The demultiplexer 211 demultiplexes the bit stream, and extracts the encoded data of the auxiliary patch information to provide the extracted encoded data to the auxiliary patch information decoding unit 212. In addition, the demultiplexer 211 extracts the encoded data of the video frame of the position information (geometry) from the bit stream by performing demultiplexing, and provides the extracted encoded data to the video decoding unit 213. In addition, the demultiplexer 211 extracts the encoded data of the video frame of the attribute information (texture) from the bit stream by performing demultiplexing, and provides the extracted encoded data to the video decoding unit 214. In addition, the demultiplexer 211 extracts the encoded data of the occupancy map from the bit stream by performing demultiplexing, and provides the extracted encoded data to the OMap decoding unit 215. Furthermore, the demultiplexer 211 extracts control information about packetization from the bit stream by performing demultiplexing, and supplies the extracted control information to the depacketizing unit 216 .
[0222] The auxiliary patch information decoding unit 212 performs processing related to decoding of the encoded data of the auxiliary patch information. For example, the auxiliary patch information decoding unit 212 acquires the encoded data of the auxiliary patch information provided from the demultiplexer 211. In addition, the auxiliary patch information decoding unit 212 decodes (decompresses) the encoded data of the auxiliary patch information included in the acquired data. The auxiliary patch information decoding unit 212 provides the auxiliary patch information obtained by decoding to the 3D reconstruction unit 217.
[0223] The video decoding unit 213 performs processing related to decoding of the encoded data of the video frame of the position information (geometry). For example, the video decoding unit 213 obtains the encoded data of the video frame of the position information (geometry) provided from the demultiplexer 211. In addition, the video decoding unit 213 decodes the acquired encoded data by any decoding method (for example, AVC or HEVC) for two-dimensional images, for example, to obtain the video frame of the position information (geometry). The video decoding unit 213 provides the obtained video frame of the position information (geometry) to the depacketizing unit 216.
[0224] The video decoding unit 214 performs processing related to decoding of the encoded data of the video frame of the attribute information (texture). For example, the video decoding unit 214 acquires the encoded data of the video frame of the attribute information (texture) provided from the demultiplexer 211. In addition, the video decoding unit 214 decodes the acquired encoded data by any decoding method (for example, AVC or HEVC) for two-dimensional images, for example, to obtain the video frame of the attribute information (texture). The video decoding unit 214 supplies the acquired video frame of the attribute information (texture) to the depacketizing unit 216.
[0225] The OMap decoding unit 215 performs processing related to decoding of the encoded data of the occupancy map. For example, the OMap decoding unit 215 acquires the encoded data of the occupancy map provided from the demultiplexer 211. In addition, for example, the OMap decoding unit 215 decodes the acquired encoded data by any decoding method such as arithmetic decoding corresponding to arithmetic coding to obtain an occupancy map. The OMap decoding unit 215 supplies the acquired occupancy map to the depacketizing unit 216.
[0226] The unpacking unit 216 performs processing related to unpacking. For example, the unpacking unit 216 obtains the video frame of the position information (geometry) from the video decoding unit 213, obtains the video frame of the attribute information (texture) from the video decoding unit 214, and obtains the occupancy map from the OMap decoding unit 215. In addition, the unpacking unit 216 unpacks the video frame of the position information (geometry) and the video frame of the attribute information (texture) based on the control information about the packing. The unpacking unit 216 provides the data of the position information (geometry) (e.g., geometry patch), the data of the attribute information (texture) (e.g., texture patch), the occupancy map, etc. obtained by unpacking to the 3D reconstruction unit 217.
[0227] The 3D reconstruction unit 217 performs processing related to reconstruction of the point cloud. For example, the 3D reconstruction unit 217 reconstructs the point cloud based on the auxiliary patch information provided from the auxiliary patch information decoding unit 212, the data of the position information (geometry) provided from the unpacking unit 216 (e.g., geometry patch), the data of the attribute information (texture) (e.g., texture patch), the occupancy map, etc. The 3D reconstruction unit 217 outputs the reconstructed point cloud to the outside of the decoding device 200.
[0228] The point cloud is provided to a display unit and imaged, for example, and the image is displayed, recorded on a recording medium, or provided to another device via communication.
[0229] In such a decoding device 200 , the 3D reconstruction unit 217 performs a three-dimensional smoothing filter process on the reconstructed point cloud.
[0230] <3D reconstruction unit>
[0231] Fig.24 It is shown Fig.23 A block diagram of an example of the main configuration of the decoding unit 217 in FIG. Fig.24 As shown, the 3D reconstruction unit 217 includes a geometric point cloud (PointCloud) generation unit 231, a three-dimensional position information smoothing processing unit 232 and a texture synthesis unit 233.
[0232] The geometric point cloud generation unit 231 performs processing related to the generation of the geometric point cloud. For example, the geometric point cloud generation unit 231 acquires the geometric patch provided from the unpacking unit 216. In addition, the geometric point cloud generation unit 231 reconstructs the geometric point cloud (position information about the point cloud) using the acquired geometric patch and other information such as auxiliary patch information. The geometric point cloud generation unit 231 provides the generated geometric point cloud to the three-dimensional position information smoothing processing unit 232.
[0233] The three-dimensional position information smoothing processing unit 232 performs processing related to the three-dimensional smoothing filter processing. For example, the three-dimensional position information smoothing processing unit 232 acquires the geometric point cloud provided from the geometric point cloud generating unit 231. In addition, the three-dimensional position information smoothing processing unit 232 acquires the occupancy map provided from the unpacking unit 216.
[0234] The three-dimensional position information smoothing processing unit 232 performs three-dimensional smoothing filter processing on the acquired geometric point cloud. At this time, as described above, the three-dimensional position information smoothing processing unit 232 performs three-dimensional smoothing filter processing using a representative value for each local area obtained by dividing the three-dimensional space. In addition, the three-dimensional position information smoothing processing unit 232 uses the acquired occupancy map to perform three-dimensional smoothing filter processing only on points in a partial area corresponding to the end of a patch in the acquired occupancy map. By performing the three-dimensional smoothing filter processing in this way, the three-dimensional position information smoothing processing unit 232 can perform filtering processing at a higher speed.
[0235] The three-dimensional position information smoothing processing unit 232 supplies the geometric point cloud subjected to the filter processing (smoothed geometric point cloud) to the texture synthesizing unit 233 .
[0236] The texture synthesis unit 233 performs processing related to geometry and texture synthesis. For example, the texture synthesis unit 233 obtains the smoothed geometry point cloud provided from the three-dimensional position information smoothing processing unit 232. In addition, the texture synthesis unit 233 obtains the texture patch provided from the unpacking unit 216. The texture synthesis unit 233 synthesizes the texture patch (i.e., attribute information) into the smoothed geometry point cloud, and reconstructs the point cloud. The position information of the smoothed geometry point cloud changes due to three-dimensional smoothing. In other words, strictly speaking, there may be a part where the position information and the attribute information do not correspond to each other. Therefore, the texture synthesis unit 233 synthesizes the attribute information obtained from the texture patch into the smoothed geometry point cloud, while reflecting the change in the position information on the part subjected to three-dimensional smoothing.
[0237] The texture synthesis unit 233 outputs the reconstructed point cloud to the outside of the decoding device 200 .
[0238] <Three-dimensional position information smoothing processing unit>
[0239] Fig.25 It is shown Fig.24 2 is a block diagram showing a main configuration example of the three-dimensional position information smoothing processing unit 232. Fig.25 As shown, the three-dimensional position information smoothing processing unit 232 includes a sending information acquisition unit 251, an area division unit 252, an area representative value deriving unit 253, a processing target area setting unit 254 and a smoothing processing unit 255.
[0240] When there is transmission information transmitted from the encoding side, the transmission information acquisition unit 251 acquires the transmission information provided as auxiliary patch information or the like. The transmission information acquisition unit 251 supplies the acquired transmission information to the region division unit 252, the intra-region representative value derivation unit 253, and the processing target region setting unit 254 as necessary. For example, when information about a local region is provided as the transmission information, the transmission information acquisition unit 251 supplies the provided information about the local region to the region division unit 252. Furthermore, when information indicating a representative value for each local region is provided as the transmission information, the transmission information acquisition unit 251 supplies the provided information indicating a representative value for each local region to the intra-region representative value derivation unit 253. Furthermore, when information indicating a processing target region is provided as the transmission information, the transmission information acquisition unit 251 supplies the provided information indicating the processing target region to the processing target region setting unit 254.
[0241] The region division unit 252 acquires the position information about the point cloud (geometric point cloud) provided from the geometric point cloud generation unit 231. The region division unit 252 divides the region of the three-dimensional space including the acquired geometric point cloud, and sets the local area (grid). At this time, the region division unit 252 divides the three-dimensional space and sets the local area by the method described above in <#1. Acceleration using a representative value for each local area>. Note that when information about the local area sent from the encoding side is provided from the transmission information acquisition unit 251, the region division unit 252 adopts the setting of the local area indicated by the provided information (for example, the shape and size of the local area).
[0242] The region dividing unit 252 supplies information about the set local region (for example, information about the shape and size of the local region) and information about the geometric point cloud to the in-region representative value deriving unit 253 .
[0243] The intra-region representative value deriving unit 253 acquires information about the local region and the geometric point cloud provided from the region dividing unit 252. The intra-region representative value deriving unit 253 derives the representative value of the geometric point cloud in each local region set by the region dividing unit 252 based on this information. At this time, the intra-region representative value deriving unit 253 derives the representative value by the method described above in <#1. Acceleration using the representative value for each local region>. Note that when information indicating the representative value for each local region that has been transmitted from the encoding side is provided from the transmission information acquiring unit 251, the intra-region representative value deriving unit 253 adopts the representative value for each local region indicated by the provided information.
[0244] The intra-region representative value deriving unit 253 supplies information on the local region, the geometric point cloud, and the representative value derived for each local region to the smoothing processing unit 255 .
[0245] The processing target area setting unit 254 acquires the occupancy map. The processing target area setting unit 254 sets the area to which the filtering process is to be applied based on the acquired occupancy map. At this time, the processing target area setting unit 254 sets the area by the method described above in <#2. Simplification of three-dimensional filtering process>. In other words, the processing target area setting unit 254 sets the partial area corresponding to the end of the patch in the occupancy map as the processing target area for the filtering process. Note that when the information indicating the processing target area that has been transmitted from the encoding side is supplied from the transmission information acquisition unit 251, the processing target area setting unit 254 adopts the processing target area indicated by the supplied information.
[0246] The processing target region setting unit 254 supplies information indicating the set processing target region to the smoothing processing unit 255 .
[0247] The smoothing processing unit 255 acquires information on the local area, the geometric point cloud, and the representative value for each local area supplied from the in-area representative value derivation unit 253. Furthermore, the smoothing processing unit 255 acquires information indicating the processing target area that has been supplied from the processing target area setting unit 254.
[0248] The smoothing processing unit 255 performs three-dimensional smoothing filter processing based on this information. In other words, as described above in <Accelerated Three-Dimensional Filter Processing>, the smoothing processing unit 255 uses the representative value of each local area as a reference value to perform three-dimensional smoothing filter processing on the points of the geometric point cloud in the processing target area. Therefore, the smoothing processing unit 255 can perform three-dimensional smoothing filter processing at a higher speed.
[0249] The smoothing processing unit 255 supplies the geometric point cloud subjected to the three-dimensional smoothing filter process (smoothed geometric point cloud) to the texture synthesis unit 233 .
[0250] <Decoding Process Flow>
[0251] Next, we will refer to Fig.26 The flowchart describes an example of the flow of the decoding process performed by the decoding device 200.
[0252] Once the decoding process starts, the demultiplexer 211 of the decoding device 200 demultiplexes the bit stream in step S201.
[0253] In step S202 , the auxiliary patch information decoding unit 212 decodes the auxiliary patch information extracted from the bitstream in step S201 .
[0254] In step S203 , the video decoding unit 213 decodes the encoded data of the geometry video frame (the video frame of the position information) extracted from the bit stream in step S201 .
[0255] In step S204 , the video decoding unit 214 decodes the encoded data of the color video frame (the video frame of the attribute information) extracted from the bit stream in step S201 .
[0256] In step S205 , the OMap decoding unit 215 decodes the encoded data of the occupancy map extracted from the bit stream in step S201 .
[0257] In step S206, the unpacking unit 216 unpacks the geometry video frame obtained by decoding the encoded data in step S203 to generate a geometry patch. In addition, the unpacking unit 216 unpacks the color video frame obtained by decoding the encoded data in step S204 to generate a texture patch. In addition, the unpacking unit 216 unpacks the occupancy map obtained by decoding the encoded data in step S205 to extract the occupancy map corresponding to the geometry patch and the texture patch.
[0258] In step S207 , the 3D reconstruction unit 217 reconstructs a point cloud based on the auxiliary patch information obtained in step S202 and the geometry patch, texture patch, occupancy map, etc. obtained in step S206 .
[0259] Once the processing in step S207 ends, the decoding process ends.
[0260] <Point cloud reconstruction process>
[0261] Next, we will refer to Fig. 27 The flowchart is described in Fig.26An example of the flow of the point cloud reconstruction processing performed in step S207.
[0262] Once the point cloud reconstruction process starts, in step S221 , the geometric point cloud generation unit 231 of the 3D reconstruction unit 217 reconstructs a geometric point cloud.
[0263] In step S222 , the three-dimensional position information smoothing processing unit 232 performs smoothing processing, and performs three-dimensional smoothing filter processing on the geometric point cloud generated in step S221 .
[0264] In step S223 , the texture synthesis unit 233 synthesizes the texture patch into the smoothed geometric point cloud.
[0265] Once the processing in step S223 is completed, the point cloud reconstruction processing is completed, and the processing returns to Fig.26 .
[0266] <Flow of Smoothing Process>
[0267] Next, we will refer to Fig.28 The flowchart is used to describe the Fig. 27 An example of the flow of the smoothing process performed in step S222 of FIG.
[0268] Once the smoothing process is started, the transmission information acquisition unit 251 acquires transmission information about the smoothing in step S241. Note that when there is no transmission information, this process is omitted.
[0269] In step S242, the region division unit 252 divides the three-dimensional space including the point cloud into local regions. The region division unit 252 divides the three-dimensional space and sets the local regions by the method described above in <#1. Acceleration using a representative value for each local region>. Note that when information about the local region is acquired as the transmission information in step S241, the region division unit 252 adopts the setting of the local region (shape, size, etc. of the local region) indicated by the acquired information.
[0270] In step S243, the in-region representative value deriving unit 253 derives a representative value of the point cloud for each local region set in step S242. The in-region representative value deriving unit 253 derives the representative value by the method described above in <#1. Acceleration using the representative value for each local region>. Note that when information indicating the representative value for each local region is acquired as the transmission information in step S241, the in-region representative value deriving unit 253 adopts the representative value for each local region indicated by the acquired information.
[0271] In step S244, the processing target region setting unit 254 sets a range for performing smoothing processing. The processing target region setting unit 254 sets the region by the method described above in <#2. Simplification of three-dimensional filtering processing>. In other words, the processing target region setting unit 254 performs reference Fig. 22 The smoothing range setting process described in the flowchart in is performed, and the processing target range for the filtering process is set. Note that when information indicating the processing target area is acquired as the transmission information in step S241, the processing target area setting unit 254 adopts the setting of the processing target area indicated by the acquired information.
[0272] In step S245, the smoothing processing unit 255 performs smoothing processing on the processing target range set in step S244 by referring to the representative value of each area. As described above in <Accelerated Three-Dimensional Filter Processing>, the smoothing processing unit 255 uses the representative value of each local area as a reference value to perform three-dimensional smoothing filtering processing on the points of the geometric point cloud in the processing target area. Therefore, the smoothing processing unit 255 can perform three-dimensional smoothing filtering processing at a higher speed.
[0273] Once the processing in step S245 ends, the smoothing process ends, and the process returns to Fig. 27 .
[0274] By performing each process as described above, an increase in the processing time of the filter process on the point cloud data can be suppressed (the filter process can be performed at a higher speed).
[0275] <4. Variation>
[0276] In the first and second embodiments, the three-dimensional smoothing filter processing has been described for the position information related to the point cloud, but the three-dimensional smoothing filter processing can also be performed on the attribute information related to the point cloud. In this case, for example, since the attribute information has been smoothed, the color of the point, etc., also changes.
[0277] For example, in the case of the encoding device 100, only the patch decomposition unit 111 ( Fig.17 ) provides a smoothing processing unit (e.g., a three-dimensional attribute information smoothing processing unit) that performs smoothing processing on the texture patch provided to the texture correction unit 134.
[0278] In addition, for example, in the case of the decoding device 200, it is only necessary to perform the 3D reconstruction in the 3D reconstruction unit 217 ( Fig.24 ) provides a smoothing processing unit (e.g., a three-dimensional attribute information smoothing processing unit) that performs smoothing processing on the texture patch provided to the texture synthesis unit 233.
[0279] <5. Additional Notes>
[0280] <Control Information>
[0281] Control information related to the present technology described in each of the above embodiments may be transmitted from the encoding side to the decoding side. For example, control information (e.g., enabled_flag) that controls whether to allow (or prohibit) the application of the above technology may be transmitted. In addition, for example, control information that specifies a range (e.g., an upper limit or a lower limit of a block size, or both an upper limit and a lower limit, a slice, a picture, a sequence, a component, a view, a layer, etc.) that allows (or prohibits) the application of the above technology may be transmitted.
[0282] <Computer>
[0283] The above series of processes can also be performed by using hardware and can also be performed by using software. When a series of processes are performed by software, the program constituting the software is installed in a computer. Here, the computer includes a computer built into dedicated hardware, a computer capable of performing various functions when various programs are installed, such as a general-purpose personal computer, etc.
[0284] Fig.29 : is a block diagram showing a hardware configuration example of a computer that executes the above-described series of processes using a program.
[0285] exist Fig.29 In a computer 900 shown in , a central processing unit (CPU) 901 , a read only memory (ROM) 902 , and a random access memory (RAM) 903 are interconnected via a bus 904 .
[0286] Furthermore, an input / output interface 905 is also connected to the bus 904. An input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected to the input / output interface 905.
[0287] For example, the input unit 911 includes a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. For example, the output unit 912 includes a display, a speaker, an output terminal, etc. For example, the storage unit 913 includes a hard disk, a RAM disk, a non-volatile memory, etc. For example, the communication unit 914 includes a communication interface. The drive 915 drives a removable medium 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0288] In the computer configured as described above, for example, the above-described series of processing is performed so that the CPU 901 loads the program stored in the storage unit 913 into the RAM 903 for execution via the input / output interface 905 and the bus 904. Data and the like required by the CPU 901 when executing various processing are also appropriately stored in the RAM 903.
[0289] For example, the program executed by the computer (CPU 901) can be applied by recording it in the removable medium 921 serving as a package medium, etc. In that case, by mounting the removable medium 921 in the drive 915, the program can be installed in the storage unit 913 via the input / output interface 905.
[0290] In addition, the program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 914 to be installed in the storage unit 913.
[0291] Alternatively, the program may also be installed in the ROM 902 or the storage unit 913 in advance.
[0292] <Application target of this technology>
[0293] In the above, the case where the present technology is applied to the encoding and decoding of point cloud data has been described, but the present technology is not limited to these examples and can be applied to the encoding and decoding of 3D data of any standard. In other words, as long as there is no contradiction with the above-mentioned present technology, various processes such as encoding and decoding technology and specifications of various data such as 3D data and metadata are optional. In addition, as long as there is no contradiction with the present technology, some of the above-mentioned processes and specifications can be omitted.
[0294] The present technology can be applied to any configuration. For example, the present technology can be applied to a variety of electronic devices, such as transmitters and receivers for satellite broadcasting, cable broadcasting such as cable television, distribution on the Internet, distribution to terminals via cellular communications, etc. (e.g., television receivers and mobile phones), or devices that record images on media such as optical disks, magnetic disks, and flash memories and reproduce images from these storage media (e.g., hard disk recorders and camera devices).
[0295] In addition, for example, the present technology can also be implemented as a partial configuration of an apparatus such as a processor used as a system large-scale integration (LSI) (e.g., a video processor), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a collection of units that further adds another function (e.g., a video collection).
[0296] In addition, for example, the present technology can also be applied to a network system composed of multiple devices. For example, the present technology can be implemented as cloud computing, in which processing is shared and collaboratively performed by multiple devices via a network. For example, the present technology can be implemented in a cloud service that provides services related to images (moving images) to any terminal such as a computer, an audio-visual (AV) device, a portable information processing terminal, and an Internet of Things (IoT) device.
[0297] Note that in this specification, a system refers to a collection of multiple constituent components (e.g., devices and modules (components)), and it is not considered important whether all constituent components are arranged in the same cabinet. Therefore, multiple devices housed in separate cabinets so as to be connected to each other via a network and a device in which multiple modules are housed in one cabinet are both systems.
[0298] <Fields where this technology can be applied and purposes of use>
[0299] For example, the device, processing unit, etc. to which the present technology is applied can be used in any field such as transportation, medical care, crime prevention, agriculture, animal husbandry, mining, beauty, factories, household appliances, meteorology, and natural monitoring. In addition, the purpose of use of the above system, etc. is also optional.
[0300] <Others>
[0301] Note that in this specification, "flag" refers to information used to identify between multiple states, and includes not only information used when identifying two states of true (1) and false (0), but also information that can identify three or more states. Therefore, the value that the "flag" can take can be a binary value such as 1 or 0, or a ternary or more value. That is, the number of bits constituting the "flag" is arbitrary, and 1 bit or more bits can be used. In addition, it is assumed that the identification information (including the flag) not only has a form in which the identification information is included in the bit stream, but also has a form in which the difference information of the identification information relative to specific reference information is included in the bit stream. Therefore, in this specification, "flag" and "identification information" not only imply all the information therein, but also imply the difference information relative to the reference information.
[0302] In addition, various information (metadata, etc.) about the coded data (bitstream) can be sent or recorded in any form, as long as the information is associated with the coded data. Here, the term "association" means, for example, ensuring that one data is available (linkable) when processing another data. In other words, data associated with each other can be collected as one data, or data associated with each other can be processed as separate data. For example, information associated with the coded data (image) can be sent on a transmission path different from the transmission path of the associated coded data (image). In addition, for example, information associated with the coded data (image) can be recorded on a recording medium (or a recording area of the same recording medium) different from the recording medium of the associated coded data (image). Note that the "association" can be performed on a part of the data rather than the entire data. For example, an image and information corresponding to the image can be associated with each other in arbitrary units such as multiple frames, one frame, or a part of a frame.
[0303] In addition, in this specification, terms such as "synthesize", "multiplex", "add", "integrate", "include", "save", "merge", "put in", and "insert" mean collecting multiple objects into one object, such as collecting encoded data and metadata into one data, and mean one of the above-mentioned "association" methods.
[0304] Furthermore, the embodiment according to the present technology is not limited to the above-described embodiment, and various modifications may be made without departing from the scope of the present technology.
[0305] For example, a configuration described as one device (or processing unit) may be divided to be configured as a plurality of devices (or processing units). In contrast, a configuration described as a plurality of devices (or processing units) above may be collected to be configured as one device (or one processing unit). Furthermore, of course, configurations other than those described above may be added to the configurations of the respective devices (or respective processing units). Furthermore, a portion of the configuration of a particular device (or a particular processing unit) may be included in the configuration of another device (or another processing unit), as long as the configuration and operation of the system remain substantially unchanged overall.
[0306] In addition, for example, the above-mentioned program can be executed by any device. In that case, those devices only need to have necessary functions (functional blocks, etc.) so that those necessary information can be obtained.
[0307] In addition, for example, one device may perform each step of a flowchart, or multiple devices may share and perform the steps. In addition, when multiple processes are included in one step, the multiple processes may be performed by a single device, or may be shared and performed by multiple devices. In different aspects, multiple processes included in one step may also be performed as a process having multiple steps. In contrast, a process described as multiple steps may also be collected as one step and performed.
[0308] In addition, for example, the program executed by the computer may be designed so that the processing of the steps describing the program is performed along the time series according to the order described in this specification, or is performed in parallel or individually at the necessary timing (for example, when called). In other words, as long as there is no contradiction, the processing of the various steps may be performed in an order different from the above order. In addition, these processings describing the steps of the program may be performed in parallel with the processing of another program, or may be performed in combination with the processing of another program.
[0309] In addition, for example, each of the multiple technologies related to the present technology can be independently performed separately as long as there is no inconsistency. Of course, any multiple technologies can also be implemented simultaneously. For example, part or all of the present technology described in any embodiment can be combined with part or all of the present technology described in another embodiment. In addition, another technology not mentioned above can be used to perform part or all of any one of the above technologies at the same time.
[0310] In addition, the present technology also includes the following embodiments.
[0311] Embodiment 1: An image processing device, comprising:
[0312] a filter processing unit that performs filter processing on the point cloud data using a representative value of the point cloud data of each local area obtained by dividing the three-dimensional space; and
[0313] An encoding unit that encodes a two-dimensional plane image on which the point cloud data subjected to the filter processing by the filter processing unit is projected and generates a bit stream.
[0314] Embodiment 2: The image processing device according to embodiment 1, wherein the local area includes a cubic area with a predetermined size.
[0315] Embodiment 3: The image processing device according to embodiment 1, wherein the local area includes a rectangular area with a predetermined size.
[0316] Embodiment 4: The image processing device according to embodiment 1, wherein the local area includes an area obtained by dividing the three-dimensional space so that each area contains a predetermined number of points in the point cloud data.
[0317] Embodiment 5: The image processing device according to embodiment 1, wherein the encoding unit generates the bit stream including information about the local area.
[0318] Embodiment 6: The image processing device according to Embodiment 5, wherein the information about the local area includes information about the size of the local area, or information about the shape of the local area, or information about the size and shape of the local area.
[0319] Embodiment 7: The image processing device according to embodiment 1, wherein the representative value includes an average value of the point cloud data contained in the local area.
[0320] Embodiment 8: The image processing device according to embodiment 1, wherein the representative value includes a median of the point cloud data contained in the local area.
[0321] Embodiment 9: The image processing device according to embodiment 1, wherein the filtering process includes a smoothing process, and the smoothing process uses a representative value of the local area around the processing target point in the point cloud data to smooth the data of the processing target point.
[0322] Embodiment 10: The image processing device according to embodiment 1, wherein the filter processing unit performs filter processing on position information related to points in the point cloud data.
[0323] Embodiment 11: The image processing device according to embodiment 1, wherein the filter processing unit performs filter processing on attribute information related to points in the point cloud data.
[0324] Embodiment 12: An image processing method, comprising:
[0325] performing filtering processing on the point cloud data using a representative value of the point cloud data of each local area obtained by dividing the three-dimensional space; and
[0326] A two-dimensional plane image on which the point cloud data subjected to the filtering process is projected is encoded and a bit stream is generated.
[0327] Embodiment 13: An image processing device, comprising:
[0328] A decoding unit that decodes the bit stream and generates encoded data of a two-dimensional plane image on which the point cloud data is projected; and
[0329] A filter processing unit performs a filter process on the point cloud data restored from the two-dimensional plane image generated by the decoding unit, using a representative value of the point cloud data of each local area obtained by dividing a three-dimensional space.
[0330] Embodiment 14: An image processing method, comprising:
[0331] decoding the bit stream and generating encoded data of a two-dimensional plane image on which the point cloud data is projected; and
[0332] Using a representative value of the point cloud data of each local area obtained by dividing the three-dimensional space, a filtering process is performed on the point cloud data restored from the generated two-dimensional plane image.
[0333] Embodiment 15: An image processing device, comprising:
[0334] a filtering processing unit that performs filtering processing on some points in the point cloud data; and
[0335] An encoding unit that encodes a two-dimensional plane image on which the point cloud data subjected to the filter processing by the filter processing unit is projected and generates a bit stream.
[0336] Embodiment 16: The image processing device according to Embodiment 15, wherein the filtering processing unit performs the filtering processing on points in the point cloud data corresponding to ends of patches included in the two-dimensional plane image.
[0337] Embodiment 17: The image processing device according to Embodiment 15, wherein the filtering process includes a smoothing process, and the smoothing process uses data of points around the processing target point in the point cloud data to smooth the data of the processing target point.
[0338] Embodiment 18: An image processing method, comprising:
[0339] performing filtering on some points in the point cloud data; and
[0340] A two-dimensional plane image on which the point cloud data subjected to the filtering process is projected is encoded and a bit stream is generated.
[0341] Embodiment 19: An image processing device, comprising:
[0342] A decoding unit that decodes the bit stream and generates encoded data of a two-dimensional plane image on which the point cloud data is projected; and
[0343] A filter processing unit that performs filter processing on some points in the point cloud data restored from the two-dimensional plane image generated by the decoding unit.
[0344] Embodiment 20: An image processing method, comprising:
[0345] decoding the bit stream and generating encoded data of a two-dimensional plane image on which the point cloud data is projected; and
[0346] A filtering process is performed on some points in the point cloud data restored from the generated two-dimensional plane image.
[0347] Reference numerals list
[0348] 100 Encoding device
[0349] 111 Patch Decomposition Unit
[0350] 112 Packing Units
[0351] 113OMap generation unit
[0352] 114 Auxiliary patch information compression unit
[0353] 115 Video Encoding Unit
[0354] 116 Video Encoding Unit
[0355] 117OMap coding unit
[0356] 118 Multiplexer
[0357] 131 Patch decomposition processing unit
[0358] 132 Geometry decoding unit
[0359] 133 3D position information smoothing processing unit
[0360] 134 Texture Correction Unit
[0361] 141 Regional Division Unit
[0362] 142 Area representative value output unit
[0363] 143 Processing target area setting unit
[0364] 144 Smoothing Units
[0365] 145 Sending information generation unit
[0366] 200 Decoding device
[0367] 211 Demultiplexer
[0368] 212 Auxiliary patch information decoding unit
[0369] 213 Video decoding unit
[0370] 214 Video decoding unit
[0371] 215OMap decoding unit
[0372] 216 Unpacking Unit
[0373] 217 3D reconstruction unit
[0374] 231 Geometric point cloud generation unit
[0375] 232 3D position information smoothing processing unit
[0376] 233 Texture Synthesis Unit
[0377] 251 Send information acquisition unit
[0378] 252 Regional Division Unit
[0379] 253 Area representative value output unit
[0380] 254 Processing target locale unit
[0381] 255 smoothing units
Claims
1. An image encoding device, comprising: The circuit is configured as: Obtain a patch of 3D data representing a three-dimensional 3D structure using a plurality of points; Performing filtering processing on the plurality of points, wherein the filtering processing includes the following processing: determining whether the current position in the filtering process is at the end of the patch, On the basis of determining that the current position is located at the end of the patch, setting the processing range of the filtering process to the current position, upon determining that the current position is not located at an end of the patch, excluding the current position from the filtering process; and encoding a two-dimensional plane image on which the 3D data subjected to the filtering process is projected, and A bit stream including the encoded two-dimensional planar image is generated.
2. The image encoding device according to claim 1, wherein: The circuit is configured as: dividing the three-dimensional space including the 3D data into a plurality of local areas; and performing the filtering process on the processing target point at the current position using a representative value of the 3D data of a nearby local area, wherein the plurality of local areas include the nearby local area, and The nearby local area is located around the processing target point.
3. The image encoding device according to claim 1, wherein: The 3D data is represented as point cloud data, The point cloud data includes an occupancy map indicating whether position information and attribute information for the plurality of points exist at each position on the two-dimensional plane, and The patches are arranged in an occupancy map of the point cloud data.
4. The image encoding device according to claim 2, wherein: The local area includes a cubic area or a rectangular parallelepiped area having a predetermined size.
5. The image encoding device according to claim 2, wherein: The circuit is configured to generate the bitstream including information about the local area, The information about the local area includes information about the size of the local area, or information about the shape of the local area, or information about the size and shape of the local area.
6. The image encoding device according to claim 2, wherein: The representative value includes an average value or a median value of the 3D data included in the nearby local area.
7. An image encoding method, comprising: Obtain a patch of 3D data representing a three-dimensional 3D structure using a plurality of points; Performing filtering processing on the plurality of points, wherein the filtering processing includes the following processing: determining whether the current position in the filtering process is at the end of the patch, On the basis of determining that the current position is located at the end of the patch, setting the processing range of the filtering process to the current position, upon determining that the current position is not located at an end of the patch, excluding the current position from the filtering process; and encoding a two-dimensional plane image on which the 3D data subjected to the filtering process is projected, and A bit stream including the encoded two-dimensional planar image is generated.
8. An image decoding device, comprising: The circuit is configured as: Decode the bitstream; generating encoded data of a two-dimensional plane image, wherein 3D data representing a three-dimensional 3D structure using a plurality of points is projected onto the two-dimensional plane image; and performing filtering processing on the 3D data restored from the two-dimensional planar image, The filtering process includes the following processes: determining whether a current position in the filtering process is at an end of a patch of the 3D data, On the basis of determining that the current position is located at the end of the patch, setting the processing range of the filtering process to the current position, Upon determining that the current position is not located at an end of the patch, the current position is excluded from the filtering process.
9. The image decoding device according to claim 8, wherein: The circuit is configured as: performing the filtering process on the processing target point at the current position using a representative value of the 3D data of a nearby local area, The three-dimensional space including the 3D data is divided into a plurality of local areas, and the nearby local area is included in the plurality of local areas and is located around the processing target point.
10. The image decoding device according to claim 8, wherein: The 3D data is represented as point cloud data, The point cloud data includes an occupancy map indicating whether position information and attribute information for the plurality of points exist at each position on the two-dimensional plane, and The patches are arranged in an occupancy map of the point cloud data.
11. An image decoding method, comprising: Decode the bitstream; generating encoded data of a two-dimensional plane image, wherein 3D data representing a three-dimensional 3D structure using a plurality of points is projected onto the two-dimensional plane image; and performing filtering processing on the 3D data restored from the two-dimensional planar image, The filtering process includes the following processes: determining whether a current position in the filtering process is at an end of a patch of the 3D data, On the basis of determining that the current position is located at the end of the patch, setting the processing range of the filtering process to the current position, Upon determining that the current position is not located at an end of the patch, the current position is excluded from the filtering process.