Image processing device and method, program, and storage medium
The image processing device addresses inaccuracies in 3D data by region-based coordinate adjustments and surface replacements, enhancing the viewing experience of 3D data with occluded areas.
Patent Information
- Application Number
- JP2024078074
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2025-11-26
AI Technical Summary
Existing 3D data generation methods using deep learning for inpainting occluded areas can result in inaccuracies and high computational costs, leading to a sense of incongruity when viewing 3D data from different viewpoints.
An image processing device that divides image data into first and second regions, adjusting coordinate information to minimize differences between regions, and optionally replacing low-reliability areas with primitive surface shapes to reduce the sense of incongruity.
Reduces the sense of discomfort and incongruity when viewing 3D data with occluded areas by minimizing perspective conflicts and filling gaps with accurate, continuous surfaces.
Smart Images

Figure 2025172523000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device that generates three-dimensional data. [Background technology]
[0002] Conventionally, technologies capable of acquiring distance distribution information (hereinafter referred to as distance information acquisition technologies), such as stereo cameras and time-of-flight cameras, have been known. A point cloud can be obtained by performing perspective projection transformation on the distance distribution information. Furthermore, by converting the point cloud into polygons, it is possible to generate a 3D surface model with a surface.
[0003] By acquiring images (color distribution information) simultaneously with distance distribution information, it is possible to generate 3D data with texture information. Unlike ordinary 2D images, 3D data has the advantage that it can be enjoyed from any viewpoint.
[0004] In 3D data generated based on distance distribution information from a single viewpoint, there may be areas that exist in the actual subject but cannot be restored as 3D data. For example, when creating 3D data of a scene in which a hand is placed in front of a body, the area of the body that is occluded by the hand from the front is considered an occluded area and information is lost. When such 3D data is rendered from a viewpoint other than the front, the loss of the occluded area becomes visible, creating an unnatural viewing experience.
[0005] Non-Patent Document 1 discloses a technology that applies inpainting processing to repair defects that occur in parts of the background due to occlusion of the foreground. The inpainting processing based on deep learning infers shape and texture information of the occluded area and repairs the defect. This makes the occluded area invisible even when viewing 3D data from different viewpoints. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] J. Kopf, "One Shot 3D Photography", ACM Trans. Graph., Vol. 39, No. 4, Article 76., 2020 Summary of the Invention [Problem to be solved by the invention]
[0007] However, the above-mentioned prior art infers information about the missing part based on deep learning, which may result in information that is significantly different from the actual subject, potentially creating a stronger sense of incongruity. Furthermore, the use of deep learning generally involves high computational costs.
[0008] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to provide an image processing device that can reduce the sense of discomfort when viewing three-dimensional data that includes an occluded area. [Means for solving the problem]
[0009] The image processing device according to the present invention is characterized by comprising a division means for dividing image data into a first region and a second region, and a conversion means for converting at least one of first coordinate information, which is coordinate information of a subject in the first region, and second coordinate information, which is coordinate information of a subject in the second region, based on the first coordinate information and the second coordinate information, so that the difference between a first representative value representing the first coordinate information and a second representative value representing the second coordinate information becomes small. [Effects of the Invention]
[0010] According to the present invention, it is possible to reduce the sense of discomfort felt when viewing three-dimensional data that includes an occluded area. [Brief explanation of the drawings]
[0011] [Figure 1]1 is a block diagram showing the configuration of an image processing apparatus according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing the configuration of an imaging unit. [Figure 3] 4 is a flowchart showing a three-dimensional data generation process according to the first embodiment. [Figure 4] FIG. 10 is a diagram for explaining region division. [Figure 5] FIG. 10 is another diagram for explaining region division. [Figure 6] FIG. 10 is another diagram for explaining region division. [Figure 7] 10A and 10B are diagrams for explaining reduction in exposure of occluded areas; [Figure 8] FIG. 1 is a conceptual diagram of a first neighborhood region and a second neighborhood region. [Figure 9] 10 is a flowchart showing a three-dimensional data generation process according to a first modification of the first embodiment. [Figure 10] 10 is a flowchart showing a three-dimensional data generation process in a second modification of the first embodiment. [Figure 11] 10A and 10B are diagrams for explaining how the exposure of the occluded area is reduced by the second modification of the first embodiment. [Figure 12] 10 is a flowchart showing a three-dimensional data generation process according to a second embodiment. [Figure 13] FIG. 10 is a diagram for explaining region division according to the second embodiment. [Figure 14] 10A and 10B are diagrams for explaining how the sense of incongruity in shape and the appearance of occluded areas are reduced. [Figure 15] FIG. 10 is a diagram illustrating replacement of distance values according to the second embodiment. [Figure 16] A conceptual diagram of a sphere set as the primitive surface shape. [Figure 17] 10 is a flowchart showing a three-dimensional data generation process according to a first modified example of the second embodiment. [Figure 18] FIG. 10 is a diagram showing distance distribution information in Modification 2 of the second embodiment. [Figure 19] FIG. 10 is a diagram for explaining the effect of the second modification of the second embodiment. [Figure 20] A conceptual diagram of setting a primitive surface shape that is rotationally asymmetric about the Z axis. [Figure 21] 10A and 10B are side views of the three-dimensional data in the first state, the transition state, and the second state of the motion data generation. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0013] First, the terms and prerequisites used in this specification and drawings will be explained.
[0014] "3D data" refers to data that includes 3D shape information, including point clouds and 3D surface models. It may also include color information in addition to shape information.
[0015] When rendering a point cloud, it is common to represent it by placing point-like objects of finite size at the vertices of the point cloud. This representation will be used when showing the rendering results of a point cloud in the accompanying drawings.
[0016] "Coordinate information" includes both distance distribution information and coordinate values of the point cloud.
[0017] The axis extending from the origin of the distance distribution information acquisition means along the optical axis is defined as the Z axis, and the two axes that intersect at right angles in a plane perpendicular to the optical axis are defined as the X and Y axes. The "depth axis" corresponds to the Z axis, and the "distance value" corresponds to the Z coordinate value.
[0018] "Distance distribution information" refers to information in which the distance values of a subject are projected onto a two-dimensional plane (XY plane) perpendicular to the optical axis direction. When distance distribution information is shown in the attached drawings, it is expressed graphically by corresponding the distance values to different shades of color, with closer distances being white and longer distances being black. In addition, the term "pixel" may be used to refer to one element of distance distribution information.
[0019] The terms "image" and "two-dimensional image" refer to texture information (color distribution information) of a subject.
[0020] An "occluded area" refers to an area of a subject that is not visible from the viewpoint where distance distribution information is acquired, and corresponds to an area that is missing in 3D data. An "occluded area" refers to an area that is visible from the viewpoint and occludes an occluded area. In distance distribution information projected onto a 2D plane, the occluded area and the occluded area coincide.
[0021] The "first region" and "second region" refer to two-dimensional regions for distance distribution information and three-dimensional regions for point clouds.
[0022] "Primitive surface shapes" refer to the shapes of surfaces included in basic three-dimensional figures, such as planes, spheres, ellipsoids, cylinders, and tori.
[0023] "Reliability" refers to the degree of reliability of coordinate information. For example, a high (low) reliability coordinate value indicates that the measurement accuracy of that coordinate value is high (low). Also, a high (low) reliability area refers to an area consisting of high (low) reliability coordinate values.
[0024] (First embodiment) FIG. 1 is a block diagram showing the configuration of an image processing apparatus 600 according to the first embodiment of the present invention.
[0025] The image processing device 600 has an imaging unit 601, which functions as an acquisition unit that acquires distance distribution information and an image of a subject. The control unit 602 is a control device such as a CPU, and controls the operation of each block of the image processing device 600. For example, it inputs information acquired by the imaging unit 601 to an image processing unit 605, and executes control such as storing the processing results in a storage device 604.
[0026] The memory 603 is a storage device such as a volatile memory used for temporary information storage. The storage device 604 is a device capable of permanently storing information, such as an HDD, and stores imaging information acquired by the imaging unit 601, parameters used by the image processing unit 605, and the like.
[0027] The image processing unit 605 applies predetermined image processing, including a point cloud generation method described below, to the information (image data) acquired by the imaging unit 601. The image processing unit 605 may be configured as an integrated circuit or may be a functional module realized by software. In this embodiment, the image processing unit 605 functions as a point cloud generation means.
[0028] Next, the detailed configuration of the imaging unit 601 will be described. The imaging unit 601 can have any configuration as long as it functions as an acquisition means capable of acquiring distance distribution information. More preferably, it should be configured to be capable of simultaneously acquiring color distribution information from the same viewpoint as the distance distribution information. Examples of configurations that realize this function include a stereo camera equipped with two imaging systems each consisting of an optical system capable of acquiring RGB images and an imaging element, and a ToF camera equipped with a ToF module capable of acquiring distance distribution information and an imaging system capable of acquiring RGB images. Below, as an example, a configuration of the imaging unit 601 that can simultaneously acquire distance distribution information and color distribution information using an imaging surface phase difference detection method will be described.
[0029] 2 is a diagram showing the configuration of the imaging unit 601. The imaging unit 601 is configured to include an imaging optical system 701 and an imaging element 702. Light from a subject located on the object plane is collected by the imaging optical system 701 and imaged on the imaging element 702, which is located in a position that is approximately optically conjugate with the object plane, thereby capturing an image.
[0030] The image sensor 702 is configured with, for example, tens of millions of pixels 703, each equipped with a photoelectric conversion element, arranged in a grid pattern. Each pixel is provided with a color filter that transmits specific wavelengths such as red, green, and blue, and by arranging these in, for example, a Bayer array, it becomes possible to acquire color distribution information.
[0031] Each pixel 703 is configured with a microlens 704, a first photoelectric conversion unit 705, and a second photoelectric conversion unit 706, and acquires different optical information according to its position using the two photoelectric conversion units 705, 706 arranged in the horizontal direction (X direction). From the received light information of all pixels, a first image configured as the luminance distribution of light received by the first photoelectric conversion unit 705 and a second image configured as the luminance distribution of light received by the second photoelectric conversion unit 706 are obtained.
[0032] The incident surface of the image sensor 702 and the light-receiving surfaces of the first photoelectric conversion unit 705 and the second photoelectric conversion unit 706 are in a nearly Fourier conjugate relationship due to the action of the microlens 704. Therefore, the light-receiving surfaces of the photoelectric conversion units 705 and 706 and the exit pupil of the image pickup optical system 701 are in a nearly optically conjugate relationship. The position distribution in the direction perpendicular to the optical axis at the exit pupil corresponds to the position distribution in the direction perpendicular to the optical axis at the light-receiving surfaces of the photoelectric conversion units. Therefore, by providing two different photoelectric conversion units, it is possible to separate and receive light beams that have passed through different pupil regions of the image pickup optical system 701. The first and second images described above are luminance distribution information obtained by light beams that have passed through different pupil regions.
[0033] Light rays that form an image on the imaging surface (ideally) enter the same point on the image sensor 702 regardless of the position on the pupil through which they pass. However, for defocused light rays, the position at which they enter the image sensor 702 changes depending on the position on the pupil through which they pass. In other words, an image shift occurs according to the amount of defocus. If the amount of this image shift can be calculated, it is possible to finally obtain distance distribution information by converting the amount of image shift into the amount of defocus.
[0034] The amount of image shift can be calculated, for example, by stereo matching between the first and second images. A patch of a small area in one image is matched with the other image along the epipolar line, and the image shift amount is found by identifying the position with the highest correlation. The image shift amount can be converted to actual distance using optical parameters such as focal length based on the image shift amount.
[0035] Next, the image processing performed by the image processing device 600 configured as described above will be described. Fig. 3 is a flowchart showing the operation of the image processing device 600 in this embodiment. The processing shown in Fig. 3 is realized by the control unit 602 of the image processing device 600 executing a control program stored in the storage device 604.
[0036] <Region division step S101> In the region dividing step S101, the control unit 602 divides the distance distribution information in the image data into a first region and a second region.
[0037] The region division is performed so that the foreground region including the occluded region is the first region, and the other region is the second region. For example, in the case of a scene in which a circular object is positioned in the foreground and a rectangular object is positioned in the background so that they overlap, as shown in Figure 4(a), an occluded region (and an occluded region) like the shaded area in Figure 4(b) will occur. In this case, as shown in Figure 4(c), the region division is performed so that the circular object including the occluded region is the first region (shaded area), and the visible region of the rectangular object is the second region (dotted area).
[0038] The above-mentioned scene includes multiple subjects, but even with a single subject, the subject may be occluded by a portion of the subject. Figure 5(a) shows distance distribution information for a scene in which a person holds out the part of their right hand from their elbow in front of their body. In this scene, the area from their elbow to their hand (the shaded area in Figure 5(b)) is the occluded area, which results in an occluded area in parts of their arm and body. In this case, as shown in Figure 5(c), the area including the occluded area from their elbow to their hand is the first area (shaded area), and the other area on the side of their body is the second area (dotted area).
[0039] In addition, in the case of a scene in which a person holds both hands out in front of their body, as shown in Figure 6(a), multiple regions that are not connected to each other become occluded regions, as shown in Figure 6(b). In this case, as shown in Figure 6(c), a single occluded region from the elbow to the hand is defined as a sub-region. The collection of multiple sub-regions is defined as the first region (shaded area), and the remaining region on the side of the body is defined as the second region (dotted area).
[0040] The above-mentioned region segmentation can be performed, for example, by referencing a two-dimensional image and performing segmentation based on deep learning. When a person is the subject, the hands and arms (hereinafter referred to as "hands and arms") tend to occlude the body side. Therefore, by using a model that has been trained to label the hand and arm regions differently from other regions, it is possible to segment the hand and arm regions that may be located in the foreground and the body regions that may be located in the background.
[0041] <Offset Addition Step S102> In the offset addition step S102, the control unit 602 adds an offset value to the distance distribution information of the first region. Here, the offset value is set so as to reduce the distance between a first representative coordinate value representing the coordinate values of the first region and a second representative coordinate value representing the coordinate values of the second region. In this embodiment, the coordinate value refers to the distance value in the distance distribution information. The representative coordinate value representing the coordinate values of a certain region is calculated by applying some rule to a set of coordinate values belonging to the certain region. For example, a rule may be used such that the distance value of the point closest to the center of gravity in the XY plane of the certain region is adopted as the representative coordinate value, but this rule is not limited to this.
[0042] In a point cloud that includes a perspective conflict edge (a boundary where distance values change discontinuously), the distance between vertices becomes large around the perspective conflict edge. When the point cloud is viewed obliquely (from a viewpoint not parallel to the optical axis), a gap is visible in the area where the distance between vertices is large. When the distance distribution information shown in Figure 4(a) is converted into a point cloud and viewed obliquely, a large gap can be seen between the foreground subject and the background subject, as shown in Figure 7(a). This is the area of the background that is occluded by the foreground, i.e., the occluded area. When an occluded area appears, there is no information in an area where information is (expected to) actually exist, which creates a significant sense of incongruity.
[0043] In this case, as in the present embodiment, by adding an offset value to the coordinate information so as to reduce the distance between the first representative coordinate value and the second representative coordinate value, the distance difference at the boundary between the first area and the second area can be reduced. In other words, the degree of perspective conflict is reduced, and the gap that is visually recognized when converted into 3D data is reduced as shown in Figure 7(b). This makes it possible to reduce the sense of incongruity when viewing.
[0044] In this embodiment, the addition of an offset value generates 3D data that differs from the actual positional relationship of the subject when the distance distribution information was acquired. However, because the shape information within the first and second regions is maintained, the sense of incongruity caused by the change in positional relationship is not significant. The sense of incongruity caused by the change in positional relationship is sufficiently small compared to the sense of incongruity caused by the appearance of an occluded region. Therefore, by applying the method of this embodiment, it is possible to reduce the sense of incongruity in 3D data that includes perspective conflict.
[0045] As mentioned above, the 3D data in FIG. 7 is shown as a point cloud. When this is further surfaced, polygons are connected between the points of the perspective conflict edge, and the gaps in FIG. 7 are filled with polygons. In the case of textured 3D data, the texture on this polygon is a stretched version of the color information near the perspective conflict edge. The polygons connecting the edges do not exist in the actual subject, and the larger the distance difference at the perspective conflict edge, the larger the area becomes, which increases the sense of incongruity when viewed from an angle. Therefore, even in such a form of 3D data, the method of this embodiment can be applied to reduce the distance difference and thereby achieve the effect of reducing the sense of incongruity.
[0046] More preferably, the first representative coordinate value (representative value) is a statistic of coordinate values in a first neighboring region corresponding to the vicinity of the boundary between the first region and the second region, and the second representative coordinate value (representative value) is a statistic of coordinate values in a second neighboring region corresponding to the vicinity of the boundary between the second region and the first region. The aforementioned gap can be reduced by reducing the perspective difference around the perspective conflict edge. Therefore, it is desirable to calculate the representative coordinate value from the coordinate values near the boundary so as to reduce the distance difference near the boundary between the first region and the second region.
[0047] The first (second) neighboring region refers to the region encompassed by pixels from the pixel corresponding to the boundary between the first and second regions to a pixel a predetermined distance away toward the first (second) region. Figure 8 is a conceptual diagram of the first (second) neighboring region, with the dark gray area representing the first neighboring region and the light gray area representing the second neighboring region. Here, the "predetermined length" may be, for example, an index of four pixels from the boundary pixel, or it may be 20 pixels. Alternatively, the diagonal length of the entire distance distribution information may be 12 pixels, and the distance from the boundary pixel may be 0.05 × 12 pixels. The statistical quantity may refer to, for example, the mean, median, mode, etc.
[0048] When the first region contains multiple sub-regions, it is desirable to add a different offset value to each sub-region, as shown in Figure 6. Since the perspective difference between the second region and multiple sub-regions may differ, adding a different offset value to each sub-region makes it possible to reduce gaps after conversion into 3D data, regardless of the magnitude of the perspective difference.
[0049] <Point cloud formation process S103> In the point cloud generation step S103, the control unit 602 converts the distance distribution information processed in S102 into a point cloud. Generally, distance distribution information acquired by a distance measurement device (in this embodiment, a distance measurement device using an image-pickup-surface phase-difference detection method) is perspectively projected. Therefore, by backprojecting based on the distance distribution information and the conditions at the time of acquisition, the distance distribution information can be converted into a point cloud, which is a set of coordinate values of three-dimensional points. Among the XYZ coordinates, the Z coordinate is the distance value of the distance distribution information itself. Regarding the X and Y coordinates, the pixel coordinates based on the end points of the distance distribution information are (u, v), the pixel coordinates corresponding to the center of the distance distribution information are (uc, vc), the pixel interval is p, the focal length of the distance measurement device is f, and the distance value at pixel coordinates (u, v) is z(u, v), and they can be calculated using the following formula:
[0050] X=Z(u,v)×(u-uc)p / f …Equation (1) Y = Z(u,v) × (v-vc)p / f ...Equation (2) <Surface processing> Although not shown in Figure 3, a surfacing process may be performed after the point cloud generation process. In the surfacing process, three adjacent points are connected as polygons to form a mesh, thereby creating 3D data with a surface. Furthermore, if there is a two-dimensional image (color distribution information) from the same viewpoint as the distance distribution information, the color distribution information can be assigned to the surface as a texture to create 3D data with color information.
[0051] In surfacing, the presence or absence of polygon connections may be determined based on the distance between points. In the example of FIG. 7(b), there is no physical continuity between the foreground and background objects. However, when the distance between vertices at the boundary is reduced by applying the method of this embodiment, generating a continuous surface at the boundary may result in more natural-looking three-dimensional data.
[0052] [Variation 1] In the above embodiment, the process of dividing distance distribution information into regions and adding an offset value to the first region has been described. However, the effects of this embodiment can also be obtained with other processes. Below, the contents of the process shown in the flowchart of FIG. 9 will be described as Modification 1 of the first embodiment.
[0053] S201 is a step of converting distance distribution information into a point cloud, and the same processing as in step S103 in FIG. 3 is performed.
[0054] S202 is a step of dividing the point cloud into a first region and a second region. Region division can be performed, for example, by referencing point cloud data and performing segmentation based on deep learning. When a person is the subject, by using a model trained to assign different labels to the hand / arm region and other regions based on the point cloud data, it is possible to assign different labels to the hand / arm region that may be located in the foreground and the body region that may be located in the background.
[0055] Step S203 is a step of adding an offset value to the coordinate information of the first region. In this modification, the coordinate information refers to the coordinate values of the point cloud. The offset value may be calculated based on the distance distribution information or the coordinate values of the point cloud.
[0056] When calculating the offset value based on the distance distribution information, the calculation may be performed in the same manner as in the process in S102 of Fig. 3. The offset value is added by uniformly adding this offset value to the coordinate (Z coordinate) values along the optical axis of the point group belonging to the first region.
[0057] When calculating offset values based on coordinate values of a point cloud, offset values in the XYZ directions may be calculated so as to reduce gaps at perspective conflict edges. An example of this processing will be described below.
[0058] Let U1 be the set of points belonging to the first region that are close to the perspective conflict edge, and U2 be the set of points belonging to the second region that are close to the perspective conflict edge. For each point in U1, the closest point in U2 is identified and the Euclidean distance to that closest point is evaluated. The average value of this distance value for all points in U1 is used as the index. By adding offset values to the points in the first region, XYZ offset values can be adopted that reduce this index.
[0059] The above process is an example and is not limited to this process. For example, the index value does not need to be evaluated at all points within U1, but may be evaluated at randomly sampled points. Also, the evaluation value may be the coordinate difference in the Z direction instead of the Euclidean distance. In this case, the XY offset value will be 0.
[0060] [Variation 2] In addition to the processing of the first embodiment, by applying filtering processing such as smoothing to the neighboring region, it becomes possible to generate 3D data that looks less unnatural. Below, as Modification 2 of the first embodiment, the content of the processing shown in the flowchart of Fig. 10 will be described.
[0061] The area dividing step S301 and the offset value adding step S302 are similar to the area dividing step S101 and the offset value adding step S102 in FIG. 3, respectively.
[0062] In S303, the control unit 602 applies filtering to the region including the first neighboring region and the second neighboring region. Adding the offset value reduces the distance difference around the boundary compared to before the processing, but in most cases it does not become zero, and some level difference may remain. This filtering process smooths the level difference, making it possible to create 3D data that is less unnatural. For example, a smoothing filter or a median filter is used for the filtering process.
[0063] Fig. 11 shows a perspective view of 3D data generated by applying the processing of Fig. 10 to the distance distribution information of Fig. 4(a). Compared to Fig. 7(b), which shows the result of the first embodiment, the distance values around the near and far edges change continuously, so gaps become almost invisible, further reducing the sense of incongruity.
[0064] [Variation 3] If the distance distribution information of the first region is Z1(u,v), the distance distribution information obtained by adding an offset to Z1 is Z1'(u,v), and the distance distribution information used in the step of calculating coordinate values in the direction perpendicular to the optical axis direction in the first region in the point group generation step is Z1"(u,v), it is desirable that the difference between Z1 and Z1" be smaller than the difference between Z1 and Z1'. The reason for this will be explained below.
[0065] As mentioned above, in the point grouping step S103, the coordinate values of the X and Y coordinates, which are in the direction perpendicular to the optical axis, are calculated according to formulas (1) and (2). When Z1(u,v), Z1'(u,v), and Z1"(u,v) are substituted for Z(u,v) in formula (1), the left-hand sides are set to X1, X1', and X1", respectively. In this case, since Z1≠Z1', X≠X'.
[0066] Z1 represents the distance distribution of the subject when the distance distribution information was acquired, and therefore X1 calculated using Z1 also represents the X position of the subject. On the other hand, Z1' is obtained by adding an offset value to Z1 to correct the distance distribution of the subject, and therefore X1' calculated using Z1' does not faithfully represent the X position of the subject. X1' obtained from distance distribution information to which a positive offset value has been added is larger than X1 according to equation (1). In other words, the 3D data is expanded in the X and Y directions. Compared to the 3D data of the second region, the 3D data of the first region is expanded only in the X and Y directions, which may cause a strange viewing experience in some cases.
[0067] From the above, it is desirable that the distance distribution information Z1" used when calculating the XY coordinates of the point cloud in the first region is closer to Z1 than Z1' after adding the offset value. It is more desirable that Z1"=Z1.
[0068] Examples of evaluation indices for the difference include average error and root mean square error. When the condition of this modification is expressed using the average error as an index, it becomes as follows:
[0069]
number
[0070] (Second embodiment) In the first embodiment, it is assumed that the coordinate information of the first and second regions is highly reliable enough to be used for generating three-dimensional data. In the second embodiment, a case will be described in which the second region includes low-reliability coordinate information.
[0071] Generally, distance information acquisition means have a limited range within which they can acquire distance information with high accuracy (hereinafter referred to as the ranging range). While it is common to generate 3D data from only highly reliable coordinate information, depending on the application, it may be desirable to record color information from low-reliability areas outside the ranging range as 3D data. For example, when photographing a person as a subject, such as a portrait, it may be desirable to record information about the distant background, which is outside the ranging range, together with the image. However, because distance information outside the ranging range contains significant random and systematic errors, generating 3D data based on low-reliability coordinate information is likely to cause discomfort when viewed visually.
[0072] Therefore, for low-reliability areas, it is possible to convert them into 3D data by adding virtual shapes (primitive surface shapes) such as planes or spheres instead of actual coordinate information, and then combine them with the 3D data of high-reliability areas.
[0073] However, since information is lost in the background behind the subject due to occlusion by the subject, simply arranging the three-dimensional data of the subject and the three-dimensional data of the background side by side can cause the loss of information in the background to be visible, which can create a sense of discomfort when viewed.
[0074] In the second embodiment, a method for solving this problem will be described with reference to the flowchart in FIG.
[0075] <Region division process S401> In the region division step S401, the control unit 602 divides the distance distribution information into a first region and a second region. Here, the second region is divided so that it has a relatively low reliability compared to the first region. "Relatively low reliability" means that the second region has a statistically low reliability compared to the first region. For example, the region may be divided so that the average reliability of the coordinate information included in the first region is higher than the average reliability of the coordinate information included in the second region. Because this is a statistical comparison, low-reliability distance information may be included in part of the first region, and conversely, high-reliability distance information may be included in part of the second region.
[0076] As an example, consider a scene in which a subject having a three-dimensional shape obtained by cutting a sphere (hereinafter referred to as a spherical subject) is within the ranging area, and a planar background is outside the ranging area. FIG. 13(a) shows distance distribution information for this scene. Because the background is outside the ranging area, the distance values of the background area are not acquired correctly, and as shown in FIG. 13(b), the distance distribution information contains a large amount of error, resulting in distance distribution information that differs from the actual distance distribution information (FIG. 13(a)). In such an example, as shown in FIG. 13(c), the area corresponding to the spherical subject within the ranging area and having a relatively high reliability is designated as the first area, and the area corresponding to the background outside the ranging area and having a relatively low reliability is designated as the second area.
[0077] The region division may be performed by, for example, performing threshold processing based on reliability information corresponding to the accuracy of distance measurement, and dividing the region into high-reliability regions with reliability equal to or greater than the threshold as the first region and low-reliability regions with reliability less than the threshold as the second region. In this case, the reliability of the second region is always lower than that of the first region.
[0078] Alternatively, a semantic region detection method based on deep learning that detects a target object region may be used to divide the detected object region into a first region and other regions into a second region. In this case, the first region may contain low-reliability distance information, and the second region may contain high-reliability distance information, but as described above, it is sufficient that the reliability of the second region is statistically lower than that of the first region.
[0079] <Coordinate value replacement step S402> In the coordinate value replacement step S402, the control unit 602 replaces some of the distance values of the second region and the first region with distance values corresponding to the primitive surface shape. The reason why this step can reduce the sense of incongruity when viewed visually will be explained below with reference to Fig. 14. Fig. 14(c) shows a case where the method of the second embodiment is used, Fig. 14(a) shows Comparative Example 1, and Fig. 14(b) shows Comparative Example 2. Here, a plane perpendicular to the Z axis is set as the primitive surface shape.
[0080] First, we will explain why replacing the distance values in the second region with distance values corresponding to primitive surface shapes can reduce the sense of incongruity. When 3D data is generated based on distance information in the second region that includes low-reliability distance values, random errors can cause high-frequency uneven shapes that do not actually exist. This can cause discomfort to the viewer, as in Comparative Example 1 in Figure 14(a). Therefore, the distance information in the second region is replaced so that it corresponds to primitive surface shapes, which are basic figures. This can reduce the sense of incongruity caused by distance value measurement errors, as in Comparative Example 2 in Figure 14(b).
[0081] In addition, in this embodiment, the distance values of some of the first regions are replaced. This makes it possible to further reduce the sense of incongruity felt when viewed visually. The reason for this will be explained below.
[0082] 14(b), because there is a distance difference at the boundary between the first and second regions in the distance distribution information, a gap is visually perceived between the first and second regions in the 3D data generated based on this distance distribution information. In this way, when the 3D data of the first region and the 3D data of the second region converted into a primitive surface shape are separated, the loss of information in the second region caused by the occlusion of the first region is visually perceived as a gap, creating an unnatural viewing experience.
[0083] On the other hand, in this embodiment, not only the distance values of the second region but also some of the distance values of the first region are replaced with distance values corresponding to the primitive surface shape. FIG. 14(c) shows the state in which the area of the first region whose Z coordinate is equal to or greater than ZP and the entire second region are replaced with a plane of Z=ZP. Here, the area of the first region before replacement is represented by a dotted line at the bottom of FIG. 14(c). In this way, a primitive surface is set near the coordinate value of the first region, and part of the first region is replaced with the primitive surface. This results in at least a part of the three-dimensional data of the first region and the three-dimensional data of the second region, which is a primitive surface, being continuously connected. As a result, gaps due to missing information in the second region are less visible, making it possible to reduce the sense of incongruity felt when viewing.
[0084] The method for generating the distance distribution information shown in the upper part of FIG. 14(c) in this embodiment will be described in more detail with reference to FIG. 15. For the second region, as shown in the upper part of FIG. 15(a), the distance distribution information has a large error, and the region to be replaced is the entire second region, as shown in the diagonal line in the middle part of FIG. 15(a). When replacing the primitive surface shape with a plane perpendicular to the Z axis, all of the second region is replaced with a constant distance value, as shown in the lower part of FIG. 15(a). On the other hand, since the first region has relatively reliable distance distribution information, only a portion of the first region needs to be replaced with a primitive surface shape to maintain continuity with the second region. When a circular region shown in the middle part of FIG. 15(b) is set as the region to be replaced, the distance values of this region can be replaced with the primitive surface shape (in this case, a uniform distance value because the plane is perpendicular to the Z axis).
[0085] The distance value to be replaced in the first region is preferably either the distance value of a point in the distance distribution information of the first region that is greater than or smaller than the Z coordinate of the primitive surface shape. In other words, it is preferable to set the primitive surface so that at least some of the points in the first region are located on the opposite side of the primitive surface. This is because, in order for the three-dimensional data of the first region and the primitive surface shape to be continuously connected, they need to be close enough to intersect.
[0086] 14 and 15, the background on the far side of the spherical subject is replaced with a primitive surface shape, so the coordinate values of the first region that are on the far side of the primitive surface shape (those with larger Z coordinates) are targeted for replacement with the primitive surface shape. The case where the coordinate values of the first region that are on the near side of the primitive surface shape are replaced with the primitive surface shape will be described in Modification 2.
[0087] It is desirable that the setting of the primitive surface is based on the coordinate values of a first neighboring region of the first region, which corresponds to the vicinity of the boundary with the second region. This is because, in order to prevent the occurrence of gaps such as those in Comparative Example 2 of Fig. 14(b), it is necessary to replace the coordinate values of the boundary portion of the first region with the primitive surface. Here, the method of determining the first neighboring region is as described in the first embodiment.
[0088] When a primitive surface is a plane perpendicular to the Z axis, only one parameter (ZP where Z=ZP) is required to determine the surface. As an example, ZP can be determined as the average value of the distance values contained in the first neighborhood. The replacement target area in the first area is an area whose distance value is greater than ZP. Here, the average value is just an example, and ZP may be determined as the maximum distance value contained in the first neighborhood. In this case, the replacement target area is smaller than when ZP is set as the average value, but an improvement can be achieved by ensuring that at least a portion of the area is continuously connected to the primitive surface.
[0089] When a primitive surface is part of a sphere centered at the origin, the parameter required to determine the surface is the same (X 2 +Y 2 +Z 2 =RP 2 From equations (1) and (2) and the equation of a sphere, the radius RP of the sphere that passes through the point at pixel coordinates (u, v) can be calculated using equation (4) below.
[0090]
number
[0091] The RP may be calculated for the points included in the first neighborhood using equation (4), and the spherical surface may be determined based on the average value, mode, etc. Alternatively, the spherical surface may be determined so as to minimize the squared error for the points included in the first neighborhood.
[0092] Alternatively, it is possible to use a more general surface shape, but if the surface shape is determined parametrically, it is possible to determine an equation for the primitive surface shape by applying the least squares method or the like to the points contained in the first neighborhood region so that some of the points contained in the first region are replaced.
[0093] <Point cloud formation process S403> In the point cloud generation step S403, the control unit 602 converts the distance distribution information processed in S402 into a point cloud. The specific process content is the same as S103 in FIG. 3 in the first embodiment.
[0094] <Surface processing> Although not shown in Fig. 12, a surface generation process may be performed after the point group generation process. The specific process contents are as described in the first embodiment.
[0095] [Variation 1] In the above embodiment, the process of dividing distance distribution information into regions and replacing distance values has been described. However, the effects of this embodiment can also be obtained with other processes. Below, the contents of the process shown in the flowchart of FIG. 17 will be described as a first modification of the second embodiment.
[0096] S501 is a step of converting distance distribution information into a point cloud, and is the same as the step S403 described above.
[0097] S502 is a step of dividing the point cloud into a first region and a second region. Region division may be performed by, for example, storing reliability information for each point in the point cloud and using threshold processing to define a set of highly reliable coordinate points whose reliability is equal to or greater than a threshold as the first region and a set of low reliability coordinate points whose reliability is less than the threshold as the second region. Alternatively, segmentation may be performed by referring to the point cloud data and using deep learning-based segmentation.
[0098] S503 is a step of replacing the coordinate values of the second region and part of the coordinate values of the first region with primitive faces.
[0099] When a primitive surface is a plane perpendicular to the Z axis, only one parameter (ZP where Z=ZP) is required to determine the surface. As an example, ZP can be determined as the average value of the Z coordinates of the coordinate points included in the first neighborhood region. Then, the Z coordinates of the coordinate points included in the second region and the Z coordinates of the coordinate points included in the first region whose Z coordinates are greater than ZP are replaced with ZP. This achieves the effect of this embodiment.
[0100] [Variation 2] In the explanation so far, the second region has been a single region, but it may be composed of multiple sub-regions, and it may be desirable to replace each of these sub-regions with a different primitive surface. The reasons for this and the processing method will be explained below.
[0101] As an example, consider a scene in which a spherical object is within the ranging area, and a rectangular object and background are outside the ranging area. Here, the rectangular object is located in front of the ranging area, and the background is located behind the ranging area. Figure 18(a) shows actual distance distribution information for this scene. Because the rectangular object and background are outside the ranging area, the distance values for these areas are not correctly obtained. Therefore, the distance values for these areas contain a large amount of error, as shown in Figure 18(b), and the resulting distance distribution information differs from the actual distance distribution information (Figure 18(a)).
[0102] When region segmentation is performed according to the method described in the second embodiment, the low-reliability rectangular object and background are segmented into the same second region, as shown in Fig. 19(a). Therefore, these are replaced with the same primitive surface shape, but when this is converted into 3D data, the rectangular object that is actually in front of the spherical object will be positioned behind the spherical object, which may result in unnatural 3D data.
[0103] 19(b), the rectangular subject is treated as sub-region 1 in the second region, the background is treated as sub-region 2 in the second region, and these are replaced with different primitive surface shapes for each sub-region. When replacing both with planes perpendicular to the Z axis, the parameters required to determine the surface are one each (ZP1, ZP2).
[0104] The plane Z=ZP1 of subregion 1 can be determined based on information about the region of the first neighborhood that is close to subregion 1 of the second region. Then, coordinate values in the first region that are located on the far side of plane Z=ZP1 can be replaced with Z=ZP1. Similarly, the plane Z=ZP2 of subregion 2 can be determined based on information about the region of the first neighborhood that is close to subregion 2 of the second region. Then, coordinate values in the first region that are located on the near side of plane Z=ZP2 can be replaced with Z=ZP2. As a result, three-dimensional data can be obtained that roughly maintains the anteroposterior relationship of the original distance distribution information, reducing the sense of discomfort felt when viewed.
[0105] Region division can be performed, for example, by using distance distribution information together with reliability information. By treating a region with a nearer distance value among low-reliability regions and a region with a farther distance value among low-reliability regions as different regions, it is possible to distinguish between subregions.
[0106] If the distance distribution information in the low-reliability region of the acquired distance distribution information has a large random error, distance distribution information to which a smoothing filter such as a median filter has been applied may be used for region segmentation. Alternatively, region segmentation may be performed using distance distribution information estimated by deep learning based on an image, rather than distance distribution information actually acquired by an acquisition means.
[0107] [Variation 3] In the above explanation, an example was described in which the primitive surface shape is rotationally symmetric about the Z axis, but in some cases, it may be possible to reduce the sense of incongruity by setting the primitive surface shape so that it is rotationally asymmetric about the Z axis. The reason for this and the processing method will be explained below.
[0108] FIG. 20(c) shows a side view of the three-dimensional data in Modification 3, and FIGS. 20(a) and 20(b) show side views of the three-dimensional data in a comparative example.
[0109] FIG. 20 illustrates a case where the spherical object exemplified above is rotated around the X axis. When a plane is set so that a part of the first neighborhood region is replaced with a plane rotationally symmetric about the Z axis (Z=ZP), many parts are separated from the plane, as shown in FIG. 20(a), and the loss of information in the second neighborhood region is easily visible. Furthermore, when a plane is set so that the entire first neighborhood region is replaced with a plane rotationally symmetric about the Z axis (Z=ZP), as shown in FIG. 20(b), the number of replaced coordinate points increases, and the three-dimensional effect is lost. Therefore, by increasing the degree of freedom of the primitive surface and creating a plane like that shown in FIG. 20(c) (a plane rotationally asymmetric about the Z axis), it is possible to generate three-dimensional data in which the first neighborhood region and the primitive surface are continuously connected but the three-dimensional effect is not lost.
[0110] When the primitive surface is a plane that is rotationally asymmetric with respect to the Z axis, a least-squares plane may be generated for the coordinate points of the first neighboring region, or the plane may be generated using a three-dimensional Hough transform.
[0111] Alternatively, it is possible to use a more general surface shape, but if the surface shape is determined parametrically, it is possible to determine an equation for the primitive surface shape by applying the least squares method or the like to the points contained in the first neighborhood region so that some of the points contained in the first region are replaced.
[0112] [Variation 4] In the above explanation, it is assumed that the primitive surface shape is determined based on the coordinate information of the first neighboring region. However, it may also be determined based on the coordinate information of a low-reliability region within the first region. As mentioned above, the coordinate points included in the first region are relatively more reliable than the coordinate points included in the second region, but the first region may include low-reliability coordinate points. Such low-reliability coordinate points may contain more errors than highly reliable coordinate points, and if such coordinate points are converted into 3D data as is, it may cause a strange feeling when viewed.
[0113] Therefore, if primitive surfaces are set so that low-reliability coordinate points in the first region are replaced with primitive surfaces, three-dimensional data can be constructed using only highly reliable coordinate points, thereby reducing the sense of incongruity felt when viewed. A specific method involves, for example, determining a temporary primitive surface by applying the least-squares method or the like to the coordinate information of the low-reliability region in the first region. Then, the temporary primitive surface may be offset in the Z-axis direction so that all of the coordinate points of the low-reliability region in the first region are replaced with primitive surfaces, thereby determining a formal primitive surface.
[0114] [Variation 5] In generating the three-dimensional data of this embodiment, motion data may also be generated that continuously transitions between a first state in which all of the coordinate information of the first region is replaced with a primitive surface shape, and a second state in which part of the coordinate information of the first region is replaced with a primitive surface shape. Here, motion data refers to information sufficient to generate an animation that continuously transitions between the three-dimensional data in the first state and the three-dimensional data in the second state. For example, the motion data may be in the form of information on the time-dependent changes in the coordinate information of each vertex in the three-dimensional data, or information on the time-dependent changes in the scale of the three-dimensional data in the first region along the Z-axis direction. Hereinafter, the state of the three-dimensional data between the first state and the second state will be referred to as a transition state.
[0115] Using the scene in FIG. 14 as an example, the motion data of this modified example will be described with reference to FIG. 21. FIG. 21(a) is a diagram showing a first state in which all of the coordinate information of the first region has been replaced with a primitive surface shape. In this example, the primitive surface is a plane perpendicular to the Z axis. FIG. 21(c) is a diagram showing a second state in which part of the coordinate information of the first region has been replaced. The second state refers to the three-dimensional data of this embodiment. FIG. 21(b) is a diagram showing three-dimensional data in a transition state between the first and second states. This is an intermediate state in which the three-dimensional shape of the first region transitions from a plane to the three-dimensional shape of this embodiment, and is compressed in the Z axis direction.
[0116] The first state is a complete primitive surface, and when color distribution information is added to this 3D data, it gives the viewer the impression of a photograph. The second state is a state in which the subject area has unevenness, and by continuously transitioning from the first state to the second state, it is possible to give the viewer the impression that the subject is popping out of the photograph. In this way, the effect of this modified example is that it can provide a new visual experience in which the subject pops out of the image.
[0117] [Variation 6] In the above description, it is assumed that no operations other than replacing a portion of the coordinate information of the first region with a primitive surface are performed. However, the present invention is not limited to this. For example, by enlarging the coordinate information of the first region that is not replaced with a primitive surface in the XY direction, it is possible to make the loss of information in the second region less visible, thereby further enhancing the effects of the present invention.
[0118] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0119] The disclosure of this specification includes the following image processing device, image processing method, program, and storage medium.
[0120] (Item 1) a dividing means for dividing the image data into a first region and a second region; a conversion means for converting at least one of first coordinate information, which is coordinate information of the subject in the first region, and second coordinate information, which is coordinate information of the subject in the second region, so that a difference between a first representative value representing the first coordinate information and a second representative value representing the second coordinate information becomes small, based on the first coordinate information and the second coordinate information; An image processing device comprising:
[0121] (Item 2) further comprising a calculation means for calculating an offset value to be added to the first coordinate information so as to reduce a difference between the first representative value and the second representative value; 2. The image processing device according to item 1, wherein the conversion means converts the first coordinate information by adding the offset value to the first coordinate information.
[0122] (Item 3) 3. The image processing device according to item 2, wherein the dividing means divides the image data so that the first region includes an occluding region that occludes a part of the subject in the second region.
[0123] (Item 4) 4. The image processing device according to item 2 or 3, characterized in that the first region includes a plurality of sub-regions that are not connected to each other, and the calculation means calculates a different offset value for each of the plurality of sub-regions.
[0124] (Item 5) 5. The image processing device according to any one of items 2 to 4, further comprising a point cloud conversion means for converting the coordinate information of the subject into a point cloud after the conversion means adds the offset value to the first coordinate information.
[0125] (Item 6) Item 6. The image processing device according to item 5, wherein the first coordinate information is Z1, the coordinate information obtained by adding the offset value to Z1 is Z1', and the coordinate information used when calculating coordinate values in a direction perpendicular to the direction of the distance to the subject in the first area in the point cloud generation is Z1'', the difference between Z1 and Z1'' is smaller than the difference between Z1 and Z1'.
[0126] (Item 7) 7. The image processing device according to any one of items 1 to 6, wherein the first representative value is a statistic of coordinate values in a first neighboring region located near the boundary between the first region and the second region, and the second representative value is a statistic of coordinate values in a second neighboring region located near the boundary between the second region and the first region.
[0127] (Item 8) 8. The image processing device according to item 7, wherein the first representative value is one of the average value, median value, and mode value of the coordinate values in the first vicinity area, and the second representative value is one of the average value, median value, and mode value of the coordinate values in the second vicinity area.
[0128] (Item 9) 9. The image processing device according to any one of items 1 to 8, further comprising a filtering unit that performs a filtering process on coordinate information of an area of the subject, the area including a first neighboring area located near the boundary between the first area and the second area, and a second neighboring area located near the boundary between the second area and the first area.
[0129] (Item 10) a dividing means for dividing the image data into a first region and a second region; a conversion means for converting second coordinate information, which is coordinate information of the object in the second region, and a part of first coordinate information, which is coordinate information of the object in the first region, into coordinate information corresponding to a primitive surface shape; The image processing device is characterized in that the conversion means converts either the coordinate values of the second coordinate information and the first coordinate information that are larger or smaller than the coordinate values of the primitive surface shape into coordinate information corresponding to the primitive surface shape.
[0130] (Item 11) Item 11. The image processing device according to item 10, wherein the reliability of the second coordinate information is lower than the reliability of the first coordinate information.
[0131] (Item 12) 12. The image processing device according to item 10 or 11, wherein the conversion means determines the primitive surface shape based on coordinate values in a first neighboring region located in the first region near the boundary with the second region.
[0132] (Item 13) 13. The image processing device according to any one of items 10 to 12, wherein the conversion means determines coordinate values of the primitive surface shape based on the reliability of the coordinates of the subject in the first region.
[0133] (Item 14) The image processing device described in any one of items 10 to 13, characterized in that the second area consists of a plurality of sub-areas, and the conversion means converts coordinate information in each of the plurality of sub-areas into primitive surface shapes that are different from each other.
[0134] (Item 15) 15. The image processing device according to any one of items 10 to 14, wherein the primitive surface shape is rotationally asymmetric with respect to an axis in the direction of the distance to the subject.
[0135] (Item 16) 16. The image processing device according to any one of items 10 to 15, further comprising means for generating motion data that continuously transitions between a first state in which all of the first coordinate information is converted into the primitive surface shape and a second state in which a portion of the first coordinate information is converted into the primitive surface shape.
[0136] (Item 17) A dividing step of dividing the image data into a first region and a second region; a conversion step of converting at least one of first coordinate information, which is coordinate information of the subject in the first region, and second coordinate information, which is coordinate information of the subject in the second region, so that a difference between a first representative value representing the first coordinate information and a second representative value representing the second coordinate information becomes small, based on the first coordinate information and the second coordinate information; An image processing method comprising:
[0137] (Item 18) A dividing step of dividing the image data into a first region and a second region; a conversion step of converting second coordinate information, which is coordinate information of the object in the second region, and a part of first coordinate information, which is coordinate information of the object in the first region, into coordinate information corresponding to a primitive surface shape; An image processing method characterized in that in the conversion step, either the coordinate values of the second coordinate information and the first coordinate information that are larger or smaller than the coordinate values of the primitive surface shape are converted into coordinate information corresponding to the primitive surface shape.
[0138] (Item 19) A program for causing a computer to function as each of the means of the image processing device according to any one of items 1 to 16.
[0139] (Item 20) A computer-readable storage medium storing a program for causing a computer to function as each of the means of the image processing device according to any one of items 1 to 16.
[0140] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0141] 600: Image processing device, 601: Image capture unit, 602: Control unit, 603: Memory, 604: Storage device, 605: Image processing unit
Claims
1. a dividing means for dividing the image data into a first region and a second region; a conversion means for converting, based on first coordinate information that is coordinate information of the subject in the first region and second coordinate information that is coordinate information of the subject in the second region, at least one of the first coordinate information and the second coordinate information so that a difference between a first representative value that represents the first coordinate information and a second representative value that represents the second coordinate information becomes small; An image processing device comprising:
2. further comprising a calculation means for calculating an offset value to be added to the first coordinate information so as to reduce a difference between the first representative value and the second representative value; 2. The image processing apparatus according to claim 1, wherein said conversion means converts said first coordinate information by adding said offset value to said first coordinate information.
3. 3. The image processing apparatus according to claim 2, wherein the dividing means divides the image data so that the first region includes a masked region that masks a part of the subject in the second region.
4. 3. The image processing device according to claim 2, wherein the first region includes a plurality of sub-regions that are not connected to one another, and the calculation means calculates a different offset value for each of the plurality of sub-regions.
5. 3. The image processing apparatus according to claim 2, further comprising a point cloud generating unit that generates a point cloud of the coordinate information of the subject after the conversion unit adds the offset value to the first coordinate information.
6. The image processing device according to claim 5, characterized in that the difference between Z1 and Z1" is smaller than the difference between Z1 and Z1', where Z1 is the first coordinate information, Z1' is the coordinate information obtained by adding the offset value to Z1, and Z1" is the coordinate information used when calculating coordinate values in a direction perpendicular to the direction of the distance to the subject in the first area in the point cloud generation.
7. 2. The image processing device according to claim 1, wherein the first representative value is a statistic of coordinate values in a first neighboring region located near the boundary between the first region and the second region, and the second representative value is a statistic of coordinate values in a second neighboring region located near the boundary between the second region and the first region.
8. 8. The image processing device according to claim 7, wherein the first representative value is one of the average value, median value, and mode value of coordinate values in the first neighboring region, and the second representative value is one of the average value, median value, and mode value of coordinate values in the second neighboring region.
9. The image processing device according to claim 1, further comprising a filtering means for performing a filtering process on coordinate information of the subject, the coordinate information of which includes a first nearby region located near the boundary between the first region and the second region and a second nearby region located near the boundary between the second region and the first region.
10. a dividing means for dividing the image data into a first region and a second region; a conversion means for converting second coordinate information, which is coordinate information of the object in the second region, and a part of the first coordinate information, which is coordinate information of the object in the first region, into coordinate information corresponding to a primitive surface shape; The image processing device is characterized in that the conversion means converts either the coordinate values of the second coordinate information and the first coordinate information that are larger or smaller than the coordinate values of the primitive surface shape into coordinate information corresponding to the primitive surface shape.
11. 11. The image processing apparatus according to claim 10, wherein the reliability of the second coordinate information is lower than the reliability of the first coordinate information.
12. 11. The image processing apparatus according to claim 10, wherein the conversion means determines the primitive surface shape based on coordinate values in a first neighboring region located in the first region near the boundary with the second region.
13. 11. The image processing apparatus according to claim 10, wherein the conversion means determines the coordinate values of the primitive surface shape based on the reliability of the coordinates of the subject in the first region.
14. 11. The image processing device according to claim 10, wherein the second region is made up of a plurality of sub-regions, and the conversion means converts the coordinate information in each of the plurality of sub-regions into a primitive surface shape that is different from each other.
15. 11. The image processing device according to claim 10, wherein the primitive surface shape is rotationally asymmetric with respect to an axis in the direction of the distance to the subject.
16. 11. The image processing device according to claim 10, further comprising means for generating motion data that continuously transitions between a first state in which all of the first coordinate information is converted into the primitive surface shape and a second state in which only a portion of the first coordinate information is converted into the primitive surface shape.
17. A dividing step of dividing the image data into a first region and a second region; a conversion step of converting at least one of first coordinate information, which is coordinate information of the subject in the first region, and second coordinate information, which is coordinate information of the subject in the second region, so that a difference between a first representative value representing the first coordinate information and a second representative value representing the second coordinate information becomes small, based on the first coordinate information and the second coordinate information; An image processing method comprising:
18. A dividing step of dividing the image data into a first region and a second region; a conversion step of converting second coordinate information, which is coordinate information of the object in the second region, and a part of first coordinate information, which is coordinate information of the object in the first region, into coordinate information corresponding to a primitive surface shape; An image processing method characterized in that, in the conversion step, either the coordinate values of the second coordinate information and the first coordinate information that are larger or smaller than the coordinate values of the primitive surface shape are converted into coordinate information corresponding to the primitive surface shape.
19. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 16.
20. 17. A computer-readable storage medium storing a program for causing a computer to function as each of the means of the image processing apparatus according to claim 1.