Image processing apparatus, image processing method, and program
Patent Information
- Application Number
- JP2023161507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-06-22
- Filing Date
- 2023-09-25
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2043-09-25
AI Technical Summary
In three-dimensional computer graphics, the positional relationship between polygons and camera viewpoints can result in fragments with collapsed shapes, leading to incomplete color information and low-quality texture images due to fragments not being fully captured by camera viewpoints.
An image processing device that acquires three-dimensional shape data, divides it into groups, associates specific photographing viewpoints, generates a two-dimensional map, and creates a texture image using both real and virtual camera viewpoints to ensure complete color information across all fragments.
This approach enables the generation of high-quality texture images by ensuring all fragments are colored, even when real camera viewpoints are not optimal, by utilizing virtual cameras to fill in missing color information.
Smart Images

Figure 00000026_0000 
Figure 00000027_0000 
Figure 00000027_0001
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a technique for generating a texture image used for coloring three-dimensional shape data of an object. [Background technology]
[0002] In recent years, three-dimensional computer graphics have been developed. In three-dimensional computer graphics, the shape of an object is represented by mesh data (generally called a "3D model") composed of multiple polygons. The color, texture, etc. of the object are represented by attaching a texture image to the surface of each polygon constituting the 3D model. The texture image can be generated by dividing the 3D model into multiple groups, and assigning color information to a UV map in which two-dimensional pieces (called "fragments") generated by UV unwrapping for each group are arranged. For example, Patent Document 1 discloses a method in which one camera viewpoint is selected from multiple camera viewpoints for an object for each polygon, and UV unwrapped for each group. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2021-64334 A Summary of the Invention [Problem to be solved by the invention]
[0004] When the camera is only located near the in-plane direction of the polygons that make up the 3D model, the shape of the fragment generated by UV unwrapping becomes a squashed shape that is close to a straight line, and color information cannot be added. Here, a specific example will be used to explain this. FIG. 1(a) illustrates a camera viewpoint 1103 that is placed at a position where two polygons 1101 and 1102 that make up the mesh data of a human face are captured. FIG. 1(b) illustrates the positional relationship between these two polygons 1101 and 1102 and the camera viewpoint 1103. Now, the line of sight of the camera viewpoint 1103 and the polygon 1101 are in a nearly perpendicular relationship, while the line of sight of the camera viewpoint 1103 and the polygon 1102 are in a nearly parallel relationship. Here, when the polygon 1101 is projected at the camera viewpoint 1103 to generate a fragment 1106, several pixel centers indicated by black circles 1108 are included inside. Therefore, color information can be added to the fragment 1106 using the pixel values of each of these pixels in the image captured by the camera viewpoint 1103. In contrast, when polygon 1102 is projected at camera viewpoint 1103 to generate fragment 1107, the center of any pixel in the image captured at camera viewpoint 1103 is not included within it. As a result, color information cannot be added to fragment 1107. Thus, depending on the positional relationship between the elements constituting the 3D model (polygons in the above example) and the camera viewpoint, it is not possible to color the fragment generated by UV unwrapping, resulting in a low-quality texture image that includes missing colors. [Means for solving the problem]
[0005] The image processing device according to the present disclosure comprises an acquisition means for acquiring three-dimensional shape data of an object captured in a plurality of captured images captured from different viewpoints; a division means for dividing elements constituting the three-dimensional shape data into a plurality of groups; a control means for associating viewpoint information representing a specific capturing viewpoint with each of the plurality of groups; a first generation means for generating a two-dimensional map based on the plurality of groups and the viewpoint information representing the specific capturing viewpoint associated with each group; and a second generation means for generating a texture image representing the color of the object based on the two-dimensional map, wherein the specific capturing viewpoint includes a virtual capturing viewpoint different from the capturing viewpoint of each of the plurality of captured images. Effect of the Invention
[0006] According to the present disclosure, a high-quality texture image can be obtained. [Brief description of the drawings]
[0007] [Figure 1] 1A is a diagram showing a camera viewpoint placed at a position where the polygons that make up the mesh data are reflected, and FIG. 1B is a diagram showing the positional relationship between the polygons and the camera viewpoint. [Diagram 2] FIG. 4A is a block diagram showing an example of a hardware configuration of an image processing apparatus, and FIG. 4B is a block diagram showing an example of a software configuration of the image processing apparatus. [Diagram 3] 5 is a flowchart showing an outline of the operation of the image processing device according to the first embodiment. [Figure 4] 5 is a flowchart showing details of a mesh division process according to the first embodiment. [Diagram 5] 10 is a flowchart showing details of a camera selection process. [Figure 6] 5 is a flowchart showing details of a camera parameter setting process according to the first embodiment. [Figure 7] 5 is a flowchart showing details of a texture image generation process according to the first embodiment. [Figure 8](a) shows mesh data and the real cameras corresponding to the multiple captured images used to generate it; (b) shows the mesh data divided into groups; (c) shows fragments corresponding to the groups; and (d) shows the fragments placed on a two-dimensional map. [Figure 9] 11A and 11B are diagrams showing how camera parameters of a virtual camera are set for a group; [Figure 10] 5A to 5C are diagrams for explaining a process of generating a texture image. [Figure 11] 10 is a flowchart showing an outline of the operation of an image processing device according to a second embodiment. [Figure 12] 10 is a flowchart showing details of a mesh division process according to the second embodiment. [Figure 13] 10 is a flowchart showing details of a camera parameter setting process according to the second embodiment. [Figure 14] FIG. 4A is a diagram showing mesh data, and FIG. 4B is a diagram showing the mesh data divided into groups. [Figure 15] 10 is a flowchart showing details of a texture image generation process according to the second embodiment. [Figure 16] 10 is a flowchart showing details of a camera selection process. [Figure 17] 11 is a flowchart showing an outline of the operation of an image processing device according to a third embodiment. [Figure 18] FIG. 1 is a diagram showing an example of the arrangement of virtual cameras. [Figure 19] 11 is a flowchart showing details of a mesh division process according to a third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following embodiments do not limit the present disclosure, and not all of the combinations of features described in the present embodiments are necessarily essential to the solution of the present disclosure. Note that the same components will be described with the same reference numerals.
[0009] [Embodiment 1] First, the hardware configuration and software configuration of an image processing apparatus according to the present embodiment will be described with reference to the drawings.
[0010] <Hardware configuration> 2A is a block diagram showing an example of a hardware configuration of the image processing device 100 of this embodiment. The image processing device 100 includes a CPU 101, a ROM 102, a RAM 103, an auxiliary storage device 104, a display unit 105, an operation unit 106, a communication unit 107, and a bus 108.
[0011] The CPU 101 is an arithmetic processing device that performs overall control of the image processing device by executing a program stored in the ROM 102 or the RAM 103. The image processing device 100 further includes one or more dedicated hardware components different from the CPU 101, and the dedicated hardware components may execute at least a part of the processing performed by the CPU 101 in place of or in cooperation with the CPU 101. Examples of the dedicated hardware components include an ASIC, an FPGA, and a DSP (digital signal processor). The ROM 102 is a memory that stores programs that do not require modification. The RAM 103 is a memory that temporarily stores programs or data supplied from the auxiliary storage device 104, or data supplied from the outside via the communication unit 107. The auxiliary storage device 104 is, for example, a hard disk drive, and stores various data such as image data or audio data.
[0012] The display unit 105 is configured, for example, by a liquid crystal display or an LED, and displays a GUI (Graphical User Interface) or the like for the user to operate the image processing device 100 or to view the state of processing in the image processing device 100. The operation unit 106 is configured, for example, by a keyboard, a mouse, a joystick, a touch panel, and the like, and inputs various instructions to the CPU 101 in response to operations by the user. The CPU 101 also operates as a display control unit that controls the display unit 105, and an operation control unit that controls the operation unit 106.
[0013] The communication unit 107 performs communication such as sending and receiving data between the image processing device 100 and an external device. For example, when the image processing device 100 is connected to an external device by wire, a communication cable is connected to the communication unit 107. When the image processing device 100 has a function of wirelessly communicating with an external device, the communication unit 107 is equipped with an antenna. The bus 108 connects each unit of the image processing device 100 to transmit information. In this embodiment, the display unit 105 and the operation unit 106 are described as being present inside the image processing device 100, but at least one of the display unit 105 and the operation unit 106 may be present as a separate device outside the image processing device 100.
[0014] <Software configuration> 2(b) is a block diagram showing an example of a software configuration (functional configuration) of the image processing device 100 of this embodiment. The image processing device 100 includes a data acquisition unit 201, a mesh division unit 202, a camera parameter setting unit 203, a UV development unit 204, and a texture image generation unit 205. Each of these functional units is realized by the CPU 101 executing a predetermined program, or by dedicated hardware such as an ASIC or FPGA.
[0015] The data acquisition unit 201 acquires a 3D model (three-dimensional shape data) that represents the three-dimensional shape of an object. The data format of the 3D model includes mesh data that represents the surface shape of an object as a set of polygons, volume data that represents the three-dimensional shape of an object as voxels, and point cloud data that represents a set of points. In this embodiment, mesh data that represents the surface shape of an object as a set of triangular polygons is acquired as a 3D model. Note that the shape of the polygon is not limited to a triangle, and may be other polygons such as a quadrangle or a pentagon. In addition, the data acquisition unit 201 acquires viewpoint information related to the shooting viewpoints of a plurality of imaging devices (hereinafter referred to as "real cameras") arranged in the shooting space. This viewpoint information is generally called "camera parameters" and includes internal parameters such as focal length and image center, and external parameters that represent the position and attitude of the camera.
[0016] The mesh division unit 202 performs processing to group elements that make up the acquired 3D model (polygons that make up the mesh data in this embodiment) into groups that correspond to the shooting viewpoints of specific real cameras.
[0017] The camera parameter setting unit 203 sets camera parameters corresponding to an image capture device having an appropriate camera viewpoint (shooting viewpoint) for each group obtained by the mesh division unit 202. At this time, if there is no camera having an appropriate camera viewpoint among the multiple real cameras arranged, the camera parameter setting unit 203 sets camera parameters corresponding to a non-existent virtual image capture device having the appropriate camera viewpoint (hereinafter referred to as a "virtual camera").
[0018] The UV unfolding unit 204 unfolds the mesh data, which is three-dimensional information, onto a two-dimensional plane, and generates a UV map that corresponds to two-dimensional coordinates (abscissa U and ordinate V). Specifically, first, each polygon included in the divided groups is unfolded onto a two-dimensional plane, and two-dimensional pieces (fragments) corresponding to the groups are generated. Then, the generated fragments for each group are placed on a two-dimensional map (UV map) that corresponds to the mesh data.
[0019] The texture image generation unit 206 calculates the pixel values of each pixel constituting a UV map based on a plurality of captured images obtained by a plurality of real cameras, and generates a texture image.
[0020] <Operation flow of image processing device> Next, the operation of the image processing device 100 according to this embodiment will be described with reference to the flowcharts shown in Figures 3 to 7. Figure 3 is a flowchart of a main routine showing a rough operational flow of the image processing device 100, and the remaining Figures 4 to 7 are flowcharts of subroutines corresponding to specific processes within the main routine. Below, the operational flow of the image processing device 100 will be described along with each flowchart. In the following description, the symbol "S" means a step.
[0021] In S301, the data acquisition unit 201 acquires mesh data of an object. In this case, it is assumed that mesh data is acquired by applying the marching cubes method to voxel data generated by applying the visual volume intersection method. Fig. 8(a) shows the acquired mesh data 801 and three real cameras 802, 803, and 804 corresponding to the multiple captured images used for generating the mesh data. Note that the mesh data may be acquired by generating the voxel data within the image processing device 100, or may be acquired by receiving mesh data generated by an external device (not shown).
[0022] In S302, the data acquisition unit 201 acquires the plurality of captured images and the camera parameters of the plurality of real cameras corresponding to each captured image.
[0023] In S303, the mesh division unit 202 performs a process of dividing the mesh data acquired in S301 based on the camera parameters of the real camera acquired in S302. By this process (hereinafter referred to as "mesh division process"), the input mesh data is divided into a plurality of groups corresponding to the respective camera viewpoints. Fig. 8(b) shows a state in which the mesh data 801 shown in Fig. 8(a) has been divided into three groups 805, 806, and 807. In Fig. 8(b), the areas surrounded by thick lines indicate the respective groups, and it can be seen that each group is made up of a plurality of polygons. The mesh division process will be described in detail later.
[0024] In S304, the camera parameter setting unit 203 sets the camera parameters of a camera having a camera viewpoint suitable for use in the coloring process described later for each group divided in S303. If there is no real camera that corresponds to the "camera having the suitable camera viewpoint", a virtual camera that is not actually placed is created. In the example of FIG. 8(b), the real camera 802 is set to the group 806, the real camera 804 is set to the group 807, and the virtual camera 808 is set to the group 805 as the "camera corresponding to the suitable camera viewpoint". In this case, the camera parameters of the virtual camera 808 are set to the group 805, the camera parameters of the real camera 802 are set to the group 806, and the camera parameters of the real camera 804 are set to the group 807. The camera parameter setting process will be described in detail later.
[0025] In S305, the UV unwrapping unit 204 performs UV unwrapping for each divided group. Specifically, based on the camera parameters set for each group, a fragment is generated by projecting from the camera viewpoint represented by the camera parameters, and a process of arranging the fragment on a two-dimensional map is performed. In this way, a two-dimensional map (UV map) corresponding to the mesh data on which the fragments for each group are arranged is obtained. FIG. 8(c) shows fragments 809-811 corresponding to the three groups 805-807 shown in FIG. 8(b), and FIG. 8(d) shows the state in which the fragments 809-811 are arranged on a two-dimensional map corresponding to the entire mesh data.
[0026] In S306, the texture image generating unit 205 calculates the pixel values of the pixels included in the fragment based on the two-dimensional map generated in S305, using a captured image showing the group corresponding to the fragment, to generate a texture image. Details of the texture image generation process will be described later.
[0027] The above is a rough outline of the operation flow in the image processing device 100. Note that after the process of S306, a process of adding color information to the surfaces of polygons of the mesh data corresponding to each pixel of the texture image may be further performed.
[0028] <Details of mesh division process> Next, the mesh division process (S303) according to this embodiment will be described in detail with reference to the flowchart in FIG.
[0029] In S401, a process (camera selection process) is performed to select a real camera that best captures each polygon (hereinafter referred to as a "best camera"), which is an element that constitutes the mesh data to be processed. Details of this camera selection process will be described later, but it is also determined whether the shooting viewpoint of the real camera selected for each polygon is suitable as a camera viewpoint to be used in grouping, which will be described later, and a correspondence table including the results of the viewpoint determination is created and saved. In the example of FIG. 8(a) described above, the following correspondence table is created. Note that in the correspondence table below, the reference symbols of the real cameras in FIG. 8(a) are used as they are as the identifiers of the best cameras (best camera IDs).
[0030] [Table 1] In S402, an identifier (group ID) for managing the group to be created is initialized (here, a process of setting an initial value=0).
[0031] In S403, one polygon that does not yet belong to any group (not associated with any group ID) is selected as a polygon of interest from among the polygons that make up the mesh data.
[0032] In S404, the result of the camera selection process in S401 (here, the correspondence table described above) is referenced, and the process to be executed next is allocated depending on whether the viewpoint judgment result of the real camera of the best camera ID corresponding to the polygon ID of the polygon of interest is "OK" or "NG." If the flag value of the polygon of interest is now "1" indicating OK, the process of S405 is executed next, and if the flag value is "0" indicating NG, the process of S408 is executed next.
[0033] In S405, adjacent polygons that satisfy a certain condition with the polygon of interest are grouped together with the polygon of interest. Specifically, first, a real camera associated as the best camera of the polygon of interest is specified. Then, among the polygons to which the same real camera as the specified real camera is associated as the best camera, all polygons whose viewpoint determination result is "OK" and which are adjacent to the polygon of interest are extracted. Then, a group consisting of the polygon of interest and all the extracted adjacent polygons is created. Here, the condition for adjacent polygons is that the polygons share one or more of the same vertices.
[0034] In S406, a group ID for identifying the group is set for the polygon of interest and its adjacent polygons that belong to the group created in S405. Accordingly, a new column for group IDs is added to the above-mentioned correspondence table, and the correspondence table is updated. In the example of FIG. 8(a) described above, the correspondence table is updated, for example, as follows:
[0035] [Table 2] In S407, for the group whose group ID was set in S406, the real camera that is associated as the best camera with the polygon of interest and its adjacent polygons belonging to the group is set as the real camera (valid camera) having an appropriate viewpoint for the group. Accordingly, the above-mentioned correspondence table is updated, a new column of "valid camera ID" that is an identifier of the valid camera is added, and information that identifies the real camera as a valid camera is entered there. In the example of FIG. 8(a) above, the correspondence table is updated as follows. Note that in the following correspondence table, the reference symbols of the real cameras in FIG. 8(a) are used as valid camera IDs as they are.
[0036] [Table 3] In S408, adjacent polygons that satisfy a certain condition with the polygon of interest are grouped together with the polygon of interest. Specifically, first, a real camera associated as the best camera of the polygon of interest is specified. Then, among the polygons to which the same real camera as the specified real camera is associated as the best camera, all polygons whose viewpoint determination result is "NG" and which are adjacent to the polygon of interest are extracted. Then, a group consisting of the polygon of interest and all the extracted adjacent polygons is created. Here, the condition for adjacent polygons is that the polygons share one or more of the same vertices, as in the above-mentioned S405.
[0037] In S409, a group ID for identifying the group is set for the polygon of interest and its adjacent polygons that belong to the group created in S408. Accordingly, the correspondence table described above is updated and a new column for group IDs is added. In the example of FIG. 8(a) described above, the correspondence table is updated as follows, for example.
[0038] [Table 4] In S410, it is determined whether the above-mentioned processing has been completed for all polygons constituting the mesh data. If the processing has been completed for all polygons, the above-mentioned correspondence table at the time of completion is saved in RAM 102 as a result of the mesh division processing, and then the flow ends. On the other hand, if there are any unprocessed polygons, the group ID number is incremented (+1) in S411, and the process returns to S403 to repeat the same processing.
[0039] The above is the content of the mesh division process (S303). Next, the camera selection process (S401) according to this embodiment will be described in detail with reference to the flowchart of FIG.
[0040] <Details of camera selection process> In S501, one polygon of interest is selected from among the polygons that make up the mesh data. In the next step S502, one real camera of interest is selected from among the multiple real cameras.
[0041] In S503, the angle between the normal vector of the polygon of interest selected in S501 and the direction vector of the real camera of interest selected in S502 is calculated. Here, the normal vector nv of the triangular polygon can be obtained by the cross product of the side vectors (v1, v2) based on the vertices (p0, p1, p2) of the polygon, and is represented by the following formula (1).
[0042]
number
[0043]
number
[0044]
number
[0045] The direction vector cv of the real camera is expressed by the position coordinates p cam and the center of gravity of the polygon p tri and is expressed by the following equation (4).
[0046]
number
[0047]
number
[0048]
number
[0049] In S505, the real camera with the smallest calculated "angle" among all the placed real cameras is associated with the polygon of interest as the best camera. As a result, the polygon ID of the polygon of interest and the best camera ID indicating the real camera as the best camera are associated and saved, for example, as in the above-mentioned correspondence table.
[0050] In S506, the process to be executed next is determined based on whether the angle of the "angle" determined to be the smallest in S505 (the minimum angle) is equal to or smaller than a preset threshold. The threshold is preset based on experience and is preferably around 70 degrees, for example. If the minimum angle is equal to or smaller than the threshold, the process of S507 is executed next, and if it is larger than the threshold, the process of S508 is executed next.
[0051] In S507, information indicating that the viewpoint of the real camera determined to be the best camera for the polygon of interest is an appropriate camera viewpoint is set for the polygon of interest. Specifically, in the aforementioned correspondence table, a flag value "1" indicating that the camera viewpoint is appropriate is set in the "viewpoint determination result" column in the record of the polygon ID of the polygon of interest.
[0052] In S508, information indicating that the viewpoint of the real camera determined to be the best camera for the polygon of interest is not an appropriate camera viewpoint is set for the polygon of interest. Specifically, in the correspondence table described above, a flag value "0" indicating that the viewpoint is not an appropriate camera viewpoint is set in the record of the polygon ID of the polygon of interest.
[0053] In S509, it is determined whether the above-mentioned processing has been completed for all polygons constituting the mesh data. If the processing has been completed for all polygons, the above-mentioned correspondence table at the time of completion is saved in RAM 102 as the result of the camera selection processing, and then the flow ends. On the other hand, if there are any unprocessed polygons, the flow returns to S501 and the same processing is repeated.
[0054] The above is the content of the camera selection process (S401).
[0055] <Details of camera parameter setting process> Next, the camera parameter setting process (S304) according to this embodiment will be described in detail with reference to the flowchart in FIG.
[0056] In S601, one group of interest is selected from all the groups obtained by the mesh division process (S303) described above.
[0057] In S602, the above-mentioned correspondence table is referenced, and the process to be executed next is determined based on the setting value of the "valid camera ID" column for the selected group of interest. Specifically, if the camera ID of a real camera is set in the "valid camera ID" column in the record of the group of interest, the process of S603 is executed next, and if nothing is set and the column is blank, the process of S604 is executed next.
[0058] In S603, the above-mentioned correspondence table is referenced, and the camera parameters of the real camera specified by the camera ID listed in the "effective camera ID" column are set for all polygons belonging to the target group.
[0059] In S604, a normal vector is calculated for each polygon belonging to the group of interest. The method of calculating the normal vector is as described in S503 above. In the following S605, an average vector of the normal vectors calculated for all polygons belonging to the group of interest is calculated. Then, in S606, camera parameters corresponding to a virtual camera having a camera viewpoint with a line of sight in the opposite direction along the direction of the average vector calculated in S605 are set. At this time, parameter values such as the distance from the virtual camera to the object and the resolution of the virtual camera may be matched to those of the surrounding real cameras that are actually placed. FIG. 9 is a diagram for explaining a state in which the camera parameters of the virtual camera are set for the group of interest. Now, the group of interest 901 is composed of four triangular polygons 902, and a normal vector 903 indicated by a thin arrow is calculated for each triangular polygon 902, and an average vector 904 indicated by a thick arrow is calculated based on these four normal vectors 903. In this case, camera parameters corresponding to a virtual camera 905 having a camera viewpoint with a line of sight in the opposite direction along the direction of the average vector 904 are set.
[0060] In S607, it is determined whether the above-mentioned processing has been completed for all groups obtained by the mesh division processing. If the processing has been completed for all groups, the flow ends. On the other hand, if there are any unprocessed groups, the flow returns to S601 and the same processing is repeated.
[0061] The above is the content of the camera parameter setting process (S304).
[0062] <Details of texture image generation process> Next, the texture image generation process (S306) according to this embodiment will be described in detail with reference to the flowchart in Fig. 7. At the time of execution of this step, the above-mentioned correspondence table has been updated by the preceding UV unwrapping (S305) as shown in Table 5 below. That is, the same setting value as the corresponding group ID is stored as the identifier (fragment ID) of the fragment generated for each group.
[0063] [Table 5] In S701, a fragment of interest is selected from among the fragments in the UV map. Figure 10 is a diagram for explaining the process of generating a texture image, and it is assumed that a fragment 1011 in the UV map is selected as the fragment of interest.
[0064] In S702, the above-mentioned correspondence table is referenced and a group corresponding to the fragment of interest is obtained. In the example of Fig. 10 above, group 1001 is obtained as the group corresponding to fragment 1011.
[0065] In S703, a triangle of interest is first selected from among the triangles included in the fragment of interest. Then, by referring to the correspondence table described above, a polygon corresponding to the triangle of interest is obtained from the group of polygons corresponding to the group obtained in S702. In the example of Fig. 10 described above, a triangle 1012 is selected as the triangle of interest from among the multiple triangles included in the fragment 1011, and a polygon 1002 corresponding to the triangle 1012 is obtained from the group of polygons.
[0066] In S704, the two-dimensional coordinates of the pixel on the UV map that is inside the target triangle selected in S703 are obtained. In the example of Fig. 10 described above, the two-dimensional coordinates (x, y) of the point p that represents the pixel center of the pixel 1013 included inside the triangle 1012 are obtained.
[0067] In S705, three-dimensional coordinates on the polygon acquired in S703 corresponding to all two-dimensional coordinates acquired in S704 are calculated. In the example of Fig. 10 described above, three-dimensional coordinates (X, Y, Z) of point P on polygon 1002 corresponding to two-dimensional coordinates (x, y) of point p representing the pixel center of acquired pixel 1013 are calculated. In this calculation, first, the barycentric coordinates of pixel 1013 are calculated based on the three vertices pa, pb, and pc of triangle 1012. Then, the three-dimensional coordinates (X, Y, Z) on polygon 1002 can be obtained by applying the barycentric coordinates of pixel 1013 thus obtained to the three vertices PA, PB, and PC of polygon 1002.
[0068] <Details of calculation method>.
[0069] First, the center of gravity coordinates of pixel 1013 are calculated using the following equation (7) based on the area of the three triangles formed by connecting the two-dimensional coordinates (x, y) of point p representing the pixel center of pixel 1013 and the three vertices pa, pb, and pc of triangle 1012.
[0070]
number
[0071]
number
[0072]
number
[0073] In S706, the above-mentioned correspondence table is referenced, and two-dimensional coordinates on the captured image corresponding to the three-dimensional coordinates calculated in S705 are obtained. In the example of Fig. 10 described above, the real camera 1027 indicated by the best camera ID associated with the fragment ID of the target fragment is first identified. Then, the two-dimensional coordinates of the point pi when the three-dimensional coordinates of the point P on the polygon calculated in S705 are projected onto the captured image 1021 corresponding to the real camera 1027 from the camera viewpoint of the real camera 1027 are obtained.
[0074] In S707, the pixel value of the position indicated by the two-dimensional coordinates on the captured image acquired in S706 is set as the pixel value in the texture image corresponding to point p on the UV map.
[0075] In S708, it is determined whether the above-mentioned process is completed for all triangles included in the target fragment. If the process is completed for all triangles, S709 is executed next. On the other hand, if there are unprocessed triangles, the process returns to S703 and the same process is repeated.
[0076] In S709, it is determined whether the above-mentioned processing is completed for all fragments generated by UV unwrapping in S305 as the target fragment. If processing is completed for all fragments, this flow is exited. On the other hand, if there are unprocessed fragments, the flow returns to S701 and the same processing is repeated.
[0077] The above is the content of the texture image generation process (S306).
[0078] <Modification> In the above embodiment, the angle between the normal vector of the polygon of interest and the direction vector of the real camera of interest is used as a condition for selecting the best camera in the camera selection process (S401) (S503), but this is not limited to this. For example, the area of the polygon obtained when the polygon of interest is projected from the shooting viewpoint of the real camera of interest may be used as a condition for selection. Specifically, three vertices of the polygon of interest are obtained, and the three vertices are projected based on the camera parameters of the real camera of interest to obtain two-dimensional coordinates p0, p1, and p2. Side vectors v1 and v2 are calculated from the two-dimensional coordinates p0, p1, and p2 thus obtained, and the area s of the triangle with the two-dimensional coordinates p0, p1, and p2 as vertices is calculated. The area s in this case is expressed by the following formula (10).
[0079]
number
[0080] As described above, according to this embodiment, polygons are grouped taking into consideration the positional relationship between each polygon constituting the mesh data and a real camera placed in the shooting space, and a real camera or virtual camera with an appropriate camera viewpoint is set for each group. This makes it possible to color fragments generated by UV unwrapping even if real cameras are only present near the in-plane direction of the polygon, resulting in a high-quality texture image.
[0081] [Embodiment 2] In the first embodiment, it is determined whether the angle between the real camera direction and the normal vector of each polygon of the mesh is equal to or less than a threshold, and a virtual camera is set if there is a polygon whose angle with any real camera is equal to or greater than the threshold. Next, a mode in which mesh division processing is performed based on the direction of the normal vector of each polygon of the mesh, virtual cameras are set for all groups, and UV unwrapping processing is performed will be described as the second embodiment.
[0082] <Operation flow of image processing device> The operation of the image processing device 100 according to this embodiment will be described with reference to the flowcharts shown in Figs. 11 to 13 and 15. Fig. 11 is a main routine flowchart showing a rough operational flow of the image processing device 100, and Figs. 12, 13 and 15 are flowcharts of subroutines corresponding to specific processes within the main routine. Below, the operational flow of the image processing device 100 according to this embodiment will be described with reference to each flowchart. In the following description, the symbol "S" means step.
[0083] In the flowchart of FIG. 11, steps S301 and S302 are the same as those in the flowchart of FIG. 3 in the first embodiment, and therefore the description thereof will be omitted.
[0084] In S303' following S302, the mesh division unit 202 performs a process of dividing the mesh data acquired in S301 into a plurality of groups. Fig. 14(a) shows the acquired mesh data 1401 and two real cameras 1402, 1403 corresponding to the plurality of captured images used to generate the mesh data. Fig. 14(b) shows the state in which the mesh data 1401 shown in Fig. 14(a) is divided into five groups 1404, 1405, 1406, 1407, and 1408. In Fig. 14(b), the areas surrounded by thick lines indicate the groups, and it can be seen that each group is made up of a plurality of polygons. Details of this mesh division process will be described later.
[0085] In S304', the camera parameter setting unit 203 sets the camera parameters of a virtual camera to be used in the next UV development process (S305') for each group divided in S303'. The details of this camera parameter setting process will be described later.
[0086] In S305', the UV unwrapping unit 204 performs UV unwrapping for each divided group. In this embodiment, based on the camera parameters of the virtual camera set for each group in S304', fragments are generated by projecting from the camera viewpoint represented by the camera parameters, and the fragments are arranged on a two-dimensional map. In this way, a two-dimensional map (UV map) corresponding to the mesh data, on which the fragments for each group are arranged, is obtained.
[0087] In S306', the texture image generating unit 205 generates a texture image based on the camera parameters of the multiple real cameras corresponding to each captured image acquired in S302 and the two-dimensional map generated in S305'. Details of this texture image generation process will be described later.
[0088] <Details of mesh division process> Next, the mesh division process (S303') according to this embodiment will be described in detail with reference to the flowchart of FIG.
[0089] In the flowchart of Fig. 12, S1201 and S1202 correspond to S402 and S403, respectively, in the flowchart of Fig. 4 of embodiment 1. That is, a group ID for managing a group to be generated is initialized (S1201), and then a polygon of interest is selected from among the polygons constituting the mesh data (S1202).
[0090] In S1203, a group ID is set for the polygon of interest.
[0091] In S1204, one adjacent polygon of interest is selected from adjacent polygons that satisfy a certain condition between the polygon of interest and the adjacent polygon of interest. The certain condition here is that the polygons share one or more of the same sides.
[0092] In S1205, it is determined whether or not a group ID has already been set for the target adjacent polygon. If it is determined that a group ID has been set, the process proceeds to S1209. On the other hand, if it is determined that a group ID has not been set, the process proceeds to S1206.
[0093] In S1206, the angle between the normal vector of the polygon of interest and the normal vector of the adjacent polygon of interest is calculated. The method of calculating the angle between the vectors is as described in S503 of the first embodiment.
[0094] In S1207, it is determined whether the "angle" calculated in S1206 is equal to or less than a preset threshold value. The threshold value is set in advance based on experience, similar to S506 in the first embodiment, and is preferably around 70 degrees, for example. If it is determined to be equal to or less than the threshold value, the process proceeds to S1208. If it is determined to be greater than the threshold value, the process proceeds to S1209.
[0095] In S1208, the same group ID as the group ID set for the target polygon is set for the target adjacent polygon.
[0096] In S1209, it is determined whether the above-mentioned processing has been completed for all adjacent polygons that satisfy a certain condition between the polygon of interest. If the processing has been completed for all adjacent polygons, the process proceeds to S1210. On the other hand, if there are adjacent polygons that have not been processed yet, the process returns to S1204, the next adjacent polygon is selected, and the process continues.
[0097] In S1210, similar to S410 in the first embodiment, it is determined whether or not the above-mentioned processing has been completed for all polygons constituting the mesh data. If processing has been completed for all polygons, the results of the mesh division processing are saved in RAM 102, and the flow exits. On the other hand, if there are unprocessed polygons, in S1211, similar to S411 in the first embodiment, the group ID number is incremented (+1), and the flow returns to S1202 to repeat the same processing.
[0098] The above is the content of the mesh division process (S303'). By the above process, the following correspondence table is created.
[0099] [Table 6] <Details of camera parameter setting process> Next, the camera parameter setting process (S304') according to this embodiment will be described in detail with reference to the flowchart of FIG. 13. The difference from the camera parameter setting process (see FIG. 6) in the first embodiment is that the viewpoint determination (S602) and the camera parameter setting of the real camera (S603) are omitted. As shown in the flowchart of FIG. 13, S1301 corresponds to S601, S1302 corresponds to S604, S1303 corresponds to S605, S1304 corresponds to S606, and S1305 corresponds to S607. That is, for each group obtained by the mesh division process (S303') described above, a virtual camera is set based on the normal vector of each polygon in the group. Thus, in the example of FIG. 14(b), the correspondence table is updated, for example, as follows:
[0100] [Table 7] After the above-mentioned camera parameter setting process (S304') is completed, the process moves to UV development for each group (S305').
[0101] In S305', based on the camera parameters of the virtual camera set for each group, fragments are generated by projecting from the camera viewpoint represented by the camera parameters and arranging them on a two-dimensional map. As a result, a column of fragment ID is added to the correspondence table created in the camera parameter setting process, and the same setting value as the corresponding group ID is stored as the identifier (fragment ID) of the fragment generated for each group. In the example of Figure 14(b), the correspondence table of Table 7 is updated to the following Table 8.
[0102] [Table 8] <Details of texture image generation process> Next, the texture image generation process (S306') according to this embodiment will be described in detail with reference to the flowchart of FIG. 15. The difference from the texture image generation process (see FIG. 7) in the first embodiment is that a camera selection process (S1506) is added between S705 and S706. As shown in the flowchart of FIG. 15, S1501 to S1505 correspond to S701 to S705, respectively, and S1507 to S1510 correspond to S706 to S709, respectively. In this embodiment, when the process up to S1505 (i.e., the process up to S705 in the flowchart of FIG. 7) is completed, a process of selecting a real camera that best captures the polygon acquired in S1503 (S703) is performed in S1506. The camera selection process according to this embodiment will be described in detail with reference to the flowchart of FIG. 16.
[0103] <Details of camera selection process> The difference from the camera selection process in the first embodiment (see FIG. 5) is that the processes related to polygon selection (S501, S509) and the processes for setting flag values according to the viewpoint determination results (S506 to S508) are not necessary and therefore do not exist. The specific process flow is as follows.
[0104] In S1601 (S502), one real camera of interest is selected from among a plurality of real cameras arranged.
[0105] In S1602 (S503), the angle between the normal vector of the polygon acquired in S703 and the direction vector of the target real camera selected in S1601 is calculated.
[0106] In S1603 (S504), it is determined whether the process of S1602 has been completed for all the real cameras that are placed. If the calculation of the angle θ has been completed for all the real cameras, the process of S1604 is then executed. On the other hand, if there is a real camera for which the calculation of the angle θ has not been performed, the process returns to S1601 and the same process is repeated.
[0107] In S1604 (S505), the real camera with the smallest calculated "angle" among all the placed real cameras is associated with the polygon acquired in S703 as the best camera.
[0108] The above is the content of the camera selection process according to this embodiment. As a result, a column of real camera IDs is added to the correspondence table created in the above-mentioned UV unwrapping (S305'), and real camera IDs corresponding to each polygon ID are set. As a result, in the example of FIG. 14(b), the correspondence table of Table 8 is updated to the following Table 9.
[0109] [Table 9] Returning to the description of the texture image generation process.
[0110] In S1507 (S706), the correspondence table updated as described above is referenced, and two-dimensional coordinates on the captured image corresponding to the three-dimensional coordinates calculated in S1505 are obtained. That is, two-dimensional coordinates are obtained when the three-dimensional coordinates calculated in S1505 (S705) are projected onto the captured image corresponding to the real camera selected in S1506 from the camera viewpoint of the real camera. The following will be explained using the above-mentioned case of FIG. 10 as an example. First, the fragment ID of the attention fragment 1011 (selected in S1501) is obtained. Next, the polygon ID of the polygon 1002 corresponding to the attention triangle 1012 included in the attention fragment 1011 is obtained. Next, the real camera 1027 indicated by the real camera ID linked to the polygon ID is identified. Then, the two-dimensional coordinates of a point pi when the three-dimensional coordinates of the point P on the polygon 1002 calculated in S1505 are projected onto the captured image 1021 corresponding to the real camera 1027 from the camera viewpoint of the real camera 1027 are acquired.
[0111] In S1508, the pixel value on the captured image at the two-dimensional coordinates acquired in S1507 is set as the pixel value of the texture image.
[0112] The above is the content of the texture image generation process (S306') according to this embodiment.
[0113] According to this embodiment, the mesh data is first grouped taking into account the direction of the normal vector of each polygon that constitutes the mesh data, and then a virtual camera is set for each group. This enables UV expansion that takes into account the shape of the mesh data, and allows for more accurate fragment generation.
[0114] [Embodiment 3] In the second embodiment, a mesh division process is performed based on the direction of the normal vector of each polygon of a mesh, and virtual cameras are set for all groups. Next, as a third embodiment, a mesh division process and UV unwrapping process are performed based on each virtual camera, in which multiple virtual cameras are placed at predetermined positions.
[0115] <Operation flow of image processing device> The operation of the image processing device 100 according to this embodiment will be described with reference to the flowcharts shown in Fig. 17 and Fig. 19. Fig. 17 is a flowchart of a main routine showing a rough operational flow of the image processing device 100, and Fig. 19 is a flowchart of a subroutine corresponding to a specific process within the main routine. Below, the operational flow of the image processing device 100 according to this embodiment will be described with reference to each flowchart. Note that in the following description, the symbol "S" means a step.
[0116] In the flowchart of FIG. 17, steps S1701 and S1702 correspond to steps S301 and S302 in the flowchart of FIG. 3 of the first embodiment, and there are no particular differences, so a description thereof will be omitted.
[0117] In S1703 following S1702, the camera parameter setting unit 203 sets the camera parameters of each virtual camera so that the virtual cameras are placed at predetermined positions in the imaging space. Here, as an example, a case will be described in which the virtual cameras are placed at equal intervals based on the positions of the mesh data acquired in S1701. Fig. 18 is a top view of a mesh 1801 representing a three-dimensional shape of a person, in which eight virtual cameras 1802a-h are placed to surround the mesh 1801 from the front, back, left, right, and 45-degree angles.
[0118] In S1704, the mesh division unit 202 performs mesh division processing on the mesh data acquired in S1701 based on the camera parameters of the multiple virtual cameras set in S1702.
[0119] <Details of mesh division process> Now, with reference to the flowchart in FIG. 19, the mesh division process according to this embodiment will be described in detail. In S1901, for each polygon of the mesh data, a virtual camera that best captures the respective polygon is set from among the multiple virtual cameras set in the above-mentioned S1703. The following steps S1902 to S1907 correspond to S402, S403, S405, S406, S410, and S411 in the flowchart of Fig. 4 of the first embodiment, respectively. That is, first, in S1902, a group ID is initialized (S402), and then, in S1903, one polygon that does not yet belong to any group is selected as a polygon of interest from among the polygons constituting the mesh data (S403).
[0120] Then, in S1904, adjacent polygons that satisfy a certain condition with the selected polygon of interest are grouped together with the polygon of interest (S405). In this grouping, the virtual camera set for the polygon of interest in S1901 is first identified. Next, of the other polygons to which the same virtual camera as the identified virtual camera is associated, all polygons adjacent to the polygon of interest are identified. Then, a group consisting of all the identified adjacent polygons and the polygon of interest is created. In S1905, a group ID that identifies the group is set for the polygon of interest and its adjacent polygons that belong to the group created in S1904 (S406). Then, in S1906, it is determined whether all polygons that make up the mesh data have been processed (S410). If there are any unprocessed polygons, the group ID number is incremented in S1907 (S411), and the process returns to S1903 and the same process is repeated.
[0121] The above is the content of the mesh division process (S1704) according to this embodiment. This enables UV development using a virtual camera that is not dependent on the shape of the mesh data or the position of the real camera, and speeds up the fragment generation process.
[0122] (Other Examples) The present disclosure can also be realized by a process in which a program for implementing one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more functions.
[0123] The present disclosure also includes the following configurations and methods.
[0124] [Configuration 1] An acquisition means for acquiring three-dimensional shape data of an object captured in a plurality of images taken from different viewpoints; a dividing means for dividing elements constituting the three-dimensional shape data into a plurality of groups; a control means for associating viewpoint information representing a specific photographing viewpoint with each of the plurality of groups; a first generating means for generating a two-dimensional map based on the plurality of groups and viewpoint information representing the specific shooting viewpoint associated with each group; a second generating means for generating a texture image representing a color of the object based on the two-dimensional map; Equipped with The specific shooting viewpoint includes a virtual shooting viewpoint different from the shooting viewpoints of each of the plurality of captured images. 13. An image processing device comprising:
[0125] [Configuration 2] the acquiring means acquires a plurality of pieces of viewpoint information representing shooting viewpoints of the plurality of captured images, the dividing means divides elements constituting the three-dimensional shape data into a plurality of groups based on the plurality of pieces of viewpoint information; 2. The image processing device according to claim 1,
[0126] [Configuration 3] The image processing device according to configuration 2, characterized in that when none of the multiple shooting viewpoints represented by the multiple viewpoint information acquired by the acquisition means satisfies a condition, the control means performs the association by setting a virtual shooting viewpoint different from the multiple shooting viewpoints represented by the multiple viewpoint information as the specific shooting viewpoint.
[0127] [Configuration 4] The image processing device according to configuration 3, characterized in that the control means performs the association by treating a shooting viewpoint that satisfies the condition among a plurality of shooting viewpoints represented by the plurality of viewpoint information acquired by the acquisition means as the specific shooting viewpoint.
[0128] [Configuration 5] 5. The image processing device according to configuration 4, wherein the condition is that an angle between a direction vector of a shooting viewpoint corresponding to viewpoint information of interest among the plurality of viewpoint information and a normal vector of an element constituting the group is equal to or smaller than a threshold value.
[0129] [Configuration 6] 5. The image processing device according to configuration 4, wherein the condition is that an angle between a direction vector of a shooting viewpoint corresponding to viewpoint information of interest among the plurality of viewpoint information and a normal vector of an element constituting the group is equal to or smaller than a threshold value.
[0130] [Configuration 7] 5. The image processing device according to configuration 4, wherein the condition is that an area of a shape obtained when projecting elements constituting the group at a shooting viewpoint of interest among the shooting viewpoints represented by the plurality of viewpoint information is equal to or larger than a threshold value.
[0131] [Configuration 8] The image processing device according to configuration 7, characterized in that when there is no shooting viewpoint whose area is equal to or greater than the threshold value among the shooting viewpoints represented by the plurality of viewpoint information, the control means performs the association by treating a virtual shooting viewpoint having a line of sight direction along the normal vector of the elements constituting the group as the specific shooting viewpoint.
[0132] [Configuration 9] 9. The image processing device according to any one of configurations 1 to 8, wherein the dividing means divides elements constituting the three-dimensional shape data into groups based on normal vectors of the elements.
[0133] [Configuration 10] The image processing device according to configuration 9, characterized in that the control means performs the association by treating a virtual shooting viewpoint having a line of sight direction along the normal vector of the elements constituting the group as the specific shooting viewpoint.
[0134] [Configuration 11] the dividing means divides the elements constituting the three-dimensional shape data into a plurality of groups based on a positional relationship between a plurality of virtual shooting viewpoints set in advance and the elements constituting the three-dimensional shape data; the control means performs the association by setting the virtual shooting viewpoint used when the dividing means divides the images into groups as the specific shooting viewpoint. 2. The image processing device according to claim 1,
[0135] [Configuration 12] 12. The image processing device according to configuration 11, wherein the dividing means divides the three-dimensional shape data into groups based on an angle between a direction vector of the virtual shooting viewpoint and a normal vector of an element that constitutes the three-dimensional shape data.
[0136] [Configuration 13] 12. The image processing device according to configuration 11, wherein the dividing means divides the three-dimensional shape data into groups based on an area of a shape obtained when elements constituting the three-dimensional shape data are projected from the virtual shooting viewpoint.
[0137] [Configuration 14] the first generating means generates a two-dimensional map corresponding to the three-dimensional shape data, in which two-dimensional fragments obtained by performing UV development for each group are arranged; the second generation means generates the texture image by calculating a pixel value of each pixel included in each two-dimensional fragment on the two-dimensional map using the plurality of captured images. 14. The image processing device according to any one of configurations 1 to 13.
[0138] [Configuration 15] The image processing device described in configuration 14, characterized in that the second generation means uses a captured image among the multiple captured images in which a group corresponding to the two-dimensional fragment is captured to calculate pixel values of each pixel whose pixel center is included in the two-dimensional fragment, thereby generating the texture image.
[0139] [Configuration 16] the elements constituting the three-dimensional shape data are polygons, The three-dimensional shape data is mesh data that represents a surface shape of the object by a set of the polygons. 16. The image processing device according to any one of configurations 1 to 15.
[0140] [Configuration 17] 16. The image processing device according to any one of configurations 1 to 15, wherein the viewpoint information includes at least information for specifying a position and an orientation of a corresponding shooting viewpoint.
[0141] [Method 1] An acquisition step of acquiring three-dimensional shape data of an object captured in a plurality of images taken from different viewpoints; a dividing step of dividing elements constituting the three-dimensional shape data into a plurality of groups; a control step of associating viewpoint information representing a specific shooting viewpoint with each of the plurality of groups; a first generation step of generating a two-dimensional map based on the plurality of groups and viewpoint information representing the specific shooting viewpoint associated with each group; a second generation step of generating a texture image representing a color of the object based on the two-dimensional map; having The specific shooting viewpoint includes a virtual shooting viewpoint different from the shooting viewpoints of each of the plurality of captured images. 13. An image processing method comprising:
[0142] [Configuration 11] A program for causing a computer to function as the image processing device according to any one of configurations 1 to 17.
Claims
1. An acquisition means for acquiring three-dimensional shape data of an object captured in a plurality of images taken from different viewpoints; a dividing means for dividing elements constituting the three-dimensional shape data into a plurality of groups; a control means for associating viewpoint information representing a specific photographing viewpoint with each of the plurality of groups; a first generating means for generating a two-dimensional map based on the plurality of groups and viewpoint information representing the specific shooting viewpoint associated with each group; a second generation means for generating a texture image representing a color of the object based on the two-dimensional map; Equipped with The specific shooting viewpoint includes a virtual shooting viewpoint different from the shooting viewpoints of each of the plurality of captured images.
13. An image processing device comprising:
2. the acquiring means acquires a plurality of pieces of viewpoint information representing shooting viewpoints of the plurality of captured images, the dividing means divides elements constituting the three-dimensional shape data into a plurality of groups based on the plurality of pieces of viewpoint information; 2. The image processing device according to claim 1,
3. The image processing device according to claim 2, characterized in that, when none of the multiple shooting viewpoints represented by the multiple viewpoint information acquired by the acquisition means satisfies a condition, the control means performs the association with a virtual shooting viewpoint different from the multiple shooting viewpoints represented by the multiple viewpoint information as the specific shooting viewpoint.
4. The image processing device according to claim 3 , wherein the control means performs the association by setting a shooting viewpoint that satisfies the condition among a plurality of shooting viewpoints represented by the plurality of viewpoint information acquired by the acquisition means as the specific shooting viewpoint.
5. The image processing device according to claim 4 , wherein the condition is that an angle between a direction vector of a shooting viewpoint corresponding to viewpoint information of interest among the plurality of viewpoint information and a normal vector of the elements constituting the group is equal to or smaller than a threshold value.
6. The image processing device according to claim 5, characterized in that, when there is no shooting viewpoint represented by the multiple viewpoint information such that the angle formed is equal to or less than a threshold value, the control means performs the association by assuming a virtual shooting viewpoint having a line of sight direction along the normal vector of the elements constituting the group as the specific shooting viewpoint.
7. The image processing device according to claim 4 , wherein the condition is that an area of a shape obtained when projecting the elements constituting the group from a shooting viewpoint of interest among the shooting viewpoints represented by the plurality of viewpoint information is equal to or larger than a threshold value.
8. The image processing device according to claim 7, characterized in that, when there is no shooting viewpoint whose area is equal to or greater than a threshold value among the shooting viewpoints represented by the multiple viewpoint information, the control means performs the association with a virtual shooting viewpoint having a line of sight direction along the normal vector of the elements constituting the group as the specific shooting viewpoint.
9. 2. The image processing apparatus according to claim 1, wherein said dividing means divides elements constituting said three-dimensional shape data into groups based on normal vectors of the elements.
10. 10. The image processing device according to claim 9, wherein the control means performs the association by treating a virtual shooting viewpoint having a line of sight direction along a normal vector of the elements constituting the group as the specific shooting viewpoint.
11. the dividing means divides the elements constituting the three-dimensional shape data into a plurality of groups based on a positional relationship between a plurality of virtual shooting viewpoints set in advance and the elements constituting the three-dimensional shape data; the control means performs the association by setting the virtual shooting viewpoint used when the dividing means divides the images into groups as the specific shooting viewpoint.
2. The image processing device according to claim 1,
12. 12. The image processing apparatus according to claim 11, wherein said dividing means divides the image into groups based on an angle between a direction vector of the virtual shooting viewpoint and a normal vector of an element constituting the three-dimensional shape data.
13. 12. The image processing apparatus according to claim 11, wherein the dividing means divides the three-dimensional shape data into groups based on an area of a shape obtained when elements constituting the three-dimensional shape data are projected from the virtual shooting viewpoint.
14. the first generating means generates a two-dimensional map corresponding to the three-dimensional shape data, in which two-dimensional fragments obtained by performing UV expansion for each group are arranged; the second generation means generates the texture image by calculating a pixel value of each pixel included in each two-dimensional fragment on the two-dimensional map using the plurality of captured images.
2. The image processing device according to claim 1,
15. The image processing device according to claim 14, characterized in that the second generation means uses a captured image among the plurality of captured images in which a group corresponding to the two-dimensional fragment is captured, calculates pixel values of each pixel whose pixel center is contained within the two-dimensional fragment, and generates the texture image.
16. the elements constituting the three-dimensional shape data are polygons, The three-dimensional shape data is mesh data that represents a surface shape of the object by a set of the polygons.
2. The image processing device according to claim 1,
17. The image processing apparatus according to claim 1 , wherein the viewpoint information includes at least information for specifying a position and a posture of a corresponding imaging viewpoint.