System, information processing device, and method

By converting three-dimensional point cloud data into two-dimensional point cloud data using a correspondence between x and y coordinates and a parameter m, the method addresses inefficiencies in existing mesh data generation techniques, achieving faster and more efficient mesh data generation for 3D models.

JP7740048B2Active Publication Date: 2025-09-17TOYOTA JIDOSHA KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022020857
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-09-17
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

Existing techniques for generating mesh data for 3D models are inefficient and require significant processing time.

Method used

A system and method that converts three-dimensional point cloud data into two-dimensional point cloud data using a correspondence between x and y coordinates and a parameter m, allowing for faster generation of mesh data by utilizing two-dimensional point cloud data.

Benefits of technology

The method significantly speeds up the mesh data generation process by converting three-dimensional point cloud data into two-dimensional point cloud data, improving efficiency in generating mesh data for 3D models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740048000001
    Figure 0007740048000001
  • Figure 0007740048000002
    Figure 0007740048000002
  • Figure 0007740048000003
    Figure 0007740048000003
Patent Text Reader

Abstract

To improve a technique for generating mesh data for 3D models.SOLUTION: A system 1 includes a first information processing device 10a and a second information processing device 10b that can communicate with each other. The first information processing device 10a acquires three-dimensional point cloud data including a plurality of vertices each having three parameters of x-, y-, and z-coordinates, and a combination of the x- and y-coordinates is different from vertex to vertex. The first information processing device 10a or the second information processing device 10b converts the three-dimensional point cloud data into two-dimensional point cloud data in which each of a plurality of vertices has two parameters of m with a different value for each combination of the x- and y-coordinates, and the z-coordinate. The second information processing device 10b generates mesh data using the two-dimensional point cloud data and information indicating correspondence between the combination of the x- and y-coordinates and the parameter m.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a system, an information processing device, and a method. [Background technology]

[0002] Conventionally, techniques for generating mesh data of a 3D model have been known. For example, Patent Document 1 discloses a technique for generating an analytical mesh based on point sequence data expressed by (x, y, z) coordinates obtained by measuring the three-dimensional shape of an object to be analyzed. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 06-301767 Summary of the Invention [Problem to be solved by the invention]

[0004] There was room for improvement in the technology for generating mesh data for 3D models.

[0005] In view of the above circumstances, an object of the present disclosure is to improve the technology for generating mesh data for 3D models. [Means for solving the problem]

[0006] In accordance with one embodiment of the present disclosure, a system includes: A system including a first information processing device and a second information processing device that can communicate with each other, The first information processing device Acquire three-dimensional point cloud data including a plurality of vertices, each vertex having three parameters of an x-coordinate, a y-coordinate, and a z-coordinate, and the combinations of the x-coordinate and the y-coordinate of each vertex are different from each other; The first information processing device or the second information processing device converting the three-dimensional point cloud data into two-dimensional point cloud data in which each of the plurality of vertices has two parameters, m and z coordinate, which have different values ​​for each combination of x coordinate and y coordinate; The second information processing device Mesh data is generated using information indicating the correspondence between the combination of x coordinates and y coordinates and m, and the two-dimensional point cloud data.

[0007] An information processing device according to an embodiment of the present disclosure includes: An information processing device including a control unit, The control unit Acquire two-dimensional point cloud data converted from three-dimensional point cloud data including a plurality of vertices, each of which has three parameters, x-coordinate, y-coordinate, and z-coordinate, and each of which has a different combination of x-coordinate and y-coordinate, and each of which has two parameters, m and z-coordinate, whose values ​​differ for each combination of x-coordinate and y-coordinate; Mesh data is generated using information indicating the correspondence between the combination of x coordinates and y coordinates and m, and the two-dimensional point cloud data.

[0008] According to one embodiment of the present disclosure, a method comprises: A method executed by an information processing device, Acquiring two-dimensional point cloud data converted from three-dimensional point cloud data including a plurality of vertices, each of which has three parameters, x-coordinate, y-coordinate, and z-coordinate, and each of which has a different combination of x-coordinate and y-coordinate, and each of which has two parameters, m and z-coordinate, whose values ​​differ for each combination of x-coordinate and y-coordinate; and The method includes generating mesh data using information indicating the correspondence between a combination of x coordinates and y coordinates and m, and the two-dimensional point cloud data. [Effects of the Invention]

[0009] According to one embodiment of the present disclosure, a technique for generating mesh data for a 3D model is improved. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram illustrating a schematic configuration of a system according to an embodiment of the present disclosure. [Figure 2] 1A and 1B are diagrams illustrating examples of the imaging range and resolution of a visible light camera and a depth camera. [Figure 3] FIG. 10 is a diagram illustrating an example of three-dimensional point cloud data. [Figure 4] FIG. 10 is a diagram illustrating an example of two-dimensional point cloud data converted from three-dimensional point cloud data. [Figure 5] FIG. 1 is a block diagram showing a schematic configuration of an information processing device. [Figure 6] FIG. 10 is a sequence diagram illustrating an example of the operation of the system. [Figure 7] FIG. 10 is a schematic diagram showing a cross section of the generated mesh data in the yz plane. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described.

[0012] (Outline of the embodiment) An overview of a system 1 according to an embodiment of the present disclosure will be described with reference to FIG. 1. The system 1 includes a plurality of information processing devices 10. Although two information processing devices 10 (10a, 10b) are illustrated in FIG. 1, the system 1 may include three or more information processing devices 10. In the following description, when distinguishing between the two information processing devices 10 (10a, 10b), they will be referred to as a first information processing device 10a and a second information processing device 10b.

[0013] The information processing device 10 is any computer used by a user, such as a PC (Personal Computer), a smartphone, or a tablet terminal. The multiple information processing devices 10 can communicate with each other via a network 20 including, for example, the Internet. Note that communication between the information processing devices 10 may be performed via a server provided on the network 20, or may be performed using P2P.

[0014] In this embodiment, the system 1 is used to provide a remote dialogue service in which users of each information processing device 10 engage in audio dialogue while viewing video of their conversation partner (another user). Specifically, each information processing device 10 is equipped with a visible light camera and a depth camera, as described below, and is capable of capturing an image of a subject to generate a visible light image and a depth image. Each pixel in the visible light image has an x-coordinate, a y-coordinate, and color information (e.g., RGB values). Meanwhile, each pixel in the depth image has three parameters: an x-coordinate, a y-coordinate, and a z-coordinate. That is, the depth image data is 3D point cloud data consisting of N vertices P, each of which has three parameters: an x-coordinate, a y-coordinate, and a z-coordinate. For ease of explanation, this embodiment will be described using an example in which the visible light camera and the depth camera have substantially the same imaging range and resolution, and the resolution of the generated visible light image and depth image is both 640 pixels × 480 pixels, as shown in FIG. 2 . In this case, the number N of vertices P constituting the 3D point cloud data is N = 640 × 480 = 307,200.

[0015] Each information processing device 10 generates mesh data corresponding to the shape of the user who is the subject of the visible light camera and the depth camera from the 3D point cloud data, maps a visible light image onto the mesh data to create a 3D model of the user, and can display on its screen a rendered image obtained by photographing the 3D model placed in a 3D virtual space with a virtual camera. During a remote conversation, each information processing device 10 performs real-time rendering of the 3D virtual space in which the 3D models of each conversation partner (i.e., users of other information processing devices 10) are placed. In this way, during a remote conversation, the 3D models of the conversation partners displayed on the screen of each information processing device 10 move so as to follow the actual movements of the conversation partners.

[0016] An outline of this embodiment will now be described, and details will be given later. In this embodiment, mesh data is generated using two-dimensional point cloud data converted from three-dimensional point cloud data.

[0017] Specifically, the first information processing device 10a acquires three-dimensional point cloud data including a plurality of vertices P. Here, as shown in FIG. 3, each vertex P of the three-dimensional point cloud data has three parameters: an x-coordinate, a y-coordinate, and a z-coordinate, and the combinations of the x-coordinate and the y-coordinate of each vertex P are different from one another. The parameter m shown in FIG. 3 is a parameter whose value varies for each combination of the x-coordinate and the y-coordinate. Therefore, there is a one-to-one correspondence between the combination of the x-coordinate and the y-coordinate and the parameter m, and once one is determined, the other is also uniquely determined. Therefore, in the three-dimensional point cloud data, each vertex P can be expressed as Pm(xm, ym, zm) using the parameter m. For example, when the resolution of the depth image is 640 pixels × 480 pixels, m can take a value between 1 and 307200 (1≦m≦307200).

[0018] Subsequently, the first information processing device 10a or the second information processing device 10b converts the three-dimensional point cloud data into two-dimensional point cloud data. Here, each vertex P of the two-dimensional point cloud data has two parameters, m, whose value differs for each combination of x coordinate and y coordinate, and z coordinate, as shown in Fig. 4. Therefore, each vertex P in the two-dimensional point cloud data can be expressed as Pm(m, zm) using the parameter m.

[0019] The second information processing device 10b then generates mesh data using information indicating the correspondence between a combination of x and y coordinates and the parameter m, and the above-mentioned two-dimensional point cloud data. Here, the specific correspondence between the combination of x and y coordinates and the parameters can be determined arbitrarily. In the example shown in Figures 3 and 4, the information indicating the correspondence between the combination of x and y coordinates and the parameter m indicates that m = 1 corresponds to (0,0), m = 2 corresponds to (1,0), and m = 307200 corresponds to (639,479).

[0020] As described above, each vertex P included in the 3D point cloud data according to this embodiment has three parameters: an x ​​coordinate, a y coordinate, and a z coordinate. Here, since the combinations of the x coordinate and the y coordinate of each vertex P are different from each other, by predefining the correspondence between the combination of the x coordinate and the y coordinate and the parameter m, the 3D point cloud data can be converted into 2D point cloud data in which each vertex P has two parameters: m and z coordinate. According to this embodiment, mesh data is generated using the 2D point cloud data, and therefore the technique for generating mesh data of a 3D model is improved in that the mesh data generation process is faster than, for example, a configuration in which mesh data is generated using 3D point cloud data.

[0021] Next, each component of the system 1 will be described in detail.

[0022] (Configuration of information processing device) As shown in FIG. 5, the information processing device 10 includes a communication unit 11, an output unit 12, an input unit 13, a sensor unit 14, a storage unit 15, and a control unit 16.

[0023] The communication unit 11 includes one or more communication interfaces that connect to the network 20. The communication interfaces correspond to mobile communication standards such as 4G (4th Generation) or 5G (5th Generation), wireless LAN (Local Area Network), or wired LAN, but are not limited to these and may correspond to any communication standard.

[0024] The output unit 12 includes one or more output devices that output information to notify the user. The output devices include, but are not limited to, a display that outputs information as a video or a speaker that outputs information as a sound. Alternatively, the output unit 12 may include an interface for connecting an external output device.

[0025] The input unit 13 includes one or more input devices that detect user input. Examples of the input devices include, but are not limited to, physical keys, capacitive keys, a mouse, a touch panel, a touch screen integrated with the display of the output unit 12, or a microphone. Alternatively, the input unit 13 may include an interface for connecting an external input device.

[0026] The sensor unit 14 includes one or more sensors. The sensor may be, for example, but is not limited to, a visible light camera that generates a visible light image or a depth camera that generates a depth image. Alternatively, the sensor unit 14 may include an interface for connecting an external sensor. For example, a Kinect (registered trademark) having a visible light camera and a depth camera may be employed as the sensor unit 14.

[0027] In this embodiment, the sensor unit 14 includes both a visible light camera and a depth camera. The visible light camera and the depth camera may be provided, for example, close to each other so as to have substantially the same imaging range. In the example shown in FIG. 2, the visible light camera and the depth camera can image a user present at a specified position (for example, a position facing the display of the output unit 12) from the front. Alternatively, the visible light camera and the depth camera may be provided so as to be able to image a user present at a specified position from different angles. In such a case, image processing for viewpoint conversion may be performed on at least one of the visible light image and the depth image so as to obtain an image of the user photographed from the front.

[0028] The storage unit 15 includes one or more memories. The memories may be, for example, semiconductor memories, magnetic memories, optical memories, or the like, but are not limited to these. Each memory included in the storage unit 15 may function as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 15 stores any information used in the operation of the information processing device 10. For example, the storage unit 15 may store system programs, application programs, embedded software, and the like.

[0029] The control unit 16 includes one or more processors, one or more programmable circuits, one or more dedicated circuits, or a combination thereof. The processor may be, for example, a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for a specific process, but is not limited to these. The programmable circuit may be, for example, but is not limited to, an FPGA (Field-Programmable Gate Array). The dedicated circuit may be, for example, but is not limited to, an ASIC (Application Specific Integrated Circuit). The control unit 16 controls the overall operation of the information processing device 10.

[0030] (System operation flow) The operation of the system 1 according to this embodiment will be described with reference to Fig. 6. In summary, this operation is one of the operations executed during a remote dialogue between a user a of the first information processing device 10a and a user b of the second information processing device 10b, in which the second information processing device 10b generates and displays a rendering image of a 3D model of user a. This operation is repeatedly executed during the remote dialogue, for example, at a frame rate that is predetermined or in accordance with the communication speed.

[0031] Step S100: The first information processing device 10a acquires a visible light image and a depth image of the user a.

[0032] Specifically, the control unit 16a of the first information processing device 10a acquires a visible light image and a depth image of user a by capturing an image of user a using the visible light camera and the depth camera of the sensor unit 14a. Here, the description will be given assuming that the visible light camera and the depth camera capture an image of user a from the front, who is positioned opposite the display of the output unit 12a. Also, as shown in FIG. 2, the description will be given assuming that the capturing ranges of the visible light camera and the depth camera are substantially the same, and that the resolution of the captured visible light image and the depth image is 640 pixels × 480 pixels. As described above, each pixel in the captured visible light image has an x ​​coordinate, a y coordinate, and color information (e.g., RGB value). Meanwhile, each vertex P included in the captured depth image (three-dimensional point cloud data) is expressed as vertex Pm(xm, ym, zm) using parameter m, whose value varies for each combination of x coordinate and y coordinate, as shown in FIG. 3.

[0033] The control unit 16a may perform any pre-processing, such as noise removal and alignment, on the acquired depth image (three-dimensional point cloud data).

[0034] Step S101: The first information processing device 10a identifies each vertex P corresponding to a person (here, user a) who is a subject, from among all vertices P included in the depth image of step S100.

[0035] Any method can be used to identify each vertex P corresponding to a person who is a subject. For example, the control unit 16a may identify, among all vertices P included in the depth image, vertices P whose z coordinates (i.e., distances from the depth camera) are within a predetermined range as vertices P corresponding to the person. Alternatively, the control unit 16a may detect the contour of the person from at least one of the visible light image and the depth image by image recognition, and identify each vertex P located within the detected contour as vertices P corresponding to the person. In such a case, the control unit 16a may identify each vertex P for each part of the person (e.g., head, shoulders, trunk, arms, etc.). Any method can be used to estimate each part of the person, such as image recognition or skeletal detection using at least one of the visible light image and the depth image.

[0036] Here, the control unit 16a may delete vertices P that do not correspond to a person (for example, vertices P that correspond to the background of the person) from the depth image (three-dimensional point cloud data). In this case, the depth image (three-dimensional point cloud data) includes only vertices P that correspond to people, so the number of vertices P included in the depth image (three-dimensional point cloud data) becomes less than N (here, 640 × 480 = 307,200), resulting in a reduction in the amount of data.

[0037] Step S102: The first information processing device 10a determines a reference value z0 based on the plurality of vertices P identified in step S101.

[0038] In summary, the reference value z0 is a representative value of the distance between the depth camera and the person (here, user a) who is the subject. The reference value z0 is information used in step S107 for generating mesh data, and will be described in detail later.

[0039] Any method can be used to determine the reference value z0. For example, the control unit 16a may determine the z-coordinate value of any one of the multiple vertices P identified in step S101 as corresponding to the person who is the subject, as the reference value z0. Here, the control unit 16a may determine the z-coordinate value of any one of the vertices P corresponding to the person's head, shoulders, or trunk, as the reference value z0. Alternatively, the control unit 16a may determine a value (for example, an average value or a median value) calculated based on the z-coordinates of two or more vertices P of the identified multiple vertices P, as the reference value z0. Here, the control unit 16a may determine the value calculated based on the z-coordinates of two or more vertices P corresponding to the person's head, shoulders, or trunk, as the reference value z0.

[0040] Step S103: The first information processing device 10a detects the posture of the person (here, user a) who is the subject.

[0041] Any method can be used to detect the posture of a person. For example, the control unit 16a may detect the posture of a person by image recognition or skeletal detection using at least one of a visible light image and a depth image. In this embodiment, the posture detected by the control unit 16a is selected from three postures: a forward leaning posture, a backward leaning posture, and an upright posture. The forward leaning posture is a posture in which the upper body of the person leans forward. The backward leaning posture is a posture in which the upper body of the person leans backward. The upright posture is a posture in which the upper body of the person is not leaning forward or backward. However, this is not limited to this example, and the control unit 16a may detect more detailed postures. For example, the control unit 16a may detect the tilt angle of the upper body of the person in the front-to-back direction as the posture.

[0042] Step S104: The first information processing device 10a determines Δz based on the posture detected in step S103.

[0043] Roughly speaking, Δz is a value for defining the distance range z0±Δz in the z direction based on the above-mentioned reference value z0. The distance range z0±Δz is information used in step S107 for generating mesh data, and will be described in detail later.

[0044] Any method can be used to determine Δz. For example, consider a situation in which a person, who is a subject, has his or her hand thrust out so as to block the space between the head, shoulders, or trunk and the depth camera. For example, the control unit 16a may set Δz to a predetermined value (e.g., 30 cm) when the posture detected in step S103 is an upright posture, while the control unit 16a may set Δz to a larger value when the posture detected in step S103 is a forward-leaning posture or a backward-leaning posture than when the posture is an upright posture. Alternatively, when the tilt angle of the person's upper body in the front-to-back direction is detected as the posture in step S103, the control unit 16a may increase Δz as the absolute value of the tilt angle increases. Regardless of the method, by appropriately determining Δz according to the detected posture, it is possible to estimate, for example, that the head, shoulders, and trunk are within the range of z0±Δz, while the thrust hand is outside the range of z0±Δz.

[0045] Step S105: The first information processing device 10a converts the three-dimensional point cloud data into two-dimensional point cloud data.

[0046] Here, the 2D point cloud data is point cloud data in which each vertex has two parameters, m and z coordinate, which have different values ​​for each combination of x coordinate and y coordinate. As described above, each vertex P included in the 3D point cloud data has three parameters, x coordinate, y coordinate, and z coordinate, and the combinations of x coordinate and y coordinate of each vertex P are different. Therefore, each vertex P included in the 3D point cloud data can be expressed as vertex Pm(xm, ym, zm) using parameter m, as shown in FIG. 3. On the other hand, each vertex P included in the 2D point cloud data has two parameters, m and z coordinate. Therefore, each vertex P included in the 2D point cloud data can be expressed as vertex Pm(m, zm), as shown in FIG. 4.

[0047] 3 and 4, the correspondence relationship between the combination of x coordinates and y coordinates and the parameter m is defined such that m=1 corresponds to (x=0, y=0), m=2 corresponds to (x=1, y=0), and m=307200 corresponds to (x=639, y=479). However, the correspondence relationship between the combination of x coordinates and y coordinates and the parameter m is not limited to this example and can be defined arbitrarily. Information indicating the correspondence relationship between the combination of x coordinates and y coordinates and the parameter m may be stored in advance in the storage unit 15a, for example.

[0048] In this way, in step S105, the data of each vertex Pm is converted from three-dimensional data having three parameters, x coordinate, y coordinate, and z coordinate, to two-dimensional data having two parameters, parameter m and z coordinate.

[0049] Step S106: The first information processing device 10a transmits any information used to generate a 3D model of the person (here, user a) who is the subject, to the second information processing device 10b.

[0050] In this embodiment, the control unit 16a transmits the visible light image acquired in step S100, the reference value z0 determined in step S102, the Δz determined in step S104, the two-dimensional point cloud data generated in step S105, and information indicating the correspondence between combinations of x and y coordinates and the parameter m to the second information processing device 10b via the communication unit 11a and the network 20. Note that if the vertices P that do not correspond to people are not deleted from the depth image (three-dimensional point cloud data) in step S101, the control unit 16a may transmit information indicating the vertices P identified in step S101 as corresponding to people who are the subjects (for example, the parameter m corresponding to each identified vertex P) to the second information processing device 10b.

[0051] Step S107: The second information processing device 10b generates mesh data of the person (here, user a) who is the subject, using information indicating the correspondence between the combination of x coordinates and y coordinates and the parameter m, and the two-dimensional point cloud data.

[0052] As described above, in this embodiment, the combination of x coordinates and y coordinates of each vertex P is different from one another, so by defining in advance the correspondence relationship between the combination of x coordinates and y coordinates and the parameter m, it is possible to convert the three-dimensional point cloud data into two-dimensional point cloud data in which each vertex P has two parameters, m and z coordinate. According to this embodiment, mesh data is generated using the two-dimensional point cloud data, so the mesh data generation process is faster than, for example, a configuration in which mesh data is generated using three-dimensional point cloud data.

[0053] Any algorithm, such as the Delaunay algorithm or the Advancing Front algorithm, can be used to generate the mesh data. Specifically, the control unit 16b of the second information processing device 10b generates an edge connecting two vertices P, each of which is identified according to the algorithm used, from among the multiple vertices P included in the 2D point cloud data. The control unit 16b then generates polygons for each infinitesimal plane surrounded by a predetermined number of edges. Typically, each polygon formed is a triangular polygon including three vertices and three edges connecting the three vertices, but the shape of the polygon is not limited to this example. In this way, the control unit 16b generates a set of vertices P, edges, and polygons as mesh data.

[0054] When generating mesh data, if the z coordinate of one of two vertices P is within the range of z0±Δz and the z coordinate of the other vertex P is outside the range of z0±Δz, the control unit 16a may not form a polygon including those two vertices. This configuration can reduce the occurrence of a problem where a polygon is formed between the head, shoulders, or trunk and the hand, for example, when a person (here, user a) who is the subject of the image sticks out his or her hand so as to block the distance between the head, shoulders, or trunk and the depth camera. This will be described below with reference to FIG. 7.

[0055] 7 is a schematic diagram showing a cross section of the generated mesh data in the yz plane. Mesh M1 corresponds to the head (face), mesh M2 corresponds to the outstretched hand, and mesh M3 corresponds to the torso. As described above, in this embodiment, the depth camera of the first information processing device 10a captures the image of user a from the front, so no corresponding mesh is formed for parts of user a's body (specifically, the back part and the part blocked by the outstretched hand). However, for ease of explanation, the parts of user a's body for which no corresponding mesh is formed are shown with dashed lines.

[0056] Here, we focus on vertex Pa located at the bottom of mesh M1 and vertex Pb located at the top of mesh M2. Vertices Pa and Pb are adjacent to each other in the y direction. The z-coordinate of vertex Pa is within the range of z0±Δz. On the other hand, the z-coordinate of vertex Pb is outside the range of z0±Δz. Therefore, a polygon including vertices Pa and Pb is not formed, which reduces the occurrence of a problem such as an unnatural connection between mesh M1 corresponding to the head and mesh M2 corresponding to the outstretched hand. The same is true for vertex Pc located at the bottom of mesh M2 and vertex Pd located at the top of mesh M3. Vertices Pc and Pd are adjacent to each other in the y direction. The z-coordinate of vertex Pc is outside the range of z0±Δz. On the other hand, the z-coordinate of vertex Pd is outside the range of z0±Δz. Therefore, since a polygon including vertices Pc and Pd is not formed, the occurrence of a problem such as mesh M2 corresponding to an outstretched hand and mesh M3 corresponding to the torso being unnaturally connected is reduced.

[0057] Step S108: The second information processing device 10b maps the visible light image onto the mesh data generated in step S107, and creates a 3D model of the person (here, user a) who is the subject.

[0058] As described above, in this embodiment, the visible light camera and the depth camera of the first information processing device 10a have substantially the same imaging range and resolution. Therefore, each pixel of the visible light image has a one-to-one correspondence with each vertex P of the depth image (3D point cloud data). Based on this correspondence, the control unit 16b maps the area of ​​the visible light image corresponding to user a onto mesh data. In this way, a 3D model of user a is created.

[0059] Step S109: The second information processing device 10b generates and displays a rendering image of the 3D model created in step S108.

[0060] Specifically, the control unit 16b places a 3D model of a person (here, user a) who is the subject in the three-dimensional virtual space. The control unit 16b generates a rendering image by capturing an image of the 3D model placed in the three-dimensional virtual space with a virtual camera, and displays the rendering image on the display of the output unit 12b. Note that the position of the virtual camera in the three-dimensional virtual space may be determined in advance, or may be linked to the position of user b of the second information processing device 10b in the real space.

[0061] As described above, in the system 1 according to this embodiment, the first information processing device 10a acquires three-dimensional point cloud data including a plurality of vertices P. In the three-dimensional point cloud data, each vertex P has three parameters: an x-coordinate, a y-coordinate, and a z-coordinate, and the combinations of the x-coordinate and the y-coordinate of each vertex P are different from one another. The first information processing device 10a converts the three-dimensional point cloud data into two-dimensional point cloud data. In the two-dimensional point cloud data, each of the plurality of vertices P has two parameters: m and z-coordinate, whose values ​​vary for each combination of x-coordinate and y-coordinate. Then, the second information processing device 10b generates mesh data using the two-dimensional point cloud data and information indicating the correspondence between the combination of x-coordinate and y-coordinate and m.

[0062] With this configuration, the combination of x coordinates and y coordinates of each vertex P is different from one another, and so by defining in advance the correspondence between the combination of x coordinates and y coordinates and the parameter m, it is possible to convert the three-dimensional point cloud data into two-dimensional point cloud data in which each vertex P has two parameters, m and z coordinate. According to this embodiment, mesh data is generated using two-dimensional point cloud data, and therefore the technology for generating mesh data for a 3D model is improved in that the mesh data generation process is faster than, for example, a configuration in which mesh data is generated using three-dimensional point cloud data.

[0063] Although the present disclosure has been described based on the drawings and examples, it should be noted that those skilled in the art may make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included in the scope of the present disclosure. For example, the functions included in each component or step can be rearranged so as not to be logically inconsistent, and multiple components or steps can be combined or divided into one.

[0064] For example, an embodiment is also possible in which the second information processing device 10b executes some of the operations executed by the first information processing device 10a in the above-described embodiment. For example, the operations of steps S101 to S105 described above may be executed by the second information processing device 10b. In such a case, the first information processing device 10a transmits the visible light image and depth image (three-dimensional point cloud data) acquired in step S100 to the second information processing device 10b, and the second information processing device 10b executes the operations of steps S101 to S105. In other words, the operations of steps S101 to S105 may be executed by either the first information processing device 10a or the second information processing device 10b.

[0065] Furthermore, in the above-described embodiment, when generating mesh data, the second information processing device 10b may adjust the resolution of the generated mesh data according to the distance (i.e., z coordinate) between the depth camera of the first information processing device 10a and the person (here, user a) who is the subject. The resolution of the mesh data can be reduced by any method, such as thinning out some vertices P included in the 2D point cloud data. Reducing the resolution of the mesh data reduces the processing load for generating and rendering the mesh data. For example, the control unit 16b of the second information processing device 10b may reduce the resolution as the distance between the depth camera and user a increases (i.e., as the z coordinate increases). When user a is farther away from the depth camera, it is difficult to accurately represent details such as user a's facial expression in a 3D model. Therefore, intentionally reducing the resolution of the generated mesh data does not adversely affect the visibility of the screen. Alternatively, the control unit 16b may reduce the resolution as the distance between the depth camera and user a decreases (i.e., as the z coordinate decreases). According to this configuration, for example, it is possible to suppress changes in the resolution of the 3D data of user a in response to changes in the distance between the depth camera and user a during remote dialogue.

[0066] Furthermore, in the above-described embodiment, when generating mesh data, the second information processing device 10b may adjust the resolution of the generated mesh data for each body part of the subject person (here, user a). For example, the control unit 16b of the second information processing device 10b may lower the resolution of mesh data corresponding to body parts other than the head of user a (for example, shoulders, torso, arms, etc.). This configuration makes it possible to reduce the processing load of generating and rendering mesh data while maintaining a certain level of accuracy in reproducing facial expressions, which are a relatively important element in remote dialogue communication.

[0067] Furthermore, in the above-described embodiment, a person is shown as a specific example of the subject, but any object other than a person may be adopted as the subject.

[0068] Also, for example, an embodiment is possible in which a general-purpose computer functions as the information processing device 10 according to the above-described embodiment. Specifically, a program describing the processing content for realizing each function of the information processing device 10 according to the above-described embodiment is stored in the memory of the general-purpose computer, and the program is read and executed by a processor. Therefore, the present disclosure can also be realized as a program executable by a processor, or a non-transitory computer-readable medium storing the program. [Explanation of symbols]

[0069] 1 System 10. Information processing equipment 11 Communications Department 12 Output section 13 Input section 14 Sensor unit 15 Storage section 16 Control Unit 20 Network

Claims

1. A system including a first information processing device and a second information processing device that can communicate with each other, The first information processing device Acquire three-dimensional point cloud data including a plurality of vertices, each vertex having three parameters of an x-coordinate, a y-coordinate, and a z-coordinate, and the combinations of the x-coordinate and the y-coordinate of each vertex are different from one another; The first information processing device or the second information processing device converting the three-dimensional point cloud data into two-dimensional point cloud data in which each of the plurality of vertices has two parameters, m and z coordinate, whose values ​​differ for each combination of x coordinate and y coordinate; The second information processing device A system that generates mesh data using information indicating a correspondence relationship between a combination of x coordinates and y coordinates and m, and the two-dimensional point cloud data.

2. The system of claim 1 When generating mesh data, the second information processing device does not form a polygon comprising two vertices if the z coordinate of one of the two vertices is within the range of a reference value z0 ± Δz and the z coordinate of the other vertex is outside the range of the reference value z0 ± Δz.

3. The system of claim 2, the three-dimensional point cloud data is data of a depth image of a person photographed by a depth camera, The first information processing device or the second information processing device Identifying each of the plurality of vertices, among all vertices included in the depth image, corresponding to the person; The system determines the reference value z0 to be the z-coordinate value of any one of the plurality of vertices, or a value calculated based on the z-coordinates of two or more of the plurality of vertices.

4. 4. The system of claim 3, The first information processing device or the second information processing device Detecting the posture of the person; The system determines Δz based on the person's pose.

5. 5. The system according to claim 3 or 4, When generating mesh data, the second information processing device adjusts the resolution of the generated mesh data according to the distance between the depth camera and the person.

6. 6. A system according to any one of claims 3 to 5, The second information processing device adjusts a resolution of the generated mesh data for each part of the person when generating the mesh data.

7. An information processing device including a control unit, The control unit Acquire two-dimensional point cloud data converted from three-dimensional point cloud data including a plurality of vertices, each of which has three parameters, x-coordinate, y-coordinate, and z-coordinate, and each of which has a different combination of x-coordinate and y-coordinate, and each of which has two parameters, m and z-coordinate, whose values ​​differ for each combination of x-coordinate and y-coordinate; An information processing device that generates mesh data using information indicating a correspondence relationship between a combination of x coordinates and y coordinates and m, and the two-dimensional point cloud data.

8. A method executed by an information processing device, Acquiring two-dimensional point cloud data converted from three-dimensional point cloud data including a plurality of vertices, each of which has three parameters, x-coordinate, y-coordinate, and z-coordinate, and each of which has a different combination of x-coordinate and y-coordinate, and each of which has two parameters, m and z-coordinate, whose values ​​differ for each combination of x-coordinate and y-coordinate; and generating mesh data using information indicating a correspondence relationship between a combination of x coordinates and y coordinates and m, and the two-dimensional point cloud data.

Citation Information

Patent Citations

  • Mesh generation method

    JP1994301767A

  • Information processing apparatus, information processing method, and program

    JP2018116537A

  • Image processing apparatus, image processing method, and image processing program

    JP2019211965A

  • Image Processing Apparatus, Image Processing Method, And Image Processing Program

    US20190371058A1

  • Shape measuring method

    WO2012164987A1