Three-dimensional model generation method and three-dimensional model generation device
By determining a search range in three-dimensional space using subject information, the method improves the accuracy and efficiency of three-dimensional model generation, addressing the limitations of existing technologies in searching for similar points across multiple images.
Patent Information
- Application Number
- JP2023563508
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-29
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-24
AI Technical Summary
Existing three-dimensional model generation methods face challenges in improving accuracy and reducing processing time, particularly in searching for similar points across multiple images.
A method that determines a search range in three-dimensional space based on subject information without using map information, allowing for precise matching of similar points between camera images.
This approach enhances the accuracy of three-dimensional model generation and reduces processing time by focusing the search for similar points within a likely range based on subject information.
Smart Images

Figure 0007692175000004 
Figure 0007692175000005 
Figure 0007692175000006
Abstract
Description
Technical Field
[0001] The present disclosure relates to a three-dimensional model generation method and a three-dimensional model generation apparatus.
Background Art
[0002] Patent Document 1 discloses a technique for generating a three-dimensional model of a subject using a plurality of images obtained by photographing the subject from a plurality of viewpoints.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the generation process of a three-dimensional model, it is desired to improve the generation accuracy of the three-dimensional model and reduce the processing time.
[0005] The present disclosure provides a three-dimensional model generation method and the like that can improve the generation accuracy of a three-dimensional model and can shorten the processing time of the generation process of the three-dimensional model.
Means for Solving the Problems
[0006] A three-dimensional model generation method according to an aspect of the present disclosure is a three-dimensional model generation method executed by an information processing apparatus, the method including: obtaining subject information including a plurality of positions on a subject in a three-dimensional space; obtaining a first camera image of the subject taken from a first viewpoint and a second camera image of the subject taken from a second viewpoint; determining a search range in a three-dimensional space including a first three-dimensional point on the subject corresponding to a first point in the first camera image based on the subject information without using map information including a plurality of three-dimensional points each indicating a position on the subject in the three-dimensional space, the map information obtained by camera calibration performed by causing one or more cameras to photograph the subject from a plurality of viewpoints including the first viewpoint and the second viewpoint; performing matching to search for a similar point similar to the first point in a range corresponding to the search range in the second camera image; and generating a three-dimensional model using a search result of the matching Moreover, the first camera image and the second camera image are images obtained by photographing with the one or more cameras, and the subject information is generated by a sensor different from the one or more cameras. 。
[0007] A three-dimensional model generation apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to obtain subject information including a plurality of positions on a subject in a three-dimensional space, obtain a first camera image of the subject taken from a first viewpoint and a second camera image of the subject taken from a second viewpoint, determine a search range in a three-dimensional space including a first three-dimensional point on the subject corresponding to a first point in the first camera image based on the subject information without using map information including a plurality of three-dimensional points each indicating a position on the subject in the three-dimensional space, the map information obtained by camera calibration performed by causing one or more cameras to photograph the subject from a plurality of viewpoints including the first viewpoint and the second viewpoint; perform matching to search for a similar point similar to the first point in a range corresponding to the search range in the second camera image; and generate a three-dimensional model using a search result of the matching Moreover, the first camera image and the second camera image are images obtained by photographing with the one or more cameras, and the subject information is generated by a sensor different from the one or more cameras. 。
[0008] In addition, a three-dimensional model generation device according to an aspect of the present disclosure includes a memory and a processor connected to the memory. The processor acquires a first camera image generated by photographing a subject in a three-dimensional space from a first viewpoint and a second camera image generated by photographing the subject from a second viewpoint, searches for a second point similar to the first point within a search range on an epipolar line specified by projecting a straight line passing through the first viewpoint and a first point of the first camera image onto the second camera image, generates a three-dimensional model of the subject based on the result of the search, and the search range is provided based on the position of a first three-dimensional point corresponding to the first point in the three-dimensional space. The first camera image and the second camera image are images obtained by photographing with one or more cameras. The position is a reflected wave of an electromagnetic wave emitted toward the subject. a sensor to be used, which is by a sensor different from the one or more cameras is obtained.
[0009] Note that the present disclosure may be implemented as a program that causes a computer to execute steps included in the above three-dimensional model generation method. Further, the present disclosure may be implemented as a non-transitory recording medium such as a CD-ROM readable by a computer on which the program is recorded. Further, the present disclosure may be implemented as information, data, or a signal indicating the program. And those programs, information, data, and signals may be distributed via a communication network such as the Internet.
Effects of the Invention
[0010] According to the present disclosure, it is possible to provide a three-dimensional model generation method and the like that can improve the generation accuracy of a three-dimensional model and can shorten the processing time of the three-dimensional model generation process.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Embodiments for Carrying Out the Invention
[0012] (Background Leading to the Present Disclosure) In the technology disclosed in Patent Document 1, a three-dimensional model is generated by searching for similar points among a plurality of images. Generally, in searching for similar points, when searching for a similar point of one pixel of one image from other images, an epipolar line on the other image is calculated from the geometric constraints of the camera, and all pixels on the epipolar line are searched. For this reason, there is room for improving the processing speed of searching for similar points. Also, when there are similar subjects on the epipolar line, incorrect similar points may be searched, and in this case, there is a problem that it leads to a decrease in the accuracy of the search. Also, not only on the epipolar line, but also when searching for similar points from a search range such as the entire image or a predetermined region, there is a problem of searching for incorrect similar points.
[0013] Therefore, the present disclosure provides a three-dimensional model generation method and the like that can improve the generation accuracy of a three-dimensional model and can shorten the processing time of the three-dimensional model.
[0014] A three-dimensional model generation method according to an aspect of the present disclosure is a three-dimensional model generation method executed by an information processing apparatus. The method includes obtaining subject information including a plurality of positions on a subject in a three-dimensional space, obtaining a first camera image of the subject taken from a first viewpoint and a second camera image of the subject taken from a second viewpoint, and performing camera calibration executed by causing one or more cameras to capture the subject from a plurality of viewpoints including the first viewpoint and the second viewpoint. Without using map information including a plurality of three-dimensional points each indicating a position on the subject in the three-dimensional space, a search range in the three-dimensional space including a first three-dimensional point on the subject corresponding to a first point in the first camera image is determined based on the subject information. Matching is performed to search for a similar point similar to the first point in a range corresponding to the search range on the second camera image. A three-dimensional model is generated using the search result in the matching.
[0015] According to this, the search range is determined based on the subject information without using the map information, and a similar point similar to the first point on the first camera image is searched in a range corresponding to the search range on the second camera image restricted by the search range. In this way, since the search for similar points is performed in a range where similar points are likely to exist based on the subject information, the search accuracy of similar points can be improved, and the time required for the search process can be shortened. Therefore, the generation accuracy of the three-dimensional model can be improved, and the processing time of the three-dimensional model generation process can be shortened.
[0016] Also, for example, in the matching, the epipolar line corresponding to the first point in the second camera image is restricted to a length corresponding to the search range, and a similar point similar to the first point is searched on the epipolar line of the second camera image.
[0017] According to this, since a similar point similar to the first point is searched on the epipolar line restricted to a length corresponding to the search range, the search accuracy of similar points can be improved, and the time required for the search process can be shortened.
[0018] Further, for example, the subject information includes a distance image generated by measurement using a distance image sensor. The distance image has a plurality of pixels each having distance information indicating the distance from the distance image sensor to the subject. In the determination, the search range may be determined based on the distance information of the pixel corresponding to the first point in the distance image.
[0019] According to this, since the subject information includes a distance image having a plurality of pixels associated with the plurality of pixels of the first camera image, the distance information corresponding to the first point can be easily specified. Therefore, based on the specified distance information, the position of the first three-dimensional point can be estimated, and the search range can be determined accurately.
[0020] Further, for example, the subject information includes a plurality of distance images respectively generated by measurement using a plurality of distance image sensors. Each of the plurality of distance images has a plurality of pixels each having distance information indicating the distance from the distance image sensor that generated the distance image to the subject. The plurality of pixels included in each of the plurality of distance images are respectively associated with the plurality of pixels of the camera image corresponding to the distance image having the plurality of pixels among the plurality of camera images. The plurality of camera images include the first camera image and the second camera image. In the determination, the search range may be determined based on the one or more distance information of the one or more pixels respectively corresponding to the first point in one or more of the plurality of distance images.
[0021] According to this, since the subject information includes a plurality of distance images each having a plurality of pixels associated with the plurality of pixels of the first camera image, a plurality of distance information corresponding to the first point can be easily specified. Since the plurality of distance information thus specified is distance information obtained from different viewpoints, even if some of the distance information includes detection errors, the influence of the detection errors can be reduced by utilizing other distance information. Therefore, based on one or more of the plurality of distance information, the position of the first three-dimensional point can be estimated more accurately, and the search range can be determined accurately.
[0022] Also, for example, in the determination, when the detection accuracy of the first distance information of the pixel corresponding to the first point in the first distance image corresponding to the first camera image is lower than a predetermined accuracy, the third distance information corresponding to the first point calculated using two or more camera images other than the first camera image may be used as the one or more distance information to determine the search range.
[0023] Therefore, when the detection accuracy of the first distance information is low, the search range can be determined using the third distance information with high accuracy. Thus, the search range can be determined accurately.
[0024] Also, for example, the plurality of distance image sensors each correspond in position and orientation to a plurality of cameras including the one or more cameras, and in the determination, the one or more pixels respectively corresponding to the first point in the one or more distance images may be specified using the positions and orientations of the plurality of cameras obtained by the camera calibration.
[0025] Therefore, using the positions and orientations of the plurality of cameras obtained by the camera calibration, one or more distance information can be specified.
[0026] Further, for example, the one or more distance images include a first distance image corresponding to the first camera image and a second distance image corresponding to the second camera image, and the second camera image may be determined from the plurality of camera images based on the number of feature points between the second camera image and the first camera image in the feature point matching in the camera calibration.
[0027] According to this, based on the number of feature points, the second camera image to be the object of matching of similarity points with the first camera image is determined. For this reason, it is possible to specify the second distance image for specifying one or more distance information with a high possibility of not including errors, that is, with high accuracy.
[0028] Further, for example, the second camera image may be determined based on the difference in the shooting pose calculated from the first position and pose at the time of shooting of the camera that shot the first camera image and the second position and pose at the time of shooting of the camera that shot the second camera image.
[0029] According to this, based on the difference in the poses of the cameras, the second camera image to be the object of matching of similarity points with the first camera image is determined. For this reason, it is possible to specify the second distance image for specifying one or more distance information with a high possibility of not including errors, that is, with high accuracy.
[0030] Further, for example, the second camera image may be determined based on the difference in the shooting position calculated from the first position and pose at the time of shooting of the camera that shot the first camera image and the second position and pose at the time of shooting of the camera that shot the second camera image.
[0031] According to this, based on the difference in the positions of the cameras, the second camera image to be the object of matching of similarity points with the first camera image is determined. For this reason, it is possible to specify the second distance image for specifying one or more distance information with a high possibility of not including errors, that is, with high accuracy.
[0032] Further, for example, the difference between the maximum value and the minimum value of the one or more distance information may be less than a first value.
[0033] According to this, one or more distance information in which the difference between the maximum value and the minimum value is less than the first value is specified. Thereby, it is possible to specify one or more distance information that is highly likely to be free of errors, that is, has high accuracy.
[0034] Also, for example, in the determination, the search range may be widened as the accuracy of the one or more distance information is lower.
[0035] According to this, since the search range is widened as the accuracy of the one or more distance information decreases, it is possible to determine the search range according to the accuracy.
[0036] Also, for example, the accuracy may be higher as the number of the one or more distance information is larger.
[0037] According to this, it can be determined that the accuracy of the one or more distance information is higher as the number of the one or more distance information is larger, that is, as the number of the one or more similar distance information is larger. Therefore, the search range can be narrowed as the number of the one or more distance information is larger.
[0038] Also, for example, the accuracy may be higher as the variance of the one or more distance information is smaller.
[0039] According to this, it can be determined that the accuracy of the plurality of distance information is higher as the variance of the plurality of distance information is smaller, that is, as the plurality of distance information is more similar. Therefore, the search range can be narrowed as the variance of the plurality of distance information is smaller.
[0040] Also, for example, the subject information may be generated based on two or more types of sensor information.
[0041] According to this, the subject information is generated based on two or more types of sensor information that are different from each other. That is, subject information with reduced accuracy degradation due to detection errors can be obtained.
[0042] Further, for example, the two or more types of sensor information may include a plurality of two-dimensional images obtained from a stereo camera and three-dimensional data obtained from a measuring instrument that emits electromagnetic waves and acquires reflected waves reflected by the subject from the electromagnetic waves.
[0043] According to this, since the subject information is generated based on the plurality of two-dimensional images and the three-dimensional data, for example, three-dimensional data in which the three-dimensional data is densified using the plurality of two-dimensional images can be accurately obtained.
[0044] Also, a three-dimensional model generation device according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to acquire subject information including a plurality of positions on a subject in a three-dimensional space, and acquires a first camera image of the subject taken from a first viewpoint and a second camera image of the subject taken from a second viewpoint. The camera calibration is performed by causing one or more cameras to photograph the subject from a plurality of viewpoints including the first viewpoint and the second viewpoint, and without using map information including a plurality of three-dimensional points each indicating a position on the subject in the three-dimensional space, based on the subject information, a search range in the three-dimensional space including a first three-dimensional point on the subject corresponding to a first point in the first camera image is determined, and matching is performed to search for a similar point similar to the first point in a range corresponding to the search range on the second camera image. A three-dimensional model is generated using the search result in the matching.
[0045] According to this, the search range is determined based on the subject information without using the map information, and a similar point similar to the first point on the first camera image is searched in a range corresponding to the search range on the second camera image limited by the search range. In this way, since the search for similar points is performed in a range where there is a high possibility of the existence of similar points based on the subject information, the search accuracy of similar points can be improved, and the time required for the search process can be shortened. Therefore, the generation accuracy of the three-dimensional model can be improved, and the processing time of the three-dimensional model generation process can be shortened.
[0046] Also, a three-dimensional model generation device according to an aspect of the present disclosure includes a memory and a processor connected to the memory. The processor acquires a first camera image generated by photographing a subject in a three-dimensional space from a first viewpoint and a second camera image generated by photographing the subject from a second viewpoint, searches for a second point similar to the first point within a search range on an epipolar line specified by projecting a straight line passing through the first viewpoint and a first point of the first camera image onto the second camera image, generates a three-dimensional model of the subject based on the result of the search, the search range is provided based on the position of a first three-dimensional point corresponding to the first point in the three-dimensional space, and the position is obtained based on a reflected wave of an electromagnetic wave emitted toward the subject.
[0047] According to this, the search range is determined based on the position of the first three-dimensional point obtained based on the reflected wave of the electromagnetic wave, and within the range corresponding to the search range on the second camera image restricted by the search range, a similar point similar to the first point on the first camera image is searched. In this way, based on the subject information, since the search for similar points is performed within a range where there is a high possibility of the existence of similar points, the search accuracy of similar points can be improved, and the time required for the search process can be shortened. Therefore, the generation accuracy of the three-dimensional model can be improved, and the processing time of the three-dimensional model generation process can be shortened.
[0048] Further, for example, the position may be obtained based on a distance image generated by a sensor that receives the reflected wave.
[0049] For this reason, the position of the first three-dimensional point can be easily specified based on the distance image. Therefore, the search range can be determined accurately.
[0050] Hereinafter, each embodiment of the three-dimensional model generation method and the like according to the present disclosure will be described in detail with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. Therefore, the numerical values, shapes, materials, components, arrangements and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure.
[0051] Also, each figure is a schematic diagram and is not necessarily drawn precisely. In each figure, substantially the same configuration is denoted by the same reference numeral, and duplicate explanations may be omitted or simplified.
[0052] (Embodiment) [Overview] First, with reference to FIG. 1, an overview of the three-dimensional model generation method according to the embodiment will be described.
[0053] FIG. 1 is a diagram for explaining an overview of the three-dimensional model generation method according to the embodiment. FIG. 2 is a block diagram showing a characteristic configuration of the three-dimensional model generation system according to the embodiment.
[0054] In the three-dimensional model generation method, as shown in FIG. 1, a three-dimensional model of a predetermined region is generated from a plurality of images taken at different viewpoints using a plurality of cameras 310. Here, the predetermined region is a region including a stationary object that is stationary, a moving object such as a person, or both. In other words, the predetermined region is, for example, a region including at least one of a stationary object that is stationary and a moving object as a subject.
[0055] Examples of the predetermined region including a stationary object and a moving object include a venue where a sports game such as basketball is being held, a space on a road where a person or a vehicle exists, and the like. Note that the predetermined region may include not only a specific object serving as a subject but also a scenery or the like. FIG. 1 illustrates a case where the subject 500 is a building. In the following, a predetermined region including not only a specific object serving as a subject but also a scenery or the like will also be simply referred to as a subject.
[0056] As shown in FIG. 2, the three-dimensional model generation system 400 includes a camera group 300 including a plurality of cameras 310, an estimation device 200, and a three-dimensional model generation device 100.
[0057] (A plurality of cameras) The plurality of cameras 310 are a plurality of imaging devices that photograph a predetermined area. The plurality of cameras 310 each photograph a subject and output the plurality of photographed frames to the estimation device 200 respectively. The plurality of photographed frames are also referred to as multi-viewpoint images. In the present embodiment, the camera group 300 includes two or more cameras 310. Also, the plurality of cameras 310 photograph the same subject from different viewpoints. A frame is, in other words, an image.
[0058] Note that although the three-dimensional model generation system 400 is described as including the camera group 300, it is not limited thereto and may include only one camera 310. For example, in the three-dimensional model generation system 400, a subject existing in the real space may be photographed at a plurality of different timings while moving one camera 310 so as to generate a multi-viewpoint image composed of a plurality of frames with different viewpoints from one camera 310. In this case, each of the plurality of frames is associated with the position and orientation of the camera 310 at the timing when the frame is photographed. Each of the plurality of frames is a frame photographed (generated) by cameras 310 in which at least one of the position and orientation is different from each other. The cameras 310 in which at least one of the position and orientation is different from each other may be realized by a plurality of cameras 310 with fixed positions and orientations, may be realized by one camera 310 in which at least one of the position and orientation is not fixed, or may be realized by a combination of a camera 310 with a fixed position and orientation and a camera 310 in which at least one of the position and orientation is not fixed.
[0059] In addition, each camera 310 generates a camera image. The camera image has a plurality of pixels arranged two-dimensionally. Each pixel of the camera image may have color information or luminance information as a pixel value. Further, each camera 310 may be a camera including a distance image sensor 320. The distance image sensor 320 generates a distance image (depth map) by measuring the distance to the subject at the position of each pixel. The distance image has a plurality of pixels arranged two-dimensionally. Each pixel of the distance image may have distance information indicating the distance from the camera 310 to the subject at the position corresponding to the pixel as a pixel value. The distance image is an example of subject information including a plurality of positions on the subject in three-dimensional space.
[0060] In the present embodiment, the plurality of cameras 310 are cameras each including a distance image sensor 320 that generates a distance image. That is, there is a fixed correspondence between the positions and postures of the plurality of cameras 310 and the positions and postures of the plurality of distance image sensors 320. The plurality of cameras 310 generate a camera image and a distance image as frames. The plurality of pixels of the camera image generated by each camera 310 may be associated with the plurality of pixels of the distance image generated by the camera 310, respectively.
[0061] The distance image sensor 320 may be a ToF (Time of Flight) camera. Further, the distance image sensor 320 may be a sensor that emits electromagnetic waves like the measuring instrument 321 described later in Modification 1, and acquires a reflected wave obtained by reflecting the electromagnetic waves from the subject, thereby generating a distance image.
[0062] The resolution (number of pixels) of the camera image and the resolution (number of pixels) of the distance image may be the same or different. When the resolution of the camera image and the resolution of the distance image are different, one pixel with a lower resolution among the camera image and the distance image may be associated with a plurality of pixels with a higher resolution.
[0063] The plurality of cameras 310 may generate camera images and distance images with the same resolution as each other, or may generate camera images and distance images with different resolutions from each other.
[0064] The camera image and the distance image may be output from the camera 310 as an integrated image in which they are integrated. That is, the integrated image may be an image including a plurality of pixels each having color information indicating the color of the pixel and distance information as pixel values.
[0065] The plurality of cameras 310 may be directly connected to the estimation device 200 by wired communication or wireless communication so that each can output the frame captured by itself to the estimation device 200, or may be indirectly connected to the estimation device 200 via a hub (not shown) such as a communication device or a server.
[0066] Note that the frames captured by the plurality of cameras 310 may be output to the estimation device 200 in real time. Also, the frames may be output from the external storage device such as a memory or a cloud server to the estimation device 200 after being once recorded in the external storage device.
[0067] Also, the plurality of cameras 310 may each be a fixed camera such as a surveillance camera, a mobile camera such as a video camera, a smartphone, or a wearable camera, or a moving camera such as a drone with a photographing function.
[0068] (Estimation device) The estimation device 200 performs camera calibration by causing one or more cameras 310 to photograph a subject from a plurality of viewpoints. The estimation device 200 performs camera calibration for estimating the positions and postures of the plurality of cameras 310 based on, for example, a plurality of frames respectively captured by the plurality of cameras 310. Here, the posture of the camera 310 indicates at least one of the photographing direction of the camera 310 and the inclination of the camera 310. The photographing direction of the camera 310 is the direction of the optical axis of the camera 310. The inclination of the camera 310 is the rotation angle around the optical axis of the camera 310 from the reference posture.
[0069] Specifically, the estimation device 200 estimates the camera parameters of the plurality of cameras 310 based on a plurality of frames (a plurality of camera images) acquired from the plurality of cameras 310. Here, the camera parameters are parameters indicating the characteristics of the camera 310, and include internal parameters such as the focal length and image center of the camera 310, and external parameters indicating the position (more specifically, the three-dimensional position) and orientation of the camera 310. That is, the position and orientation of each of the plurality of cameras 310 can be obtained by estimating the camera parameters of each of the plurality of cameras 310.
[0070] Note that the estimation method by which the estimation device 200 estimates the position and orientation of the camera 310 is not particularly limited. The estimation device 200 may, for example, estimate the position and orientation of the plurality of cameras 310 using Visual-SLAM (Simultaneous Localization and Mapping) technology. Alternatively, the estimation device 200 may, for example, estimate the position and orientation of the plurality of cameras 310 using Structure-From-Motion technology.
[0071] Here, the camera calibration by the estimation device 200 will be described with reference to FIG. 3.
[0072] As shown in FIG. 3, the estimation device 200 uses Visual-SLAM technology or Structure-From-Motion technology to extract characteristic points as feature points 541 to 543 from each of the plurality of frames 531 to 533 captured by the plurality of cameras 310, and searches for features to extract a set of similar points that are similar among the plurality of frames from among the plurality of extracted feature points 541 to 543. By performing the search for feature points, the estimation device 200 can identify points on the subject 510 that are common to the plurality of frames 531 to 533, and thus can obtain the three-dimensional coordinates of the points on the subject 510 using the principle of triangulation with the set of extracted similar points.
[0073] In this way, the estimation device 200 can extract a plurality of sets of similarity points and use the plurality of sets of similarity points to estimate the position and orientation of each camera 310. In the process of estimating the position and orientation of each camera 310, the estimation device 200 calculates three-dimensional coordinates for each set of similarity points and generates map information 520 including a plurality of three-dimensional points represented by the calculated plurality of three-dimensional coordinates. Each of the plurality of three-dimensional points indicates a position on the subject in three-dimensional space. The estimation device 200 obtains the position and orientation of each camera 310 and the map information as estimation results. Since the obtained map information is optimized together with the camera parameters, it is information with higher accuracy than a predetermined accuracy. Also, the map information includes the three-dimensional positions of each of the plurality of three-dimensional points. Note that the map information may include not only the plurality of three-dimensional positions but also information such as the color of each three-dimensional point, the surface shape around each three-dimensional point, and information indicating by which frame each three-dimensional point was generated.
[0074] Further, in order to speed up the estimation process, the estimation device 200 may generate map information including a sparse three-dimensional point cloud by limiting the number of sets of similarity points to a predetermined number. This is because the estimation device 200 can estimate the position and orientation of each camera 310 with sufficient accuracy even with a predetermined number of sets of similarity points. Note that the predetermined number may be determined to be a number that can estimate the position and orientation of each camera 310 with sufficient accuracy. Also, the estimation device 200 may estimate the position and orientation of each camera 310 using sets that are similar with a similarity of a predetermined similarity or more among the sets of similarity points. As a result, the estimation device 200 can limit the number of sets of similarity points used in the estimation process to the number of sets that are similar with a similarity of a predetermined similarity or more.
[0075] Also, the estimation device 200 may calculate, as camera parameters, the distance between the camera 310 and the subject based on, for example, the position and orientation of the camera 310 estimated using the above technique. Note that the three-dimensional model generation system 400 may include a distance measurement sensor, and the distance between the camera 310 and the subject may be measured using the distance measurement sensor.
[0076] The estimation device 200 may be directly connected to the three-dimensional model generation device 100 by wired communication or wireless communication, or may be indirectly connected to the estimation device 200 via a hub (not shown) such as a communication device or a server. Thereby, the estimation device 200 outputs a plurality of frames received from the plurality of cameras 310 and a plurality of camera parameters of the plurality of estimated cameras 310 to the three-dimensional model generation device 100.
[0077] Note that the estimation result by the estimation device 200 may be output to the three-dimensional model generation device 100 in real time. Further, the estimation result may be output to the three-dimensional model generation device 100 from those external storage devices after being once recorded in an external storage device such as a memory or a cloud server.
[0078] The estimation device 200 includes at least a computer system including, for example, a control program, a processing circuit such as a processor or a logic circuit that executes the control program, and a recording device such as an internal memory or an accessible external memory that stores the control program.
[0079] (Three-dimensional model generation device) Based on a plurality of frames captured by the plurality of cameras 310 and the estimation result (the position and orientation of each camera 310) of the estimation device 200, the three-dimensional model generation device 100 generates a three-dimensional model of a predetermined region. Specifically, the three-dimensional model generation device 100 is a device that executes a three-dimensional model generation process of generating a three-dimensional model of a subject on a virtual three-dimensional space based on the camera parameters of each of the plurality of cameras 310 and the plurality of frames.
[0080] Note that the three-dimensional model of the subject is data including the three-dimensional shape of the subject and the color of the subject, which is restored from the frames in which the actual object of the subject is photographed onto a virtual three-dimensional space. The three-dimensional model of the subject is a set of points indicating the three-dimensional positions of the plurality of points on the subject that appear in each of the plurality of camera images photographed by the plurality of cameras 310 at a plurality of different viewpoints, that is, multi-viewpoints.
[0081] The three-dimensional position is represented by three-value information including an X component, a Y component, and a Z component indicating the respective positions of the X-axis, Y-axis, and Z-axis orthogonal to each other, for example. Note that the three-dimensional position is not limited to the coordinates indicated in the orthogonal coordinate system, and may be the coordinates indicated in the polar coordinate system. Note that the information including a plurality of points indicating the three-dimensional position may include not only the three-dimensional position (that is, the information indicating the coordinates), but also the information indicating the color of each point, the information representing the surface shape of each point and its periphery, and the like.
[0082] The three-dimensional model generation device 100 includes at least a computer system including, for example, a control program, a processing circuit such as a processor or a logic circuit that executes the control program, and a recording device such as an internal memory or an accessible external memory that stores the control program. The three-dimensional model generation device 100 is an information processing device. The functions of each processing unit of the three-dimensional model generation device 100 may be realized by software or by hardware.
[0083] Further, the three-dimensional model generation device 100 may store camera parameters in advance. In this case, the three-dimensional model generation system 400 may not include the estimation device 200. Further, the plurality of cameras 310 may be communicably connected to the three-dimensional model generation device 100 wirelessly or by wire.
[0084] Further, the plurality of frames captured by the camera 310 may be directly output to the three-dimensional model generation device 100. In this case, the camera 310 may be directly connected to the three-dimensional model generation device 100 by, for example, wired communication or wireless communication, or may be indirectly connected to the three-dimensional model generation device 100 via a hub (not shown) such as a communication device or a server.
[0085] [Configuration of Three-Dimensional Model Generation Device] Subsequently, with reference to FIG. 2, the details of the configuration of the three-dimensional model generation device 100 will be described.
[0086] The three-dimensional model generation device 100 is a device that generates a three-dimensional model from a plurality of frames. The three-dimensional model generation device 100 includes a reception unit 110, a storage unit 120, an acquisition unit 130, a determination unit 140, a generation unit 150, and an output unit 160.
[0087] The reception unit 110 receives, from the estimation device 200, a plurality of frames captured by a plurality of cameras 310 and an estimation result including the positions and postures of the respective cameras 310 obtained by the estimation device 200. By receiving the plurality of frames, the reception unit 110 acquires a first frame (a first camera image and a first distance image) of the subject captured from a first viewpoint and a second frame (a second camera image and a second distance image) of the subject captured from a second viewpoint. That is, the plurality of frames received by the reception unit 110 includes the first frame and the second frame. The reception unit 110 outputs the received plurality of frames and the estimation result to the storage unit 120.
[0088] The reception unit 110 is, for example, a communication interface for communicating with the estimation device 200. When the three-dimensional model generation device 100 and the estimation device 200 perform wireless communication, the reception unit 110 includes, for example, an antenna and a wireless communication circuit. Alternatively, when the three-dimensional model generation device 100 and the estimation device 200 perform wired communication, the reception unit 110 includes, for example, a connector connected to a communication line and a wired communication circuit. Note that the reception unit 110 may receive a plurality of frames from the plurality of cameras 310 without going through the estimation device 200.
[0089] The storage unit 120 stores the plurality of frames and the estimation results received by the reception unit 110. By storing the plurality of frames, the storage unit 120 stores a distance image, which is an example of subject information included in the plurality of frames. Further, the storage unit 120 stores the search range calculated by the determination unit 140. Note that the storage unit 120 may store the processing results of the processing units included in the three-dimensional model generation apparatus 100. The storage unit 120 stores, for example, a control program for causing a processing circuit to execute the processing by each processing unit included in the three-dimensional model generation apparatus 100. The storage unit 120 is realized by, for example, an HDD (Hard Disk Drive), a flash memory, or the like.
[0090] The acquisition unit 130 acquires from the storage unit 120 the plurality of frames stored in the storage unit 120 and the camera parameters of each camera 310 among the estimation results, and outputs them to the determination unit 140 and the generation unit 150.
[0091] Note that the three-dimensional model generation apparatus 100 may not include the storage unit 120 and the acquisition unit 130. Further, the reception unit 110 may output to the determination unit 140 and the generation unit 150 the plurality of frames received from the plurality of cameras 310 and the camera parameters of each camera 310 among the estimation results received from the estimation apparatus 200.
[0092] When a plurality of pixels of the camera image obtained by each camera 310 and a plurality of pixels of the distance image are not associated with each other, the determination unit 140 associates the plurality of pixels of the camera image with the plurality of pixels of the distance image. Note that when a plurality of pixels of the camera image obtained by each camera 310 and a plurality of pixels of the distance image are associated with each other in advance, this association process may not be performed.
[0093] The determination unit 140 determines a search range to be used for searching for a plurality of similar points between a plurality of frames based on the subject information acquired from the storage unit 120 by the acquisition unit 130 without using the map information. The search range is a range in a three-dimensional space including a first three-dimensional point on the subject corresponding to a first point on the first frame. It can also be said that the search range is a range in the three-dimensional space where the first three-dimensional point is likely to exist. Further, the search range is a range in the shooting direction from the first viewpoint at which the first frame was shot.
[0094] Note that the search range is used for searching for a plurality of similar points between the first frame and a second frame in a range corresponding to the search range on the second frame different from the first frame among the plurality of frames. The second frame is a frame for which similar points are to be searched between it and the first frame. The search for similar points may be performed between a frame different from the first frame among the plurality of frames. That is, the frame selected as the second frame is not limited to one frame and may be a plurality of frames.
[0095] For example, the determination unit 140 may estimate the position of the first three-dimensional point based on the first distance information of the pixel corresponding to the first point in the first distance image included in the first frame, and determine the search range based on the estimated position of the first three-dimensional point. For example, the determination unit 140 may determine, as the search range, a range within a predetermined distance or less from the estimated position of the first three-dimensional point. Further, in order to estimate the position of the first three-dimensional point with higher accuracy, the determination unit 140 may select one or more second frames having distance information corresponding to the position of the first three-dimensional point from among a plurality of frames other than the first frame. That is, the determination unit 140 may estimate the position of the first three-dimensional point based not only on the first distance information but also on the second distance information of the pixel corresponding to the first point in the second distance image included in the second frame. Further, the determination unit 140 may determine a plurality of second frames from among a plurality of frames, and determine the search range based on the plurality of second distance information of the plurality of pixels respectively corresponding to the first point in the plurality of second distance images of the determined plurality of second frames. At this time, the determination unit 140 may determine the search range based on the plurality of second distance information without using the first distance information. Thus, the determination unit 140 may determine the search range based on one or more distance information of one or more pixels respectively corresponding to the first point in one or more distance images among the plurality of distance images. Note that when the determination unit 140 estimates the position of the first three-dimensional point using the second distance information, the determination unit 140 estimates the position of the first three-dimensional point using the converted second distance information obtained by converting the second distance information into the coordinate system of the first frame. The coordinate system conversion is performed based on the position and orientation of the camera from which the distance information before conversion was obtained and the position and orientation of the camera from which the frame to be converted was obtained. The one or more distance information may include the first distance information or the converted second distance information.
[0096] For example, if the distance between the position of the first three-dimensional point estimated based on the first distance information and the position of the first three-dimensional point estimated based on the second distance information after conversion to the coordinate system of the first frame is less than a predetermined threshold, the determination unit 140 may estimate the midpoint of the two positions as the position of the first three-dimensional point. If the above distance is greater than or equal to the predetermined threshold, the determination unit 140 may not estimate the position of the first three-dimensional point that is the reference of the search range.
[0097] The determination unit 140 may specify, as one or more pieces of distance information, the distance information in which the difference between the maximum value and the minimum value among a plurality of pieces of distance information corresponding to the first points in a plurality of distance images is less than the first value. The determination unit 140 may estimate the representative value of the one or more pieces of distance information as the position of the first three-dimensional point. That is, the determination unit 140 may determine the search range based on the representative value of the one or more pieces of distance information. The representative value is, for example, an average value, a median value, a maximum value, a minimum value, or the like. Note that if the variation of the one or more pieces of distance information is large, the determination unit 140 may not estimate the position of the first three-dimensional point. That is, in this case, the determination unit 140 may not determine the search range. The variation may be indicated by, for example, the variance or the standard deviation of the one or more pieces of distance information. The case where the variation of the one or more pieces of distance information is large means, for example, the case where the variance of the one or more pieces of distance information is larger than a predetermined variance, or the case where the standard deviation of the one or more pieces of distance information is larger than a predetermined standard deviation. Note that the determination unit 140 specifies one or more pixels corresponding to the first point using the positions and postures of the plurality of cameras 310 obtained by camera calibration.
[0098] Next, the process of selecting a target frame for acquiring distance information corresponding to the position of the first three-dimensional point from a plurality of frames will be described with reference to FIGS. 4A to 4E. FIGS. 4A to 4E show two subjects 510 and cameras 311, 312, and 313. The cameras 311, 312, and 313 are included in the plurality of cameras 310. Here, the camera 312 is a camera that generates the first frame, which is a reference frame for searching for similar points, and the cameras 311 and 313 are other cameras.
[0099] FIG. 4A is a diagram for explaining a first example of a process of selecting a target frame.
[0100] The determination unit 140 may select, as the target frame, a frame captured by the camera 311 whose difference in shooting pose from the first position and pose at the time of shooting of the camera 312 that captured the first frame is included in a first range. The target frame in the first example is also the second frame. For example, as shown in FIG. 4A, the determination unit 140 may select, as the target frame, a frame captured by the camera 311 that shoots in the shooting direction D1 whose difference θ from the shooting direction D2 of the camera 312 is included in the first range. The first range may be determined as a range having a common field of view with the camera 312. That is, the first range may be determined as a range in which the number of feature points between the first camera image of the first frame and the first camera image is equal to or greater than a first number in feature point matching in camera calibration. For example, the first number may be a value greater than 1. In this way, the second camera image of the second frame as the target frame may be determined based on the difference in shooting pose calculated from the first position and pose at the time of shooting of the camera that captured the first camera image of the first frame and the second position and pose of the camera that captured the second camera image.
[0101] FIG. 4B is a diagram for explaining a second example of a process of selecting a target frame.
[0102] As shown in FIG. 4B, the determination unit 140 may select, as the target frame, a frame captured by the camera 311 that satisfies the condition that the angular difference between the normal direction D11 of the surface of the subject at an arbitrary point on the subject and the direction D12 to the arbitrary point is included in a second range. In this case, the determination unit 140 does not necessarily have to select the first frame as the target frame. The second range may be defined as an angular range in which the distance image sensor 320 can accurately detect the distance to the surface of the subject.
[0103] FIG. 4C is a diagram for explaining a third example of a process of selecting a target frame.
[0104] The determination unit 140 may select, as a target frame, a frame captured by the camera 311 whose second position and orientation is such that the difference in the shooting position between the first position and orientation at the time of shooting of the camera 312 that captured the first frame is included in the third range. The target frame in the third example is also the second frame. For example, as shown in FIG. 4C, the determination unit 140 may select, as the target frame, frames captured by the cameras 311 and 313 that shoot at positions where the difference ΔL in the distance from the position of the camera 312 is included in the third range. The third range may be determined as a range having a common field of view with the camera 312. That is, the third range may be determined as a range in which the number of feature points between the first camera image of the first frame is equal to or greater than the first number in the feature point matching in camera calibration. For example, the first number may be a value greater than 1. In this way, the second camera image of the second frame as the target frame may be determined based on the difference in the shooting position calculated from the first position and orientation at the time of shooting of the camera that captured the first camera image of the first frame and the second position and orientation of the camera that captured the second camera image.
[0105] FIG. 4D is a diagram for explaining a fourth example of the process of selecting a target frame.
[0106] The determination unit 140 may select, as the target frame, a frame captured by the camera 311 at a second position and orientation that is separated from the subject 510 by a distance such that the difference between the distance at the time of capturing the first frame by the camera 312 and the subject 510 is included in a fourth range. The target frame in the fourth example is also the second frame. For example, as shown in FIG. 4D, the determination unit 140 may select, as the target frame, frames captured by the cameras 311 and 313 that capture at positions separated from the subject by distances L11 and L13 such that the differences from the distance L12 between the camera 312 and the subject are included in a third range. The fourth range may be determined as a range having a common field of view with the camera 312. That is, the fourth range may be determined as a range in which the number of feature points between the first camera image of the first frame is equal to or greater than a first number in feature point matching in camera calibration. For example, the first number may be a value greater than 1.
[0107] FIG. 4E is a diagram for explaining a fifth example of the process of selecting a target frame.
[0108] As shown in FIG. 4E, the determination unit 140 may select, as the target frame, a frame in which the area where the subject included in the first frame is captured repeatedly is large. For example, the determination unit 140 may select, as the target frame, a frame having a second number or more of distance information corresponding to distance information whose difference from the first distance information at the first point in the first frame is equal to or less than a fifth value. The distance information corresponding to the distance information whose difference from the first distance information is equal to or less than the fifth value is distance information whose coordinate system is converted to the coordinate system of the first frame by projecting the distance information corresponding to the first point of the frame onto the first frame. Note that the distance information whose difference from the first distance information is equal to or less than the fifth value is referred to as distance information overlapping with the first distance information. Note that the determination unit 140 may compare the position and orientation and the angle of view of the camera 312 with the position and orientation and the angle of view of the cameras 311 and 313, and select, as the target frame, a frame captured by a camera in which the overlapping area of the shooting ranges exceeds a predetermined size.
[0109] Note that the first range, the third range, and the fourth range are determined as ranges in which the number of feature points between the first camera image of the first frame and the feature point matching in camera calibration is equal to or greater than the first number. Therefore, it can be said that the target frame is determined from a plurality of camera images based on the number of feature points between the first camera image in feature point matching.
[0110] Note that the positions and postures of the cameras 311 to 313 used in the process of selecting the target frame by the determination unit 140 are specified by the camera parameters obtained by camera calibration.
[0111] Note that if the determination unit 140 satisfies the conditions for selecting the target frame described in the first to fifth examples, it may select a plurality of target frames. In this case, the determination unit 140 may set a priority order for the plurality of target frames that satisfy the conditions, and select the target frames with the third number as the upper limit in descending order of priority. The third number is a number determined so that the load required for the similarity point search process between the first frame and the first frame is equal to or less than a predetermined load. The priority order may be determined in descending order such that the closer the position of the camera that captured the first frame is, the higher the priority order, or the closer the direction of capturing an arbitrary point of the subject is to the normal direction at the arbitrary point, the higher the priority order, or the closer the distance from the subject is to the distance from the position of the camera that captured the first frame to the subject, the higher the priority order.
[0112] The determination unit 140 determines a search range based on representative values of one or more distance information on a straight line passing through the first viewpoint and the first three-dimensional point. Specifically, the determination unit 140 determines, as the search range, a range with a predetermined size centered at a position separated from the camera 312 by the distance indicated by the representative value on the straight line. Further, specifically, the determination unit 140 projects the distance information of the target frame corresponding to each pixel of the first frame and the first frame onto the first frame, and based on the obtained distance information, obtains the distance from the first viewpoint to the point corresponding to the position of each pixel in the first frame on the subject, and determines the size of the search range according to the obtained distance. The search range is a search range for searching for points similar to the points of each pixel of the first frame from a second frame different from the first frame.
[0113] Regarding the search range determined for each of the plurality of pixels in the distance image, the determination unit 140 may widen the search range as the accuracy of the one or more distance information for estimating the position of the first three-dimensional point is lower. Specifically, the determination unit 140 may determine that the accuracy of the one or more distance information is higher as the number of the one or more distance information is larger. The one or more distance information are similar to each other, that is, it is highly likely that the values are within a predetermined range. Therefore, it can be determined that the accuracy is higher as the number of the one or more distance information is larger. Further, the determination unit 140 may determine that the accuracy of the one or more distance information is higher as the variance of the one or more distance information is smaller. Since it can be determined that the one or more distance information are similar to each other when the variance is small, it can be determined that the accuracy of the one or more distance information is higher as the variance of the one or more distance information is smaller. Note that it may be determined that the accuracy of the distance information is higher as the reflectance when the distance information is obtained is larger.
[0114] Here, a case where the determination unit 140 determines the search range based on a plurality of second distance information without using the first distance information will be described with reference to FIGS. 5A and 5B.
[0115] FIG. 5A is a diagram for explaining problems in the case of using only the first distance information. FIG. 5B is a diagram showing an example of estimating the position of the first three-dimensional point using the second distance information. In FIG. 5A, three subjects 513 and a camera 312 are shown. In FIG. 5B, three subjects 513 and three cameras 311, 312, and 313 are shown. Here, the camera 312 is a camera that generates a first frame which is a reference frame for searching for similar points, and the cameras 311 and 313 are other cameras. Also, the thick solid line in FIG. 5A indicates a detection result with high accuracy by the distance image sensor of the camera 312, and the thick dashed line in FIG. 5A indicates a detection result with low accuracy by the distance image sensor. For example, a detection result with high accuracy (a detection result with higher accuracy than a predetermined accuracy) is a result detected when the reflectance is equal to or higher than a predetermined reflectance, and a detection result with low accuracy (a detection result with lower accuracy than a predetermined accuracy) may be a result detected when the reflectance is less than a predetermined reflectance. The reflectance is, for example, the ratio of the intensity of the emitted electromagnetic wave to the intensity of the acquired reflected wave.
[0116] As shown in FIG. 5A, the detection result by one camera 312 may include not only a detection result with high accuracy but also a detection result with low accuracy. Therefore, if the determination unit 140 estimates the position of the first three-dimensional point by adopting only the distance information obtained from the detection result with low accuracy, the accuracy of the estimated position of the first three-dimensional point becomes low, and a position different from the actual position may be estimated as the position of the first three-dimensional point.
[0117] On the other hand, as shown in FIG. 5B, the detection results by the three cameras 311, 312, and 313 are likely to include a detection result with high accuracy. This possibility increases as the number of cameras increases. Therefore, in FIG. 5A, instead of the distance information in the pixel including the detection result with low accuracy, high-accuracy detection calculated using the detection results of the cameras 311 and 313 other than the camera 312 that generated the detection result can be interpolated.
[0118] For example, the determination unit 140 determines whether the accuracy of the first distance information of the first point detected by the camera 312 is lower than a predetermined accuracy. When it is determined that the accuracy of the first distance information is lower than the predetermined accuracy, the distance information of the first point may be interpolated by replacing the first distance information with the third distance information. The third distance information is the distance information corresponding to the first point and is calculated using two camera images captured by the cameras 311 and 313. The determination unit 140 associates two pixels corresponding to the first point in the two camera images captured by the cameras 311 and 313, and based on the two pixels and the positions and postures of the cameras 311 and 313 respectively, calculates the position of the first point by triangulation, and may calculate the third distance information based on the position of the first point. In this way, when the detection accuracy of the first distance information of the pixel corresponding to the first point in the first distance image is lower than the predetermined accuracy, the determination unit 140 may use the third distance information corresponding to the first point, which is calculated using two or more camera images other than the first camera image, as one or more distance information to determine the search range. Note that when the accuracy of the generated third distance information is lower than the predetermined accuracy, the determination unit 140 may change the first frame used as the reference for searching for similar points to another frame. That is, after the frame is changed, the search for similar points is performed between the changed first frame and the frames other than the changed first frame.
[0119] Further, for example, the determination unit 140 determines whether the accuracy of the first distance information of the first point detected by the camera 312 is lower than a predetermined accuracy. When it is determined that the accuracy of the first distance information is lower than the predetermined accuracy, the distance information of the first point may be interpolated by replacing the first distance information with the first conversion information. The first conversion information is distance information obtained by performing coordinate conversion so as to project the second distance information of the first point detected by the camera 311 onto the detection result of the camera 312. Further, in this case, the determination unit 140 replaces the distance information calculated using the second conversion information obtained by performing coordinate conversion so as to project the second distance information of the first point detected by the camera 313 onto the detection result of the camera 312 and the first conversion information, thereby interpolating the distance information of the first point. In this way, when the detection accuracy of the first distance information of the pixel corresponding to the first point in the first distance image is lower than the predetermined accuracy, the determination unit 140 may use the second distance information of the pixel corresponding to the first point in the second distance image as one or more distance information to determine the search range. When interpolating, distance information determined to have high accuracy is used for calculating the distance information to be replaced.
[0120] The generation unit 150 generates a three-dimensional model of the subject based on the plurality of frames acquired from the storage unit 120 by the acquisition unit 130, the camera parameters, and the search range. The generation unit 150 searches for similar points similar to the first point on the first frame in a range corresponding to the search range on another frame (for example, the second frame) different from the first frame. The generation unit 150 restricts the epipolar line corresponding to the first point in the second frame to a length corresponding to the search range, and searches for similar points similar to the first point on the epipolar line of the second frame. The generation unit 150 searches for similar points from the second frame for each of the plurality of first pixels included in the first frame. As shown in Equation 1 below, the generation unit 150 calculates the Normalized Cross Correlation (NCC) between small regions as N(I, J) in combinations between the first frame and other plurality of frames excluding the first frame, and generates matching information indicating the result of performing frame matching.
[0121] Here, the merits of restricting the search range will be specifically described with reference to FIGS. 6 and 7. FIG. 6 is a diagram for explaining the matching process when there is no restriction on the search range. FIG. 7 is a diagram for explaining the matching process when there is a restriction on the search range.
[0122] As shown in FIG. 6, for one pixel 572 in the first frame 571, when matching is performed with the frame 581 within the unrestricted search range R1, in the frame 581, the epipolar line 582 corresponding to the straight line L1 passing through the first viewpoint V1 and the pixel 572 exists across the entire frame 581 from end to end. Note that the first frame 571 is an image obtained at the first viewpoint V1, and the frame 581 is an image obtained at the second viewpoint V2. The straight line L1 coincides with the shooting direction by the camera 311 at the first viewpoint V1. The pixel 572 corresponds to the point 511 of the subject 510. Therefore, the search for pixels in the frame 581 similar to the pixel 572 is performed on the unrestricted epipolar line 582. Thus, when there are two or more pixels having features similar to the pixel 572 on the epipolar line 582, there may be a case where a pixel 583 corresponding to a point 512 different from the point 511 of the subject 510 on the frame 581 is erroneously selected as a similar point. As a result, the generation accuracy of the three-dimensional model decreases.
[0123] On the other hand, as shown in FIG. 7, by the processing of the determination unit 140, the search range R2 is determined to be a shorter search range than the search range R1 shown in FIG. 6. Therefore, for one pixel 572 in the first frame 571, matching is performed with the frame 581 within the limited search range R2. In the frame 581, the epipolar line 584 corresponding to the straight line L1 passing through the first viewpoint V1 and the pixel 572 becomes shorter than the epipolar line 582 in accordance with the search range R2. Therefore, the search for the pixels in the frame 581 similar to the pixel 572 is performed on the epipolar line 584 that is shorter than the epipolar line 582. Thus, the number of pixels having features similar to the pixel 572 can be reduced, and the possibility of determining the corresponding pixel 585 on the frame 581 as a similar point to the point 511 of the subject 510 can be increased. For this reason, the generation accuracy of the three-dimensional model can be improved. Also, since the search range can be narrowed, the processing time related to the search can be shortened.
[0124] The generation unit 150 generates a three-dimensional model by performing triangulation using the positions and postures of the respective cameras 310 and the matching information. Note that the matching may be performed for all combinations of two frames among a plurality of frames.
[0125]
Number
[0126] Note that I xy and J xy are the pixel values within the small regions of the frame I and the frame J. Also,
Number
Number
[0127] Then, the generation unit 150 generates a three-dimensional model using the search results in the matching. As a result, the generation unit 150 generates a three-dimensional model including a plurality of three-dimensional points that are more numerous and have a higher density than the plurality of three-dimensional points included in the map information.
[0128] The output unit 160 outputs the three-dimensional model generated by the generation unit 150. The output unit 160 includes, for example, a display device such as a display (not shown) and an antenna, a communication circuit, a connector, etc. for communicably connecting by wire or wirelessly. The output unit 160 outputs the integrated three-dimensional model to the display device, thereby causing the display device to display the three-dimensional model.
[0129] [Operation of the Three-Dimensional Model Generation Device] Next, the operation of the three-dimensional model generation device 100 will be described with reference to FIG. 8. FIG. 8 is a flowchart showing an example of the operation of the three-dimensional model generation device 100.
[0130] First, in the three-dimensional model generation device 100, the reception unit 110 receives a plurality of frames captured by the plurality of cameras 310 and the camera parameters of each camera 310 from the estimation device 200 (S101). Note that the reception unit 110 does not necessarily receive the plurality of frames and the camera parameters at the same timing, and may receive them at different timings. That is, the first acquisition step and the second acquisition step may be performed at the same timing or at different timings.
[0131] Next, the storage unit 120 stores the plurality of frames captured by the plurality of cameras 310 and the camera parameters of each camera 310 received by the reception unit 110 (S102).
[0132] Next, the acquisition unit 130 acquires subject information (a plurality of distance images) from the plurality of frames stored in the storage unit 120, and outputs the acquired subject information to the determination unit 140 (S103).
[0133] Based on the subject information acquired by the acquisition unit 130, the determination unit 140 determines a search range to be used for matching of a plurality of points among a plurality of frames (S104). Details of step S104 are omitted because they have been described in the explanation of the process performed by the determination unit 140.
[0134] Next, the generation unit 150 searches for similar points similar to the first point on the first frame within a range corresponding to the search range on the second frame (S105), and generates a three-dimensional model based on the search result (S106). Details of step S105 and step S106 are omitted because they have been described in the explanation of the process performed by the generation unit 150.
[0135] Then, the output unit 160 outputs the three-dimensional model generated by the generation unit 150 (S107).
[0136] [Effects, etc.] The three-dimensional model generation method according to the present embodiment acquires subject information including a plurality of positions on a subject in a three-dimensional space (S103), acquires a first camera image of the subject taken from a first viewpoint and a second camera image of the subject taken from a second viewpoint (S101), and is obtained by camera calibration performed by photographing the subject from a plurality of viewpoints including the first viewpoint and the second viewpoint with one or more cameras, and without using map information including a plurality of three-dimensional points each indicating a position on the subject in the three-dimensional space, determines a search range in the three-dimensional space including a first three-dimensional point on the subject corresponding to the first point of the first image based on the subject information (S104), performs matching to search for similar points similar to the first point within a range corresponding to the search range on the second image (S105), and generates a three-dimensional model using the search result in the matching (S106). The subject information is information different from map information obtained by camera calibration performed by photographing the subject from a plurality of viewpoints with one or more cameras and including a plurality of three-dimensional points each indicating a position on the subject in the three-dimensional space.
[0137] According to the three-dimensional model generation method, the search range is determined based on the subject information without using map information, and similar points similar to the first point on the first image are searched in the range corresponding to the search range on the second image restricted by the search range. In this way, based on the subject information, since the search for similar points is performed in a range where there is a high possibility of the existence of similar points, the search accuracy of the similar points can be improved, and the time required for the search process can be shortened. Therefore, the generation accuracy of the three-dimensional model can be improved, and the processing time of the three-dimensional model generation process can be shortened.
[0138] Also, for example, in the matching (S105), the epipolar line corresponding to the first point in the second camera image is restricted to a length corresponding to the search range, and similar points similar to the first point are searched on the epipolar line of the second camera image.
[0139] According to this, since similar points similar to the first point are searched on the epipolar line restricted to a length corresponding to the search range, the search accuracy of the similar points can be improved, and the time required for the search process can be shortened.
[0140] Also, for example, the subject information includes a distance image generated by measurement by the distance image sensor 320. The distance image has a plurality of pixels each having distance information indicating the distance from the distance image sensor 320 to the subject. In the determination, the search range is determined based on the distance information of the pixel corresponding to the first point in the distance image.
[0141] According to this, since the subject information includes a distance image having a plurality of pixels associated with the plurality of pixels of the first camera image, the distance information corresponding to the first point can be easily specified. Therefore, based on the specified distance information, the position of the first three-dimensional point can be estimated, and the search range can be determined accurately.
[0142] Further, for example, the subject information includes a plurality of distance images respectively generated by a plurality of distance image sensors 320. Each of the plurality of distance images has a plurality of pixels having distance information indicating the distance from the distance image sensor 320 that generated the distance image to the subject. The plurality of pixels included in each of the plurality of distance images are respectively associated with the plurality of pixels included in the camera image corresponding to the distance image having the plurality of pixels among the plurality of camera images. The plurality of camera images include a first camera image and a second camera image. In the determination, the search range is determined based on the distance information of one or more pixels respectively corresponding to a first point in one or more of the plurality of distance images.
[0143] According to this, since the subject information includes a plurality of distance images each having a plurality of pixels associated with the plurality of pixels included in the first camera image, the plurality of distance information corresponding to the first point can be easily specified. Since the plurality of distance information thus specified is distance information obtained from different viewpoints, even if some of the distance information includes detection errors, the influence of the detection errors can be reduced by utilizing other distance information. Therefore, based on one or more of the plurality of distance information, the position of the first three-dimensional point can be estimated more accurately, and the search range can be determined accurately.
[0144] Further, for example, in the determination, when the detection accuracy of the first distance information of the pixel corresponding to the first point in the first distance image is lower than a predetermined accuracy, the third distance information corresponding to the first point calculated using two or more camera images other than the first camera image is used as one or more of the distance information to determine the search range. Therefore, when the detection accuracy of the first distance information is low, the search range can be determined using the third distance information with high accuracy. Thus, the search range can be determined accurately.
[0145] Also, for example, the plurality of distance image sensors 320 respectively correspond in position and orientation to a plurality of cameras 310 each including one or more cameras. The plurality of distance images include a first distance image corresponding to a first camera image and a second distance image corresponding to a second camera image. In the determination, when the detection accuracy of the first distance information of the pixel corresponding to the first point in the first distance image is lower than a predetermined accuracy, the second distance information of the pixel corresponding to the first point in the second distance image is used as one or more distance information to determine the search range. Therefore, when the detection accuracy of the first distance information is low, the search range can be determined using the second distance information with high accuracy. Thus, the search range can be determined accurately.
[0146] Also, for example, the plurality of distance image sensors 320 respectively correspond in position and orientation to a plurality of cameras 310 each including one or more cameras. In the determination, one or more pixels respectively corresponding to the first point in one or more distance images are specified using the positions and orientations of the plurality of cameras obtained by camera calibration.
[0147] Therefore, one or more distance information can be specified using the positions and orientations of the plurality of cameras obtained by camera calibration.
[0148] Also, for example, one or more distance images include a first distance image corresponding to a first camera image and a second distance image corresponding to a second camera image. The second camera image is determined from a plurality of camera images based on the number of feature points between the second camera image and the first camera image in feature point matching in camera calibration.
[0149] According to this, based on the number of feature points, the second camera image to be the target of matching similar points with the first camera image is determined. Therefore, it is possible to specify the second distance image for specifying one or more distance information with a high probability of not including errors, that is, with high accuracy.
[0150] Further, for example, the second camera image is determined based on the difference in the shooting poses calculated from the first position and pose at the time of shooting the first camera image and the second position and pose at the time of shooting the second camera image.
[0151] According to this, based on the difference in the poses of the cameras, the second camera image that is the target of matching similar points with the first camera image is determined. Therefore, it is possible to identify a second distance image for identifying one or more distance information with a high probability of not including errors, that is, with high accuracy.
[0152] Further, for example, the second camera image is determined based on the difference in the shooting positions calculated from the first position and pose at the time of shooting the first camera image and the second position and pose at the time of shooting the second camera image.
[0153] According to this, based on the difference in the positions of the cameras, the second camera image that is the target of matching similar points with the first camera image is determined. Therefore, it is possible to identify a second distance image for identifying one or more distance information with a high probability of not including errors, that is, with high accuracy.
[0154] Further, for example, the difference between the maximum value and the minimum value of one or more distance information is less than a first value.
[0155] According to this, one or more distance information in which the difference between the maximum value and the minimum value is less than the first value is identified. Thereby, it is possible to identify one or more distance information with a high probability of not including errors, that is, with high accuracy.
[0156] Further, for example, in the determination, the search range is widened as the accuracy of one or more distance information is lower.
[0157] According to this, since the search range is widened as the accuracy of one or more distance information decreases, it is possible to determine the search range according to the accuracy.
[0158] Further, for example, the accuracy is higher as the number of one or more distance information is larger.
[0159] According to this, it can be determined that the higher the number of one or more distance information, that is, the larger the number of one or more similar distance information, the higher the accuracy of the one or more distance information. Therefore, the larger the number of one or more distance information, the narrower the search range can be narrowed down.
[0160] Also, the accuracy is higher as the variance of the one or more distance information is smaller.
[0161] According to this, it can be determined that the smaller the variance of the plurality of distance information, that is, the more similar the plurality of distance information, the higher the accuracy of the plurality of distance information. Therefore, the smaller the variance of the plurality of distance information, the narrower the search range can be narrowed down.
[0162] [Modification Example 1] The three-dimensional model generation system 410 according to this modification example will be described. In this modification example, a case where subject information different from the subject information described in the embodiment is used will be described. That is, the subject information used in this modification example is different from the distance image.
[0163] FIG. 9 is a block diagram showing a characteristic configuration of the three-dimensional model generation system according to Modification Example 1.
[0164] The three-dimensional model generation system 410 according to this modification example is mainly different from the three-dimensional model generation system 400 according to the embodiment in that the camera group 300 further includes a measuring instrument 321 and a sensor fusion device 210 is provided instead of the estimation device 200. The same components as those in the three-dimensional model generation system 400 according to the embodiment are denoted by the same reference numerals, and the description thereof is omitted.
[0165] FIG. 10 is a diagram showing an example of the configuration of the camera group.
[0166] As shown in FIG. 10, two cameras 310 included in the camera group 300 and the measuring instrument 321 are fixed and supported by a fixing member 330 so that their positions and postures are fixed relative to each other. An apparatus including two cameras 310 and the measuring instrument 321 whose relative positions are fixed is called a sensor apparatus. The two cameras 310 constitute a stereo camera. The two cameras 310 take pictures of images synchronously with each other and generate stereo images taken at the synchronized shooting times. The generated stereo images are given the shooting times (timestamps) at which they were taken. The stereo images are output to the sensor fusion device 210. The two cameras 310 may take a stereo video.
[0167] The measuring instrument 321 generates three-dimensional data by emitting electromagnetic waves and acquiring reflected waves obtained by reflection of the electromagnetic waves by a subject. Specifically, the measuring instrument 321 measures the time it takes for the emitted electromagnetic waves to be reflected by the subject and return to the measuring instrument 321 after being emitted, and uses the measured time and the wavelength of the electromagnetic waves to calculate the distance between the measuring instrument 321 and a point on the surface of the subject. The measuring instrument 321 emits electromagnetic waves in a plurality of predetermined radial directions from a reference point of the measuring instrument 321. For example, the measuring instrument 321 emits electromagnetic waves at a first angular interval around the horizontal direction and emits electromagnetic waves at a second angular interval around the vertical direction. Therefore, the measuring instrument 321 can calculate the three-dimensional coordinates of a plurality of points on the subject by detecting the distances to the subject in a plurality of directions around the measuring instrument 321. Thus, the measuring instrument 321 can calculate position information indicating a plurality of three-dimensional positions on the subject around the measuring instrument 321 and generate a three-dimensional model having the position information. The position information may be a three-dimensional point cloud including a plurality of three-dimensional points indicating a plurality of three-dimensional positions.
[0168] In this embodiment, the measuring device 321 is a three-dimensional laser measuring device having a laser irradiation unit (not shown) that irradiates laser light as electromagnetic waves and a laser light receiving unit (not shown) that receives the reflected light obtained when the irradiated laser light is reflected by the subject. The measuring device 321 scans the subject with laser light by rotating or swinging a unit including the laser irradiation unit and the laser light receiving unit around two different axes, or by installing a movable mirror (MEMS (Micro Electro Mechanical Systems) mirror) that swings around two axes on the path of the laser for irradiation or reception. Thereby, the measuring device 321 can generate a high-precision and high-density three-dimensional model of the subject. Here, the generated three-dimensional model is, for example, a three-dimensional model in the world coordinate system.
[0169] The measuring device 321 acquires a three-dimensional point cloud by line scanning. For this purpose, the measuring device 321 acquires a plurality of three-dimensional points included in the three-dimensional point cloud at different times respectively. That is, the measurement time by the measuring device 321 and the photographing times by the two cameras 310 are not synchronized. The measuring device 321 generates a three-dimensional point cloud that is dense in the horizontal direction and sparse in the vertical direction. That is, in the three-dimensional point cloud obtained by the measuring device 321, the interval between vertically adjacent three-dimensional points is wider than the interval between horizontally adjacent three-dimensional points. In the three-dimensional point cloud generated by the measuring device 321, the measurement time at which each three-dimensional point is measured is associated and assigned to the three-dimensional point.
[0170] The measuring device 321 has been exemplified as a three-dimensional laser measuring device (LiDAR) that measures the distance to the subject by irradiating laser light, but is not limited thereto, and may be a millimeter-wave radar measuring device that measures the distance to the subject by emitting millimeter waves.
[0171] Note that the two cameras 310 shown in FIG. 10 may be a part or all of the plurality of cameras 310 included in the camera group 300.
[0172] Next, the operation of the sensor fusion device 210 will be described with reference to FIG. 11.
[0173] FIG. 11 is a flowchart showing an example of the operation of the sensor fusion device 210 according to Modification 1.
[0174] The sensor fusion device 210 acquires a stereo video and a time-series three-dimensional point cloud (S201). The stereo video includes a plurality of stereo images respectively generated in time series.
[0175] The sensor fusion device 210 calculates the position and orientation of the sensor device (S202). Specifically, the sensor fusion device 210 calculates the position and orientation of the sensor device using the stereo video obtained by the sensor device and the stereo images and three-dimensional points generated at shooting times and measurement times within a predetermined time difference among the three-dimensional point cloud. Note that the coordinates serving as the reference for the position and orientation of the sensor device may be the camera coordinate origin of the left-eye camera of the stereo camera when using the stereo video, may be the coordinates of the rotation center of the measuring instrument 321 when using the time-series three-dimensional point cloud, or may be either the camera coordinate origin of the left-eye camera or the coordinates of the rotation center of the measuring instrument 321 when using both.
[0176] For example, as shown in FIG. 12, the sensor device may move to different positions at time t1 and time t2. FIG. 12 shows the positions of the left-eye camera 310 of the stereo camera at times t1 and t2 and the position of the measuring instrument 321 at time t1.
[0177] Then, as shown in FIG. 13, using the stereo images taken at times t1 and t2, a camera image integrated three-dimensional point cloud existing only at characteristic portions of the subject is generated.
[0178] Also, as shown in FIG. 14, by measuring with the measuring instrument 321 from time t1 to time t2, a time-series three-dimensional point cloud existing only at the portions scanned by the measuring instrument 321 is generated.
[0179] When the sensor fusion device 210 calculates the position and orientation of the sensor device using stereo video, it may calculate the position and orientation of the sensor device by Visual SLAM (Simultaneous Localization and Mapping) based on feature point matching between the stereo image and the time-series images.
[0180] The sensor fusion device 210 integrates the time-series three-dimensional point clouds using the calculated position and orientation (S203). The three-dimensional point cloud obtained by the integration is referred to as the LiDAR integrated 3D point cloud.
[0181] In step S202, when the sensor fusion device 210 calculates the position and orientation of the sensor device using the time-series three-dimensional point clouds, for example, it may calculate the position and orientation of the sensor device by NDT (Normal Distribution Transform) based on three-dimensional point cloud matching. Since the time-series three-dimensional point clouds are used, the LiDAR integrated 3D point cloud can be generated simultaneously when calculating the position and orientation of the sensor device.
[0182] Next, the case where the sensor fusion device 210 calculates the position and orientation of the sensor device using both the stereo video and the time-series three-dimensional point clouds in step S202 will be described. In this case, in advance, the camera parameters including the individual focal lengths, lens distortions, image centers of the left-eye camera and the right-eye camera of the stereo camera, and the relative positions and orientations of the left-eye camera and the right-eye camera are calculated by, for example, a camera calibration method using a checkerboard. Then, the sensor fusion device 210 performs feature point matching between the stereo images and also performs feature point matching between the temporally consecutive images of the left-eye image, and calculates the three-dimensional positions of the matching points using the in-image coordinates of the matched feature points (matching points) and the camera parameters. The sensor fusion device 210 executes this process for an arbitrary number of frames to generate the camera image integrated three-dimensional point cloud.
[0183] Then, the sensor fusion device 210 aligns the camera image integrated three-dimensional point cloud and the time-series three-dimensional point cloud acquired by the measuring instrument 321 by a method of minimizing a cost function, and generates subject information (S204).
[0184] As shown in Equation 2, the cost function consists of a weighted sum of two error functions.
[0185] cost = E1 + wx E2 (Equation 2)
[0186] As shown in FIG. 15, the first error function E1 in the cost function is the reprojection error when each three-dimensional point of the camera image integrated three-dimensional point cloud is reprojected into camera coordinates at two times. Camera parameters obtained in advance by camera calibration are used for the reprojection calculation. For any three-dimensional point at any time, this error is calculated and summed up.
[0187] The second error function E2 in the cost function is the result of calculating the distance from each three-dimensional point of the camera integrated three-dimensional point cloud to the time-series three-dimensional points of the surrounding measuring instrument 321 after converting it into the coordinate system of the time-series three-dimensional point cloud generated by the measuring instrument 321. Note that the transformation matrix between the two coordinate spaces may be calculated from the actual positional relationship between the left-eye camera and the measuring instrument 321.
[0188] For the three-dimensional points in the same time period as the error function E1, this error is calculated and summed up.
[0189] Using each element of the camera coordinates at two times and the transformation matrix from the camera coordinate system to the coordinate system of the measuring instrument 321 as variable parameters, the minimization process of the cost function is performed. The minimization may be performed by the least squares method, the Gauss-Newton method, the levenberg-marquardt method, or the like.
[0190] Note that the weight w may be the ratio of the number of time-series three-dimensional points obtained by the measuring instrument 321 to the number of three-dimensional points of the camera image integrated three-dimensional point cloud.
[0191] The transformation formula between the time-series camera position and orientation and the camera coordinate system and the measuring instrument coordinate system is determined by the minimization process. Using this, the time-series three-dimensional point clouds are integrated to generate a LiDAR-integrated three-dimensional point cloud as subject information.
[0192] The sensor fusion device 210 outputs the generated subject information to the three-dimensional model generation device 100.
[0193] In the three-dimensional model generation system 410 according to this modification example, the subject information is generated based on two or more types of sensor information. That is, the subject information is generated based on two or more types of sensor information that are different from each other. That is, subject information with reduced accuracy degradation due to detection errors can be obtained.
[0194] Also, in the three-dimensional model generation system 410, the two or more types of sensor information include a plurality of two-dimensional images obtained from a stereo camera and three-dimensional data obtained from a measuring instrument 321 that emits electromagnetic waves and acquires reflected waves reflected by the subject.
[0195] According to this, since the subject information is generated based on a plurality of two-dimensional images and three-dimensional data, for example, three-dimensional data in which the three-dimensional data is densified using a plurality of two-dimensional images can be obtained with high accuracy.
[0196] [Modification Example 2] In the three-dimensional model generation device 100 according to the above-described embodiment, the determination unit 140 determines the search range used for searching for a plurality of similar points between a plurality of frames based on subject information (for example, a distance image) without using map information. However, the present invention is not limited to this. The determination unit 140 may switch between a first method of determining the search range based on the distance image as described in the above embodiment and a second method of determining the search range based on map information according to the distance between the subject and the camera 310 that generates the first frame, and then determine the search range. For example, for each of the plurality of pixels constituting the first distance image included in the first frame among the plurality of frames, when the distance indicated by the distance information of the pixel is less than a predetermined distance (that is, when the subject and the camera 310 that generates the first frame are close), the determination unit 140 may determine the search range using the first method. When the distance between the subject and the camera 310 that generates the first frame is equal to or greater than the predetermined distance (that is, when the subject and the plurality of cameras 310 are far), the determination unit 140 may determine the search range using the second method. This is because when the distance between the subject and the plurality of cameras 310 is separated by a predetermined distance or more, the accuracy of the map information becomes higher than the accuracy of the distance image of the camera 310.
[0197] In the second method, for example, the determination unit 140 generates three-dimensional information of the subject by interpolating three-dimensional points where the subject is estimated to exist between a plurality of three-dimensional points included in the map information, and determines a search range based on the generated three-dimensional information. Specifically, the determination unit 140 estimates a rough three-dimensional position of the subject surface by filling (i.e., interpolating) between a plurality of three-dimensional points included in the sparse three-dimensional point group based on the map information with a plurality of planes, and generates the estimation result as an estimated three-dimensional model. For example, between a plurality of three-dimensional points included in the sparse three-dimensional point group may be interpolated by meshing the plurality of three-dimensional points. Next, the determination unit 140 estimates, for each of a plurality of pixels on the projection frame on which the estimated three-dimensional model is projected in the first frame, a three-dimensional position based on the first viewpoint from which the first frame was captured, which is the three-dimensional position on the subject corresponding to the pixel. Thereby, the determination unit 140 generates an estimated distance image including a plurality of pixels each including the estimated three-dimensional position. Then, the determination unit 140 estimates the position of the first three-dimensional point based on the generated estimated distance image and determines a search range based on the estimated position of the first three-dimensional point, in the same manner as in the first method.
[0198] (Other Embodiments) As described above, the three-dimensional model generation method and the like according to the present disclosure have been described based on the above-described embodiments, but the present disclosure is not limited to the above-described embodiments.
[0199] For example, in the above embodiment, it was explained that each processing unit included in the three-dimensional model generation device or the like is realized by a CPU and a control program. For example, the components of the processing unit may each be composed of one or more electronic circuits. Each of the one or more electronic circuits may be a general-purpose circuit or a dedicated circuit. The one or more electronic circuits may include, for example, a semiconductor device, an IC (Integrated Circuit), or an LSI (Large Scale Integration). The IC or LSI may be integrated on one chip or on a plurality of chips. Here, although it is called an IC or LSI, the name may change depending on the degree of integration, and it may be called a system LSI, a VLSI (Very Large Scale Integration), or a ULSI (Ultra Large Scale Integration). Also, an FPGA (Field Programmable Gate Array) programmed after the manufacture of the LSI can be used for the same purpose.
[0200] Furthermore, the general or specific aspects of the present disclosure may be realized by a system, a device, a method, an integrated circuit, or a computer program. Alternatively, it may be realized by a computer-readable non-transitory recording medium such as an optical disk, an HDD (Hard Disk Drive), or a semiconductor memory in which the computer program is stored. Also, it may be realized by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
[0201] In addition, forms obtained by applying various modifications that can be conceived by those skilled in the art to each embodiment, and forms realized by arbitrarily combining the components and functions in the embodiment without departing from the spirit of the present disclosure are also included in the present disclosure.
Industrial Applicability
[0202] The present disclosure can be applied to a three-dimensional model generation device or a three-dimensional model generation system, and can be applied to, for example, figure creation, terrain or building structure recognition, human behavior recognition, or free viewpoint video generation, etc.
Explanation of Signs
[0203] 100 Three-dimensional model generation device 110 Receiver 120 Storage unit 130 Acquisition unit 140 Decision unit 150 Generation unit 160 Output unit 200 Estimation device 210 Sensor fusion device 300 Camera group 310, 311, 312, 313 Cameras 320 Distance image sensor 321 Measuring instrument 330 Fixing member 400, 410 Three-dimensional model generation system 500, 510, 513 Subjects 511, 512 Points 520 Map information 531~533, 581 Frames 541~543 Feature points 571 First frame 572, 583, 585 Pixels 582, 584 Epipolar lines D1, D2 Shooting directions D11 Normal direction D12 Direction L1 Straight line L11, L12, L13 Distances R1, R2 Search ranges t1, t2 Times V1 First viewpoint V2 Second viewpoint
Claims
1. A three-dimensional model generation method executed by an information processing apparatus, comprising: acquiring subject information including a plurality of positions on a subject in a three-dimensional space; acquiring a first camera image of the subject taken from a first viewpoint and a second camera image of the subject taken from a second viewpoint; determining a search range in a three-dimensional space including a first three-dimensional point on the subject corresponding to a first point in the first camera image based on the subject information, without using map information obtained by camera calibration executed by causing one or more cameras to photograph the subject from a plurality of viewpoints including the first viewpoint and the second viewpoint and including a plurality of three-dimensional points each indicating a position on the subject in the three-dimensional space; performing matching to search for a similar point similar to the first point in a range corresponding to the search range in the second camera image; generating a three-dimensional model using a search result in the matching; wherein the first camera image and the second camera image are images obtained by photographing with the one or more cameras; and the subject information is generated by a sensor different from the one or more cameras. A three-dimensional model generation method.
2. In the matching, an epipolar line corresponding to the first point in the second camera image is limited to a length according to the search range, and a similar point similar to the first point is searched on the epipolar line in the second camera image. The three-dimensional model generation method according to claim 1.
3. The subject information includes a distance image generated by measurement using a distance image sensor, wherein the distance image has a plurality of pixels each having distance information indicating a distance from the distance image sensor to the subject, and in the determination, the search range is determined based on the distance information of the pixel corresponding to the first point in the distance image. The three-dimensional model generation method according to claim 1 or 2.
4. The subject information includes a plurality of distance images respectively generated by measurement using a plurality of distance image sensors, each of the plurality of distance images has a plurality of pixels each having distance information indicating a distance from the distance image sensor that generated the distance image to the subject, and the plurality of pixels included in each of the plurality of distance images are respectively associated with a plurality of pixels included in a camera image corresponding to the distance image having the plurality of pixels among the plurality of camera images. The plurality of camera images include the first camera image and the second camera image. In the determination, the search range is determined based on distance information of one or more pixels each corresponding to the first point in one or more of the plurality of distance images. The three-dimensional model generation method according to claim 1 or 2.
5. In the determination, when the detection accuracy of the first distance information of the pixel corresponding to the first point in the first distance image corresponding to the first camera image is lower than a predetermined accuracy, the third distance information corresponding to the first point calculated using two or more camera images other than the first camera image is used as the one or more distance information to determine the search range. The three-dimensional model generation method according to claim 4.
6. The plurality of distance image sensors respectively correspond in position and orientation to a plurality of cameras including the one or more cameras. In the determination, the one or more pixels each corresponding to the first point in the one or more distance images are specified using the positions and orientations of the plurality of cameras obtained by the camera calibration. The three-dimensional model generation method according to claim 4.
7. The one or more distance images include a first distance image corresponding to the first camera image and a second distance image corresponding to the second camera image. The second camera image is determined from the plurality of camera images based on the number of feature points between the second camera image and the first camera image in feature point matching in the camera calibration. The three-dimensional model generation method according to claim 6.
8. The second camera image is determined based on the difference in shooting poses calculated from the first position and orientation at the time of shooting by the camera that shot the first camera image and the second position and orientation at the time of shooting by the camera that shot the second camera image. The three-dimensional model generation method according to claim 6.
9. The second camera image is determined based on the difference in shooting positions calculated from the first position and orientation at the time of shooting by the camera that shot the first camera image and the second position and orientation at the time of shooting by the camera that shot the second camera image. The three-dimensional model generation method according to claim 6.
10. The difference between the maximum value and the minimum value of the one or more distance information is less than a first value. The three-dimensional model generation method according to claim 6.
11. In the determination, the wider the search range is as the accuracy of the one or more distance information is lower. The three-dimensional model generation method according to claim 5.
12. The higher the number of the distance information of 1 or more, the higher the accuracy The three-dimensional model generation method according to claim 11.
13. The higher the accuracy, the smaller the variance of the distance information of 1 or more The three-dimensional model generation method according to claim 11.
14. The subject information is generated based on two or more types of sensor information The three-dimensional model generation method according to claim 1 or 2.
15. The two or more types of sensor information include a plurality of two-dimensional images obtained from a stereo camera and three-dimensional data obtained from a measuring instrument that emits electromagnetic waves and acquires reflected waves reflected by the subject The three-dimensional model generation method according to claim 14.
16. A processor and A memory, comprising The processor uses the memory to Acquire subject information including a plurality of positions on a subject in three-dimensional space, Acquire a first camera image of the subject taken from a first viewpoint and a second camera image of the subject taken from a second viewpoint, Without using map information obtained by camera calibration executed by photographing a subject from a plurality of viewpoints with one or more cameras and including a plurality of three-dimensional points each indicating a position on the subject in three-dimensional space, based on the subject information, determine a search range in three-dimensional space including a first three-dimensional point on the subject corresponding to a first point in the first camera image, Perform matching to search for a similar point similar to the first point in a range corresponding to the search range on the second camera image, Generate a three-dimensional model using the search result in the matching, The first camera image and the second camera image are images obtained by photographing with the one or more cameras, The subject information is generated by a sensor different from the one or more cameras Three-dimensional model generation device.
17. A memory and A processor connected to the memory, comprising The processor Acquire a first camera image generated by photographing a subject in three-dimensional space from a first viewpoint and a second camera image generated by photographing the subject from a second viewpoint, Search for a second point similar to the first point in a search range on an epipolar line specified by projecting a straight line passing through the first viewpoint and the first point of the first camera image onto the second camera image, Generate a three-dimensional model of the subject based on the result of the search The search range is provided based on the position of a first three-dimensional point corresponding to the first point in the three-dimensional space. The first camera image and the second camera image are images obtained by photographing with one or more cameras. The position is a sensor that uses a reflected wave of an electromagnetic wave emitted toward the subject, and is obtained by a sensor different from the one or more cameras. Three-dimensional model generation device. **Claim 18** The position is obtained based on a distance image generated by a sensor that receives the reflected wave. The three-dimensional model generation device according to claim 17.
Citation Information
Patent Citations
Device and method for processing image and provision medium
JP1999265454A
Image management apparatus, image management method and program
JP2017130146A
Three-dimensional model distribution method, three-dimensional model receiving method, three-dimensional model distribution device, and three-dimensional model receiving device
WO2018123801A1
Three-dimensional model generation method and three-dimensional model generation device
WO2021193672A1