In-building location estimation system

The building interior position estimation system addresses the challenge of high memory usage and slow processing by generating reduced-data target images for indoor self-position estimation, achieving efficient and accurate position determination.

JP2025091319APending Publication Date: 2025-06-18TOYOHASHI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023206526
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing building interior position estimation systems face challenges in reducing memory resources and improving processing speed for high-precision self-position estimation indoors, especially when dealing with large environmental maps and complex matching processes.

Method used

A building interior position estimation system that generates target images as cylindrical panoramic images or circular images for each point inside a building, reducing data capacity and memory resources, and enables fast and accurate self-position estimation by comparing query images with these target images.

Benefits of technology

The system achieves reduced memory resource usage and faster processing times, enabling highly accurate self-position estimation of clients or moving bodies within a building, while maintaining low-cost imaging means and reducing the weight and size of autonomous mobile robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025091319000001_ABST
    Figure 2025091319000001_ABST
Patent Text Reader

Abstract

To provide a system that can reduce memory resource by reducing data obtained from an environment map, and enables accurate estimation of the self-location of a mobile object in a short time.SOLUTION: An in-building location estimation system comprises: processing means 110 for generating target images at a plurality of arbitrary points inside a building from a three-dimentional image of the building; a database 130 that accumulates the generated target images; and determination means 120 for, upon input of a query image captured at a mobile body 300 that moves in the building, determining a location and a direction in which the query image was captured through comparison with a target image. The processing means comprises a target image generation unit 112 that uses an integer number of planar images at several points to generate a target image in the entire circumferential direction centered at an individual one of the points. The determination means comprises: a similarity calculation unit 121 that sequentially matches the query image to the entire circumferential direction of the objected image for all of the target images at the plurality of points, and calculates a similarity for each of all of the target images; and a determination unit 122 that determines the position and the orientation where the similarity is the maximum.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a building interior position estimation system. In particular, based on a 3D image of a building, a target image is generated on computer graphics, and by comparing it with an image taken at an arbitrary position by a device, a person, or other moving object moving within the building, the system is for estimating the photographed position (the position of the moving object) and the photographing direction (the orientation of the moving object).

Background Art

[0002] In order to move a moving object for unmanned operation typified by an autonomous mobile robot, a highly accurate self-position estimation technique is required. Also, when a person moves in a vast area, a self-position estimation technique can be used to grasp the person's position. Currently, as the most widespread position estimation technique, GNSS (Global Navigation Satellite System) that uses signals from artificial satellites can be mentioned. However, in the case of GNSS, there is a problem that the position estimation accuracy indoors decreases because signals are blocked by obstacles such as building walls.

[0003] Therefore, a technique called SLAM (Simultaneous Localization and Mapping) is sometimes used for autonomous mobile robots that operate exclusively indoors. This SLAM is a method that uses LiDAR (Light Detection And Ranging) and cameras to simultaneously create an environmental map and estimate self-position, and as represented by a cleaning robot, etc., it was possible to achieve highly accurate position estimation in an indoor environment (see Non-Patent Document 1).

[0004] However, the realization of SLAM technology requires expensive sensors, and also requires an expensive processing device for the calculation of creating an environmental map and estimating self-position, so the cost per unit in the case of an autonomous mobile robot has to be high.

[0005] As a technique for estimating one's own position after preparing an environmental map in advance, a technique called C * has been developed. This technique attempts to estimate its own position by comparing a virtual image created from a color point cloud map with an image obtained by a monocular camera (see Non-Patent Document 2). However, in the case of this method, it is difficult to realize global self-position estimation, and high-precision initial value setting is required for self-position estimation, so there are still problems in terms of the scope of use and applications.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0007] Therefore, the inventors of the present application proposed a cloud-based position estimation infrastructure system (which the inventors call Universal Map or UMap), where a 3D map as an environmental map is managed by a server, and an image captured on the client side (mobile body side) is transmitted to the server side as query data, so as to locate within the 3D map on the server side. The final result is transmitted to the client side, enabling the client to confirm the self-position estimation result (see Non-Patent Document 3 and Non-Patent Document 4). According to this technology, by using a simple and low-cost imaging means for acquiring images, it is excellent in that low-cost and high-precision self-position estimation indoors can be realized.

[0008] However, in order to improve the accuracy of self-position estimation or realize self-position estimation over a wide range, a large amount of data related to environmental maps (3D maps) is required, so a large amount of memory resources is required. Also, it takes time for the matching process with the query data captured and transmitted on the client side (mobile body side).

[0009] The present invention has been made in view of the above points, and its object is to provide a system that reduces memory resources by reducing the data (target images) obtained from environmental maps and enables high-precision self-position estimation of a client (mobile body) in a short time.

Means for Solving the Problems

[0010] Therefore, the present invention provides a building interior position estimation system comprising: processing means for generating a target image at any plurality of points inside a building from a 3D image of the building; a database for storing the target images generated by the processing unit; and determination means for receiving an input of a query image taken by a moving body moving inside the building, comparing the query image with the target images, and determining the position and orientation at which the query image was taken. The processing means includes a target image generation unit that generates a target image in all circumferential directions centered on each of the plurality of points by a plurality of integral plane images at the plurality of points directed in a shooting direction assumed in advance for the moving body. The determination means includes a similarity calculation unit that sequentially matches the query image with all circumferential directions of the target images for all of the target images at the plurality of points and calculates a similarity for each of the target images; and a determination unit that specifies the position of the target image with the largest similarity among the similarities calculated for the target images at the plurality of points as the moving body position, and specifies the azimuth with the largest similarity among the similarities calculated for each angle of the target image as the moving body orientation.

[0011] According to the above configuration, the target images generated by the target image generation unit can be obtained as one image for each point as images in all circumferential directions at a plurality of points inside the building. Therefore, conventionally, when acquiring a target image, a point was specified by the x and y directions indoors, and at that point, a plurality of three-dimensional images including the yaw angle θ should have been generated as the target image. However, by excluding the yaw angle θ and specifying the point by two dimensions in the x and y directions, a target image can be obtained. As a result, the matching with the query image at one point, which was conventionally a large number according to the set angle of the yaw angle θ, can be achieved by matching one image. However, when specifying the orientation of the moving body, it is necessary to obtain the yaw angle θ. This can be achieved by specifying the position in the target image at the time of calculating the similarity and specifying the calculation result and the calculation position of the similarity in one target image, thereby enabling the calculation of the yaw angle θ.

[0012] Note that, as the image in the circumferential direction in the present invention, for each of a plurality of points, a single cylindrical panoramic image centered on the vertical direction at that point, or an image in which the ceiling image vertically upward from that point is circular is assumed.

[0013] Therefore, in the invention having the above configuration, the target image generation unit converts, for each of the individual points among the plurality of points, a plurality of planar images in the horizontal direction at that position into cylindrical partial images with a predetermined curvature whose center line is vertical, and generates a single cylindrical panoramic image generated by connecting the plurality of cylindrical partial images as the target image. The determination means includes an image deformation unit for deforming the query image into an image having the same curvature as the cylindrical panoramic image, and the similarity calculation unit may calculate the similarity while sequentially matching the deformed image changed by the image deformation unit with the cylindrical panoramic images for each of the plurality of points.

[0014] Alternatively, instead of the above configuration, the target image generation unit generates, for each of the plurality of points, a single square planar image related to the ceiling image in the vertically upward direction as the target image. The determination means includes an image deformation unit for performing an enlargement process or a reduction process so that the diagonal length of the query image is equal to the length of one side of the target image. The similarity calculation unit may calculate the similarity while sequentially matching the query image while rotating it 360 degrees with respect to the square planar images for each of the plurality of points.

[0015] In any case of the configuration, the target image can be obtained as a single image in the circumferential direction centered on its position for each of the plurality of points. If the query image taken by the moving body is taken in a similar orientation, it can be matched for each point specified in the x direction and the y direction within the building. Note that, at the time of matching, it is rotated 360 degrees for each point to perform matching for each azimuth.

[0016] Furthermore, in the invention of each of the above configurations, the processing means, the database, and the determination means are stored in a server, and the moving body is provided with a client that can access the server. The client is configured to transmit the captured image to the server and receive the determination result by the determination means.

[0017] In such a configuration, the server further includes an image processing unit, the query image is generated by the image processing unit based on the real image transmitted from the moving body, and the image processing unit detects the boundary line of the planar portion of the subject captured in the real image and converts this boundary line into line segments, thereby generating the query image composed only of line segments.

[0018] According to the above configuration, the moving body only needs to be provided with an inexpensive imaging means such as a monocular camera and the client can transmit the query image (the previous real image) to the server. Therefore, it is not necessary to move a large-scale device for data processing. Thus, when the moving body is an autonomous mobile robot, the weight and size of the moving body can be reduced, and when the moving body is a person or an object operated by a person, the movement becomes easy. On the other hand, although various data are processed in the server, reducing the data for the target image can reduce the memory resources and increase the processing speed.

Advantages of the Invention

[0019] According to the present invention, since the target image obtained from the environmental map is generated as one image data such as a cylindrical panoramic image or a circular image for each of a plurality of points, the data capacity of the entire target image inside the building can be reduced. As a result, the memory resources can be reduced, and highly accurate self-position estimation of the client (moving body) can be performed in a short time.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0021] <Outline of Configuration> Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 schematically shows the configuration of this embodiment. As shown in this FIG. 1, the in-building position estimation system of this embodiment is basically divided into a server and a client. An image (query image or real image before conversion) captured at the client is transmitted to the server, necessary processing is performed at the server, and the result information of the estimated position is returned to the client again. The server is provided with a processing means and a determination means necessary for processing. The processing means includes a target image generation unit that generates a target image, and the determination means includes a similarity calculation unit that calculates the similarity when the target image and the query image are matched, and selects the target image with the highest similarity from the calculation results, and selects the angle with the highest similarity from among the target images, and a determination unit that specifies the position and azimuth at which the query image was taken. It has a configuration including each.

[0022] FIG. 2 shows a configuration centered on the server. As shown in this FIG. 2, the server 100 is provided with a processing means 110 and a determination means 120, and further has a database 130. The processing means 110 is exclusively composed of a DBIG (Data Base Image Generator), reads three-dimensional (3D) CAD data (CAD model) related to a desired building stored in an external (or within the server) data storage unit 200, and can generate a perspective projection image at an arbitrary position. Specifically, the CAD data is converted into an image composed only of line segments by the 3D wireframe processing unit 111, and based on that data, an image at an arbitrary position is further generated by the target image generation unit 112. The target image generated by the processing means 110 is stored in the database 130 of the server 100 and is used for matching with the query image.

[0023] The determination means 120 is provided with a similarity calculation unit 121 that performs matching between the query image and the target image and calculates the similarity. The calculation result is used by the determination unit 122 to select the target image with the highest similarity, thereby specifying (estimating) the position and direction in which the query image was taken. As the determination means, for example, a graphics board manufactured by NVIDIA Corporation equipped with a GPU (Graphics Processing Unit) can be used. For calculating the similarity, for example, CUDA (Compute Unified Device Architecture) manufactured by the same company can be used. The similarity calculation method can be quantified using a similarity function (see Non-Patent Document 4).

[0024] The client 300 that can be connected (wired or wireless) to the above server 100 is a moving body that moves within a real building. By transmitting the image taken by the imaging means (e.g., a monocular camera) 301 to the server 100, it is converted into an image composed only of line segments (wireframe processed image) by the image processing unit (wireframe processing unit) 101 provided in the server 100, and the query image and the target image composed only of the same type of line segments are subjected to matching.

[0025] Note that the server 100 and the client 300 can transmit and receive image data and estimation information via input / output interfaces 102 and 302, respectively. In addition to the above, the client (moving body) 300 can be equipped with moving means (drive unit, etc.) 303, a display device 304, etc. as required.

[0026] <Generation of the target image> As described above, the target image is generated as a perspective projection image at a plurality of points (arbitrary positions) by the target image generation unit 112. In this embodiment, the target image is generated as an image in the full circumferential direction centered on an arbitrary point. For example, for each of the plurality of points, it is assumed to be a single cylindrical panoramic image centered on the vertical direction at that point, or a ceiling image vertically upward from that point is circular.

[0027] Therefore, this embodiment will be described by exemplifying a cylindrical panoramic image. In generating a cylindrical panoramic image, a planar image with a predetermined viewing angle at an arbitrary position can be obtained from CAD data. At this time, in order to align the planar image with the shooting direction of the query image (the image captured by the shooting means), the horizontal direction is set at the same height as the shooting means (such as a monocular camera). Further, by sequentially connecting planar images in four orthogonal directions with a horizontal viewing angle (φ) of 45 degrees, an image of 360 degrees (full circumference direction) in the horizontal direction can be obtained. Note that the vertical viewing angle (ψ) of the image is set to match the vertical viewing angle of the query image.

[0028] The above planar image is subjected to wireframe processing on the planar image obtained from the CAD data, deformed into an image composed of line segments, and further deformed into an image (the image to be the object) projected onto a cylindrical curved surface (a cylindrical shape centered on a vertical line at a predetermined position). What should be deformed is to deform the pixel coordinates in the horizontal direction (lateral direction) into circular coordinates with a radius of 1 centered at 0. By deforming into a cylindrical shape in this way, a cylindrical panoramic image can be obtained.

[0029] Then, in order to enable comparison between the above cylindrical panoramic image and the query image, if it is assumed (if actually developed) that the cylindrical shape is cut along an arbitrary generatrix (for example, a generatrix located in a predetermined direction) and the whole is developed, it can be recognized as an image in which the full circumference direction is displayed by a single continuous image of a rectangle. When generating an actually developed image, the developed image can be used as the final image to be the object. An example of an image in a state assuming such development (actually developed) is shown in FIG. 3.

[0030] The image illustrated in FIG. 3 has a predetermined direction (predetermined generatrix) at one end, and the entire circumference is displayed within the range reaching the other end. Therefore, when matching with the query image, the matching at that point can be completed by sequentially moving the query image within the range from the one end to the other end. In the matching in the state of the cylindrical panoramic image, the moving direction is not a linear movement in the horizontal direction but rotation at predetermined angles while performing the matching, and the matching is completed by a rotation of the entire circumference (360 degrees).

[0031] <Conversion to Cylindrical Panoramic Image> Next, the process of converting a planar image to a cylindrical panoramic image will be described. Here, since the following coordinate system, signs, and subscripts are used, they will be explained together. When imaging a horizontal plane (cross-section) as shown in FIG. 4, the pixel coordinates are in a coordinate system where the center of the image is set to "0", the upper right is "positive", and the lower left is "negative". The sign "w" represents the "number of pixels in the horizontal direction (horizontal direction in this embodiment)" of the image, and "h" represents the "number of pixels in the vertical direction (vertical direction in this embodiment)" of the image. Also, "ψ" represents the "vertical viewing angle" of the image, and "φ" represents the "horizontal viewing angle" of the image. The vector representation of "P" represents a two-dimensional vector indicating the pixel position in the image, where P row represents the row direction and P col represents the column direction. Similarly, the vector representation of "L" represents a two-dimensional vector indicating the pixel position in the coordinate system where the radius of the panoramic cylinder is normalized to 1, where L row represents the row direction and L col represents the column direction. The subscript " r " indicates that it is the "raw image" before panoramic conversion, and " p " indicates the "panoramic image" after conversion. Also, " d " indicates that it is the image (target image) stored in the database, and " q " indicates that it is the "query image". Note that each image is assumed to be an image that has undergone wireframe processing.

[0032] First, in order to match the query image with the target image, it is necessary to adjust the size of the target image to match the number of pixels and the vertical viewing angle of the query image. This is to minimize information loss when the query image is converted into a cylindrical image. That is, when there is a large difference between the two, the information to be compared during matching will be different from each other, which will cause errors in the subsequent similarity calculation results. Therefore, the size of the target image is determined based on the number of pixels and the vertical viewing angle of the query image. The reason for using the query image as a reference is that position estimation should be made based on the image taken at a specific position of the moving object during its operation (when in use).

[0033] Therefore, the vertical viewing angle (ψ rd ) before the panoramic transformation of the target image is determined. Note that the horizontal viewing angle (φ rd ) is set to 90 degrees (continuous in four directions to cover the entire circumference). The calculation formula for determining the vertical viewing angle (ψ rd ) is as follows. First, there is the following relationship between the horizontal viewing angle (φ rq ) and the vertical viewing angle (ψ rq ) of the query image.

[0034]

Equation

[0035] Here, in the coordinate system shown in FIG. 4, let the line segment AB in the figure be the number of pixels in the horizontal direction (w rq ) of the query image, and the line segment CD be the number of pixels in the horizontal direction (w rd ) of the target image before the panoramic transformation. Note that for any image, the center of the viewpoint (the position of the camera, etc.) is set to 0, and the center of the circular cross-section (the periphery of the cylindrical shape) is used. Assuming this situation, the vertical viewing angle (ψ rd ) of the target image can be calculated from the number of pixels in the vertical direction (h rq ) of the query image and the number of pixels in the vertical direction (h rd ) before the panoramic transformation by the following formula.

[0036]

Number

[0037] Next, the relationship between the two-dimensional vector of the pixel position of the image to be processed before panoramic conversion and the pixel position after conversion to a cylindrical panoramic image can be obtained by the following formula. Here, Fig. 5 shows a state assuming a state of projecting onto a cylindrical surface with the horizontal viewing angle (φ rd ) of the image to be processed being 90 degrees. The line segment AB in this Fig. 5 is the length in the state where the horizontal pixels of the image to be processed are aligned (the same as the number of horizontal pixels (w rd ), and the arc corresponding to 45 degrees of the cylindrical surface is defined as "w pd ". Since this "w pd " is also the same as the number of pixels, points A and B may not overlap on the circumference of the cylindrical surface when rounded to an integer. Therefore, in order to eliminate the influence of such quantization error, the length from the center point to point A (or B) is defined as "r d " to derive a relational expression. Note that the vector representation of L indicates the vector of the pixel position in the coordinate system where the radius of the panoramic cylinder is normalized to "1".

[0038]

Number

[0039] In this way, by converting the pixel position of the image to be processed into the pixel position on the normalized panoramic cylindrical surface and obtaining the images converted to the panoramic cylindrical surfaces in four directions for each horizontal viewing angle (φ rd = 90 degrees), a cylindrical panoramic image of 360 degrees (full circle) can be generated.

[0040] <Conversion of Query Image> On the other hand, for the query image, the planar image is also converted into the coordinate system of the panoramic cylinder normalized in the same way. Fig. 6 shows the horizontal viewing angle (φ rq) is not fixed in angle because it is determined by the angle of view of the imaging means (such as a monocular camera), and shows a state assuming a state of being projected onto a cylindrical surface. In FIG. 6, similar to the case of the target image, the line segment AB is defined as the number of pixels (w rq ) in the horizontal direction, and in order to eliminate the influence of quantization error, the length from the center point to point A (or B) is defined as "r q ". Also, an arc corresponding to the horizontal angle of view (φ rq ) of the cylindrical surface is defined as "w pq ", and a relational expression will be derived.

[0041]

Number

[0042] In this way, by converting the pixel positions of the query image input as a planar image into pixel positions on the normalized panoramic cylindrical surface, an image converted to a panoramic cylindrical surface with a horizontal angle of view (φ rq ) is obtained, enabling high-precision matching between the converted cylindrical panoramic image and the target image.

[0043] <Exemplification of Various Parameters> Here, the horizontal angle of view (φ rd ) of the target image (cylindrical panoramic image) is set to 90 degrees. For the query image, when the number of pixels in the horizontal direction (w rq ) is 320 pixels, the number of pixels in the vertical direction (h rq ) is 240 pixels, and the vertical angle of view (ψ rq ) is 52.5 degrees, the following table shows various parameters calculated.

[0044]

Table 1

[0045] <Position Estimation> Next, an estimation method for the position where the query image was taken will be described. When matching the query image against the target image, as described above, both the target image and the query image are converted into panoramic cylindrical surfaces, and the number of vertical pixels (h rd 、h rq ) are the same. Therefore, by overlapping them on the panoramic cylindrical surface and moving the pixels in the horizontal direction (substantially a rotational movement in the yaw angle direction), the query image can be matched to the entire area of the cylindrical panoramic image. At this time, by comparing the pixel positions of both, a numerical similarity can be obtained by calculating using a similarity function based on elements such as the overlapping degree between the line segment and the plane.

[0046] At this time, by moving the query image by a predetermined number of pixels in the horizontal direction with respect to the cylindrical panoramic image, it can be rotated in the yaw angle direction. When the angle of the yaw angle for a movement of 1 pixel is "dθ" degrees, this angle "dθ" can be expressed by the following formula.

[0047]

Equation

[0048] In the above, the shift angle (dθ) is calculated as the angle per 1 pixel unit. However, within the range where the matching accuracy is guaranteed, the matching speed can also be improved by setting the shift angle to an integer multiple. This can be achieved by appropriately changing the magnification according to the characteristics of the structure inside the building where the position estimation is performed and adjusting the matching speed.

[0049] In this way, by performing matching while shifting the query image converted to a cylindrical surface on the cylindrical panoramic image horizontally at predetermined angles (dθ), it becomes possible to calculate the omnidirectional similarity at intervals of the shifting angle (dθ) from a single cylindrical panoramic image. By sequentially performing this for a plurality of cylindrical panoramic images, the similarity is calculated for all of the target images at a plurality of planned points, and by determining the point with the highest similarity among them, it is estimated as the position and orientation at which the query image was taken.

Example

[0050] The experimental examples of the present invention are shown below. The experimental location was the fourth floor of Building (Building D) existing in the applicant of this application. As an example, 50 query images were taken, and position estimation was performed for each query image. These 50 query images were randomly taken within the range of the above experimental location x×y = 39.2 m × 1.2 m, θ = 0 to 360 degrees. As other shooting conditions, as the shooting means, the main camera of iPhone (registered trademark) 15 was used, fixed at a height of 1.395 m from the floor surface, and the roll angle and pitch angle were each fixed at 0 degrees. Also, for the query image, wireframe processing was performed, and an image composed only of line segments after conversion was used, and an example of the processed image is shown in FIG. 7.

[0051] In addition, as a processing device for position estimation, a personal computer equipped with an Intel Core (registered trademark) i9 processor 12900K manufactured by Intel Corporation as the CPU and an NVIDIA (registered trademark) RTX4090 manufactured by NVIDIA as the graphics board was used. Using this personal computer, the generation of the cylindrical panoramic image of the target image and the conversion of the query image into a cylindrical image were processed. Each parameter at this time was set to the same state as shown in Table 1 above. Also, as the target image, a group of line segment images with position information generated in advance from the 3D CAD data of the experimental site using DBIG for experimental purposes was used, and matching processing was performed by the GPU (Graphics Processing Unit) installed on the graphics board, and the similarity was calculated using the company's CUDA (Compute Unified Device Architecture).

[0052] On the other hand, as a comparative example, position estimation by a conventional method (see Non-Patent Document 4) was performed under the same conditions. The differences between the example and the comparative example are generally as follows. In the comparative example, instead of converting the target image into a cylindrical panoramic image, a plurality of planar images in a state facing in a plurality of azimuths at a predetermined position were prepared in advance at each point without conversion, and the query image was matched with the plurality of target images (planar images) in the state of the planar image, and the position and azimuth were estimated by selecting the target image with the maximum similarity. Therefore, the number of pixels (h rd , w rd ) in the vertical and horizontal directions of the target image is the number of pixels (h rq , w rq) It was generated in the same state as []. Also, for the target image, after generating a line segment image group by wireframe processing, an image obtained by performing gradient dilation processing (processing to dilate line segments to clarify the comparison target) was used. For the query image, only wireframe processing was performed. In addition, in order to enable comparison of processing speeds and the like between the examples and the comparative examples, the conditions were the same for images other than those to be matched. That is, the shooting location, the number of captured query images, and other shooting conditions were the same as those in the examples, and the same processing device for position estimation was used.

[0053] For both of the above, the position estimation accuracy is compared. To compare the position estimation accuracy, the cumulative relative frequency of the error norm will be used. The error norm (e) can be obtained by the following formula when the error of each axis is set to "e x " m, "e y " m, "e θ " degrees.

[0054]

Equation

[0055] The cumulative relative frequency graph of the error norm calculated after matching is shown in FIG. 8. In the graph, the horizontal axis is the cumulative error norm and the vertical axis is the accuracy (probability). The 80%ile error norm is shown by a dashed line in the figure. Also, the generation intervals (dx, dy, dθ) of the target images used in the experiment are shown in the following table (Table 2), and the number of target images, the capacity of the target images, and the average processing time required for position estimation of one query image are shown in the following table (Table 3).

[0056]

Table 2

[0057]

Table 3

[0058] From the above results, it is clear that the number and capacity of the target images used for matching are fewer in the examples than in the comparative examples, and the position estimation time per query image is also less. Specifically, the capacity of the target images has been reduced by 95.2%, and the processing time has been reduced by 63.2%. The position estimation accuracy based on the graph shown in Figure 8 was found to be a 6.7-fold improvement in accuracy with an 80%ile error. This result indicates that position estimation is being performed much more accurately and quickly than in the comparative examples.

[0059] Regarding the above examples, Figure 9 shows a graph of the error norm and similarity in the matching with all target images performed per query image (with the magnitude of the error norm on the horizontal axis), and Figure 10 shows a graph of the similarity for each angle (θ) among the target images (cylindrical panoramic images) with the maximum similarity (the horizontal axis ranges from -180 degrees to +180 degrees). For "similarity", the value calculated using a separate similarity evaluation function is used as an index.

[0060] According to Figure 9, a peak in similarity appears where the error norm is small, indicating that the relationship between the error norm and similarity is normal. Also, according to Figure 10, among the target images (cylindrical panoramic images) with the highest similarity, a peak in similarity appears at a single angle, and it can be estimated that the angular direction is the orientation of the query image. Therefore, as long as such a peak in similarity appears at one location, it is also possible to set the shift angle during matching to twice or more. In that case, it is obvious that the processing speed will improve dramatically.

[0061] <Parentheses> Since the embodiments and examples of the present invention are as described above, the target images obtained from the environmental map are generated individually as single image data as cylindrical panoramic images for each of a plurality of points and can be used for matching with the query image. Therefore, the number of target images themselves can be reduced, and as a result, the data volume of the entire target image inside the building can be reduced. Accordingly, memory resources can be reduced, and furthermore, highly accurate position estimation can be realized in a short time.

[0062] Note that although the embodiments and examples of the present invention are as described above, these are merely examples of the present invention, and the present invention is not limited to these embodiments and the like. Therefore, it may be possible to change the elements of the above embodiments or add other elements.

[0063] For example, in the above-described embodiments and examples, the target image is converted into a cylindrical panoramic image and compared with the query image. However, since the target image only needs to be generated as an image in the entire circumferential direction centered on the position to be estimated, it is not limited to a cylindrical panoramic image. Therefore, a single square planar image related to the ceiling image in the vertically upward direction may be used as the target image for the entire circumferential direction. In this case, for the matching, a ceiling image centered on the shooting location is also used for the query image. However, if a deformation process is performed such that the diagonal length of the query image is equal to the length of one side of the target image by the image deformation unit, when the angle of the query image is changed, the query image will not protrude from the target image, and matching can be performed with the planar image without performing circular processing. Through such matching, pixels will overlap for the entire query image. Therefore, in the collation with the target image, the similarity between the query image and a specific target image can be determined, and the specific image with the maximum similarity can be identified. Also, it is possible to identify the azimuth of the shooting position with the angle having the highest similarity among those specific images. Since the angle change in this case is also due to rotation in the yaw angle direction, it can be processed in the same way as in the case of a cylindrical panoramic image. Note that the case where position estimation is enabled by such a configuration may be assumed, for example, when there is a characteristic shape in the ceiling structure (such as a beam structure).

[0064] By the way, when only using the cylindrical panoramic image or the square planar image of the ceiling that should be the target image, there may be cases where the accuracy of position estimation is lacking (the similarity is maximized among multiple values). This is because in building structures, the same type of structure may be constructed at multiple locations. In such cases, the accuracy may be improved by combining the similarity determination using the cylindrical panoramic image and the similarity determination using the cylindrical image of the ceiling.

Explanation of Reference Numerals

[0065] 100 Server 101 Image Processing Unit 102 Input / Output Interface 110 Processing Means (DBIG) 111 3D Wireframe Processing Unit 112 Target Image Generation Unit 120 Judgment Means 121 Similarity Calculation Unit 122 Judgment Unit 130 Database 200 Data Storage Unit 300 Client 301 Photographing Means (Monocular Camera, etc.) 302 Input / Output Interface 303 Moving Means (Drive Unit, etc.) 304 Display Device

Claims

1. A building interior position estimation system comprising: processing means for generating a target image at any plurality of points inside the building from a 3D image of the building; a database for storing the target images generated by the processing unit; and determination means for receiving an input of a query image taken by a moving body moving inside the building, comparing the query image with the target images, and determining the position and orientation at which the query image was taken, The processing means includes a target image generation unit that generates a target image in all circumferential directions centered on each of the plurality of points by a plurality of integer planar images at the plurality of points in the shooting direction assumed in advance for the moving body. The determination means includes: a similarity calculation unit that sequentially matches the query image with all circumferential directions of the target images for all of the plurality of points and calculates the similarity for each target image; and a determination unit that specifies the position of the target image with the largest similarity among the similarities calculated for the target images at the plurality of points as the moving body position, and specifies the azimuth with the largest similarity among the similarities calculated for each angle of the target image as the moving body orientation. A building interior position estimation system characterized by the above.

2. The target image generation unit converts a plurality of planar images in the horizontal direction at each of the plurality of points into a cylindrical partial image with a predetermined curvature whose center line is vertical, and generates a single cylindrical panoramic image generated by connecting the plurality of cylindrical partial images as the target image. The determination means includes an image deformation unit for deforming the query image into an image with the same curvature as the cylindrical panoramic image. The similarity calculation unit calculates the similarity while sequentially matching the deformed image changed by the image deformation unit with the cylindrical panoramic images for each of the plurality of points. The building interior position estimation system according to Claim 1.

3. The target image generation unit generates, for each of the plurality of points, a single square planar image related to the ceiling image in the vertically upward direction as the target image. The determination means includes an image deformation unit for performing an enlargement process or a reduction process so that the diagonal length of the query image becomes equal to the length of one side of the target image. The similarity calculation unit calculates the similarity while sequentially matching the query image while rotating the query image 360 degrees for the square images for each of the plurality of points. The in-building position estimation system according to claim 1.

4. The processing means, the database, and the determination means are stored in a server, the moving body includes a client accessible to the server, and the client transmits the captured image to the server and receives the determination result by the determination means. The in-building position estimation system according to any one of claims 1 to 3.

5. The server further includes an image processing unit, and the query image is generated by the image processing unit based on the real image transmitted from the moving body. The image processing unit detects a boundary line of a planar portion of a subject photographed in the real image, and generates the query image composed only of line segments by converting this boundary line into line segments. The in-building position estimation system according to claim 4.