Object estimation device, position estimation system, object estimation method, and control program

The object estimation device addresses the challenge of accurately matching detected landmarks by using text-format information and arrangement analysis to reduce computational load, ensuring efficient object identification.

JP2026064170APending Publication Date: 2026-04-13TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing object estimation devices struggle to accurately determine whether detected landmarks correspond to predetermined landmarks, leading to increased computational load when multiple landmarks are extracted.

Method used

An object estimation device that acquires and converts text-format information about predetermined objects and their arrangement from a database, comparing it with detected objects to accurately estimate correspondence, reducing computational load by narrowing down likely matches.

Benefits of technology

Accurately estimates whether detected objects are predetermined objects, significantly reducing computational load by narrowing down potential matches, even when multiple objects are detected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064170000001_ABST
    Figure 2026064170000001_ABST
Patent Text Reader

Abstract

It accurately estimates whether an object detected from the surrounding environment is a specified object. [Solution] An object estimation device comprising: an acquisition unit that acquires text-format information related to a predetermined object and surrounding objects located around the object; an object detection unit that detects the object and its surrounding objects; a conversion unit that converts the detected object into a first feature quantity which is a feature quantity defined as data representing the object, and converts the acquired text-format information related to the predetermined object into a second feature quantity which is a feature quantity defined as data representing the object; a calculation unit that calculates first arrangement information which defines the arrangement relationship between the object detected by the object detection unit and the surrounding objects, and calculates second arrangement information which defines the arrangement relationship between the object acquired by the acquisition unit and the surrounding objects; and an object estimation unit that estimates whether the detected object is a predetermined object by comparing the first feature quantity and the second feature quantity, and by comparing the first arrangement information and the second arrangement information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an object estimation device, a position estimation system, an object estimation method, and a control program.

Background Art

[0002] Patent Document 1 discloses a self-position estimation device that can accurately recognize different landmarks when a plurality of landmarks are extracted and can estimate its own position with high accuracy.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The device disclosed in Patent Document 1 cannot accurately estimate whether the extracted landmark corresponds to a predetermined landmark stored in the storage means. Therefore, there is a problem that it is impossible to narrow down a small number of predetermined landmarks that are likely to be the extracted landmarks from a plurality of predetermined landmarks stored in the storage means. As a result, in the device disclosed in Patent Document 1, when a plurality of landmarks are extracted, the computational load required to determine the correspondence between the plurality of extracted landmarks and the plurality of predetermined landmarks stored in the storage means becomes large.

[0005] The present disclosure has been made in view of the above background, and an object thereof is to provide an object estimation device, a position estimation system, an object estimation method, and a control program that can accurately estimate whether an object detected from the surrounding environment is a predetermined object.

Means for Solving the Problems

[0006] The object estimation device according to this disclosure includes: an acquisition unit that acquires text-format information related to a predetermined object and surrounding objects located around the predetermined object from a database in which text-format information related to a predetermined object is stored in association with the location information of the object; an object detection unit that detects an object and surrounding objects that are objects around the predetermined object; and a unit that converts the object detected by the object detection unit into a first feature quantity which is a feature quantity defined as data representing the object, and converts the text-format information related to the predetermined object acquired by the acquisition unit into a second feature quantity which is a feature quantity defined as data representing the object. The object estimation device comprises: a conversion unit; a calculation unit that calculates first arrangement information defining the arrangement relationship between the object detected by the object detection unit and the surrounding objects, and calculates second arrangement information defining the arrangement relationship between the object acquired by the acquisition unit and the surrounding objects; and an object estimation unit that compares the first feature quantity representing the object detected by the object detection unit and the second feature quantity representing the predetermined object, and compares the first arrangement information calculated by the calculation unit and the second arrangement information to estimate whether the object detected by the object detection unit is the same as the predetermined object. This object estimation device can accurately estimate whether an object detected from the surrounding environment is the predetermined object by acquiring and referring to text-format information about the predetermined object and arrangement information of objects around the predetermined object. Therefore, this object estimation device can narrow down a number of predetermined objects from a plurality of predetermined objects to a number of predetermined objects that are highly likely to be objects detected from the surrounding environment. As a result, even when multiple objects are detected from the surrounding environment, this object estimation device can reduce the computational load required to determine the correspondence between those multiple objects and, for example, multiple predetermined objects registered in a map database.

[0007] The object estimation method according to this disclosure involves a computer obtaining text-format information related to a predetermined object and surrounding objects located around it from a database in which text-format information related to a predetermined object is stored in association with the object's location information; detecting the object and surrounding objects located around it; converting the detected object into a first feature, which is a feature defined as data representing the object; converting the obtained text-format information related to the predetermined object into a second feature, which is a feature defined as data representing the object; calculating first arrangement information that defines the arrangement relationship between the detected object and the surrounding objects; calculating second arrangement information that defines the arrangement relationship between the obtained object and the surrounding objects; comparing the first feature representing the detected object with the second feature representing the predetermined object; and comparing the calculated first arrangement information with the second arrangement information to estimate whether the detected object is identical to the predetermined object. This object estimation method can accurately estimate whether an object detected from the surrounding environment is a predetermined object by acquiring and referencing text-based information about a predetermined object and information about the arrangement of objects around that predetermined object. Therefore, this object estimation method can narrow down the number of predetermined objects that are likely to be detected from the surrounding environment from among multiple predetermined objects. As a result, even when multiple objects are detected from the surrounding environment, this object estimation method can reduce the computational load required to determine the correspondence between those multiple objects and, for example, multiple predetermined objects registered in a map database.

[0008] The control program according to this disclosure causes a computer to perform the following steps: acquire text-format information related to a predetermined object and surrounding objects located around the predetermined object from a database in which text-format information related to a predetermined object is stored in association with the location information of the object; detect an object and surrounding objects located around the predetermined object; convert the detected object into a first feature quantity which is a feature quantity defined as data representing the object, and convert the acquired text-format information related to the predetermined object into a second feature quantity which is a feature quantity defined as data representing the object; calculate first arrangement information which defines the arrangement relationship between the detected object and the surrounding objects, and calculate second arrangement information which defines the arrangement relationship between the acquired object and the surrounding objects; compare the first feature quantity which represents the detected object and the second feature quantity which represents the predetermined object, and estimate whether the detected object is the same as the predetermined object by comparing the calculated first arrangement information and the second arrangement information. This control program can accurately estimate whether an object detected from the surrounding environment is a predetermined object by acquiring and referencing text-formatted information about a predetermined object and information about the arrangement of objects around that predetermined object. Therefore, this control program can narrow down the number of predetermined objects from a group of predetermined objects to those that are most likely to be detected from the surrounding environment. As a result, even when multiple objects are detected from the surrounding environment, this control program can reduce the computational load required to determine the correspondence between those multiple objects and, for example, multiple predetermined objects registered in a map database. [Effects of the Invention]

[0009] This disclosure provides an object estimation device, a position estimation system, an object estimation method, and a control program that can accurately estimate whether an object detected from the surrounding environment is a predetermined object. [Brief explanation of the drawing]

[0010] [Figure 1]This figure shows an example configuration of the position estimation system according to Embodiment 1. [Figure 2] This is a flowchart showing the operation of the position estimation device according to Embodiment 1. [Figure 3] This figure shows an example of the contents registered in the map database. [Figure 4] This diagram shows the relationship between landmarks registered in a map database and objects detected from camera images. [Figure 5] This is a diagram illustrating an application example of the position estimation system according to Embodiment 1. [Modes for carrying out the invention]

[0011] The present invention will be described below through embodiments, but the claims are not limited to the following embodiments. Furthermore, not all of the configurations described in the embodiments are necessarily essential for solving the problem. For clarity of explanation, the following descriptions and drawings have been omitted and simplified as appropriate. In each drawing, the same elements are denoted by the same reference numerals, and redundant explanations have been omitted where necessary.

[0012] <Embodiment 1> Figure 1 shows an example configuration of the position estimation system 1 according to Embodiment 1. The position estimation system 1 is applied, for example, to an autonomous mobile robot and estimates the autonomous mobile robot's own position by matching objects detected from camera images, etc., with landmarks registered in a map database. Here, the position estimation system 1 can accurately estimate whether an object detected from camera images, etc., corresponds to a landmark registered in the map database by referring not only to information about landmarks detected from camera images, etc., but also to text-format information about landmarks obtained via an operating terminal, etc. Therefore, the position estimation system 1 can narrow down the number of landmarks that are likely to be objects detected from camera images, etc., from among multiple landmarks. As a result, even when multiple objects are detected from camera images, etc., the position estimation system 1 can reduce the computational load required to determine the correspondence between the multiple objects and the multiple landmarks registered in the map database. A detailed explanation follows below.

[0013] As shown in Figure 1, the position estimation system 1 comprises a position estimation device 10, an operating terminal 20, a camera 30, a map database 40, and a network 50. The position estimation device 10 can also be considered a position estimation system on its own. The position estimation device 10, the operating terminal 20, the camera 30, and the map database 40 are configured to communicate with each other via a wired or wireless network 50. In this embodiment, the case in which the position estimation system 1 is applied to an autonomous mobile robot will be described as an example.

[0014] The autonomous mobile robot moves from its current location to its destination while estimating its own position by comparing surrounding objects detected by the position estimation device 10 with landmarks, which are objects registered in the map database 40.

[0015] The position estimation device 10 comprises an object detection unit 11, an acquisition unit 12, a conversion unit 13, an object estimation unit 14, and a position estimation unit 15. The object estimation device is composed of the object detection unit 11, the acquisition unit 12, the conversion unit 13, and the object estimation unit 14. The object estimation unit 14 includes a calculation unit 16.

[0016] The object detection unit 11 detects objects around the autonomous mobile robot by analyzing images captured by a camera 30 attached to the autonomous mobile robot. A distance sensor or depth sensor may be used instead of the camera 30, but the camera 30 is suitable for use because it is lighter and less expensive than distance sensors or depth sensors. For example, the object detection unit 11 detects an object as a rectangular image portion surrounding an object in the image captured by the camera 30. In this case, the center of the rectangular image surrounding the object becomes the representative point of that object. Furthermore, the object detection unit 11 can improve the accuracy of object detection in captured images by using a trained model generated by machine learning using multiple captured images.

[0017] Here, the object detection unit 11 detects at least one object in the vicinity of the autonomous mobile robot positioned at the reference location, and registers information about the detected object (or its characteristic features) as landmark information (or its characteristic features) in the map database 40 (database). The landmark information includes information such as the shape, color, and pattern of the object that can be identified from the captured image. The landmark information also includes information such as the type of object that can be identified from the captured image. Furthermore, the landmark information also includes the position information of the object used to represent the reference position of the position estimation device 10 (autonomous mobile robot). Additional landmarks may be registered in the map database 40 as appropriate, provided that their positional relationship with already registered landmarks is clear.

[0018] The acquisition unit 12 acquires text-form information regarding landmarks registered in the map database 40 separately from the information regarding landmarks detected by the object detection unit 11. For example, when the landmark is a store, the acquisition unit 12 acquires detailed information regarding the landmark, such as the store name, in text form, and when the landmark is furniture or tableware, the specific type thereof, etc., in text form. The text-form information (or feature amount representing the same) regarding the landmark acquired by the acquisition unit 12 is registered in the map database 40 in association with the corresponding landmark.

[0019] The text-form information regarding the landmark acquired by the acquisition unit 12 is transmitted, for example, from the operation terminal 20.

[0020] The operation terminal 20 is a communicable terminal owned by the user or temporarily assigned to the user, and is, for example, a PC (Personal Computer) terminal, a mobile terminal such as a smartphone or a tablet terminal, or a dedicated communication terminal prepared for this system.

[0021] For example, the user operates the operation terminal 20 by touching the monitor of the operation terminal 20 with a touch pen or a finger, or by operating a mouse, keyboard, etc. of the operation terminal 20, and inputs text-form information regarding the landmarks registered in the map database 40 to the operation terminal 20. The operation terminal 20 receives the text-form information regarding the landmark and transmits it to the position estimation device 10 via the network 50.

[0022] Note that the acquisition unit 12 may acquire the text-form information regarding the landmark from an external device other than the operation terminal 20.

[0023] The conversion unit 13 converts information about objects detected by the object detection unit 11 (including objects registered as landmarks in the map database 40) into vector-defined feature quantities using a multimodal recognition model, such as an encoder that can accept both language and images as input. The conversion unit 13 also converts text-format information about landmarks acquired by the acquisition unit 12 into vector-defined feature quantities using a multimodal recognition model. This makes it possible to compare information about objects detected by analyzing captured images with information about objects acquired in text format. The calculation unit 16 then calculates first placement information that defines the placement relationship between each object detected by the object detection unit 11 and its surrounding objects. More specifically, the calculation unit 16 represents the objects detected by the object detection unit 11 and multiple surrounding objects located around them as multiple nodes, and represents them as a node-edge graph by connecting them with edges. Then, it constructs a semantic graph by connecting objects that are adjacent to each other beyond a certain distance threshold with edges. Then, for a node representing an object, the calculation unit 16 examines all possible paths, starting from the node representing that object and moving a fixed number of steps (e.g., 3 steps) to nodes representing adjacent surrounding objects connected by edges, and counts the patterns of object classes (types) traversed to construct a histogram (semantic histogram). By calculating the results obtained using a modeling method such as L2 normalization on the resulting histogram as descriptors, the calculation unit 16 calculates placement information (first placement information), which is information that models or quantifies the placement relationship between a given object and multiple surrounding objects. Similarly, the calculation unit 16 applies the same method to the placement relationships between objects and their surrounding objects acquired by the acquisition unit 12, thereby calculating second placement information, which is information that models or quantifies the placement relationship between an object and its surrounding objects. Such placement information (first and second placement information) can be represented as numerically computable vectors.

[0024] The object estimation unit 14 compares the object detected by the object detection unit 11 with landmarks registered in the map database 40 to estimate whether the object detected by the object detection unit 11 corresponds to a landmark registered in the map database 40.

[0025] Specifically, the object estimation unit 14 estimates whether the object detected by the object estimation unit 14 corresponds to a landmark registered in the map database 40 by comparing a feature quantity representing an object detected by the object detection unit 11 (first feature quantity) with a feature quantity representing a landmark registered in the map database 40 (second feature quantity).

[0026] More specifically, the object estimation unit 14 first estimates whether the object detected by the object estimation unit 14 corresponds to a landmark registered in the map database 40 by comparing a feature quantity (first feature quantity) representing the object detected by the object detection unit 11 with a second feature quantity representing a landmark registered in the map database 40. Here, the second feature quantity includes a feature quantity representing text-format information about the landmark obtained from the acquisition unit 12. The second feature quantity may also further include a feature quantity representing information about the landmark obtained from the captured image. In this case, the second feature quantity contains more detailed information about the landmark. Therefore, the object estimation unit 14 can accurately estimate whether the object detected by the object detection unit 11 corresponds to a landmark registered in the map database 40. Furthermore, the landmark information may use a feature quantity representing at least one of text and an image.

[0027] Furthermore, if multiple landmarks are registered in the map database 40, the object estimation unit 14 narrows down the multiple landmarks to those that are most likely to be objects detected by the object detection unit 11.

[0028] Specifically, the object estimation unit 14 compares a first feature representing an object detected by the object detection unit 11 with multiple reference feature quantities representing each of the multiple landmarks registered in the map database 40, and narrows down the landmarks of the second feature quantities whose degree of agreement with the first feature quantity is above a predetermined threshold as landmarks that are highly likely to be objects detected by the object detection unit 11. More specifically, the degree of agreement is calculated by calculating the dot product (cosine similarity) of the first feature quantity and the second feature quantity as the similarity. Alternatively, the object estimation unit 14 may compare a first feature representing an object detected by the object detection unit 11 with multiple second feature quantities representing each of the multiple landmarks registered in the map database 40, and estimate the landmark of the second feature quantity with the highest degree of agreement with the first feature quantity as a landmark that is highly likely to be an object detected by the object detection unit 11.

[0029] As already explained, the second feature includes not only a feature representing information about landmarks obtained from the captured image, but also a feature representing text-formatted information about landmarks obtained from the acquisition unit 12. In other words, the second feature contains more detailed information about landmarks. Therefore, the object estimation unit 14 can narrow down the number of landmarks from multiple landmarks to those that are highly likely to be objects detected by the object detection unit 11. Furthermore, the object estimation unit 14 can narrow down landmarks based on both the comparison result of the first feature and the second feature as described above, and the comparison result of the first placement information and the second placement information calculated by the calculation unit 16. In this case, the object estimation unit 14 calculates the dot product (cosine similarity) of the first placement information and the second placement information calculated by the calculation unit 16 as the similarity. Then, by calculating a weighted sum of the similarity between the separately calculated first feature and the second feature, the object estimation unit 14 calculates the degree of agreement by calculating a similarity based on a comparison of both the feature and the placement information. Based on this, landmarks with a degree of similarity above a predetermined threshold are narrowed down as landmarks that are highly likely to be objects detected by the object detection unit 11.

[0030] Furthermore, if the object detection unit 11 detects multiple objects, the object estimation unit 14 narrows down the multiple landmarks to the landmark corresponding to each object detected by the object detection unit 11.

[0031] The object estimation unit 14 then evaluates multiple candidate correspondences between multiple objects detected by the object detection unit 11 and multiple landmarks registered in the map database 40 using a scoring function that assigns a score to each candidate. The object estimation unit 14 then formally adopts the correspondence candidate with the highest score among the multiple candidate correspondences between multiple objects detected by the object detection unit 11 and multiple landmarks registered in the map database 40. The object estimation unit 14 can then maintain a dataset 100 that associates each of the multiple objects detected by the object detection unit 11 with a predetermined object acquired by the acquisition unit 12, which the object estimation unit 14 has estimated to be a candidate correspondence (identical). This makes it possible to have a highly reliable dataset readily available as needed, even in situations such as when the object detection unit 11 moves around the environment in real time. The mapping and retention of datasets are more specifically represented and retained using a consistency graph, where each correspondence (correspondence candidate) within the mapped dataset is represented as a node, and these nodes are connected by edges. In such a case, an edge is drawn between two correspondence candidates corresponding to two nodes when they satisfy a predefined spatial consistency condition. Here, the condition that the distance between objects included in the two correspondences is approximately equal in observation and on the map can be used. A set of mutually consistent correspondences can be efficiently found as a maximal clique of the consistency graph. A maximal clique is defined as a complete subgraph of the graph such that adding any adjacent node does not result in a complete graph. The set of multiple correspondence candidates found as a maximal clique is scored by the sum of the similarities between the correspondence candidates included in that set. This allows us to determine the most likely correspondence set with a higher overall similarity, especially when the number of correspondences found is equal.This allows us to estimate the likelihood of these highly reliable candidate data sets being used as candidates for the inlier, a highly reliable dataset used for position estimation and viewpoint transformation (coordinate transformation) of the position estimation device 10, and then utilize them.

[0032] The position estimation unit 15 estimates the position of the position estimation device 10 (in other words, the autonomous mobile robot equipped with the position estimation device 10, or the camera 30 attached to the position estimation device 10) by solving the PnP (Perspective-n-Point) problem based on the estimation results from the object estimation unit 14. In other words, the position estimation unit 15 estimates the position of the position estimation device 10 from the difference in viewpoint (amount and direction of viewpoint movement) between the object detected by the object detection unit 11 and the corresponding landmark registered in the map database 40.

[0033] Thus, the position estimation system 1 according to this disclosure can accurately estimate whether an object detected from the image captured by the camera 30 corresponds to a landmark registered in the map database 40 by acquiring and referring to not only information about landmarks obtained from the image captured by the camera 30, but also text-formatted information about landmarks obtained via the operation terminal 20, etc. Therefore, the position estimation system 1 according to this disclosure can narrow down the number of landmarks that are likely to be objects obtained from the image captured by the camera 30 from among multiple landmarks. As a result, even when multiple objects are detected from the image captured by the camera 30, the position estimation system 1 according to this disclosure can reduce the computational load required to determine the correspondence between the multiple objects and the multiple landmarks registered in the map database 40.

[0034] Next, we will explain the operation of the position estimation device 10 using Figure 2. Figure 2 is a flowchart showing the operation of the position estimation device 10.

[0035] First, the position estimation device 10 detects objects around the autonomous mobile robot positioned at a reference location from the images captured by the camera 30, and registers information about the detected objects (or features representing them) as information about landmarks (or features representing them) in the map database 40 (step S101).

[0036] Subsequently, the position estimation device 10 acquires text-format information about landmarks registered in the map database 40 via the operation terminal 20 (step S102). The acquired text-format information about landmarks (or the features representing them) is registered in the map database 40, linked to the corresponding landmarks.

[0037] Figure 3 shows an example of the contents registered in the map database 40. In the example in Figure 3, information on eight landmarks M1 to M8 is registered in the map database 40. Specifically, the map database 40 contains the location information of each landmark M1 to M8, and the associated text-formatted information (language label).

[0038] Subsequently, the position estimation device 10 moves in conjunction with the movement of the autonomous mobile robot, and then detects objects in the vicinity of the position estimation device 10 (in other words, the autonomous mobile robot on which the position estimation device 10 is mounted) (step S103).

[0039] Here, the position estimation device 10 converts information about objects detected by analyzing images captured by the camera 30 (including objects registered as landmarks in the map database 40) into first feature quantities, which are vector-defined feature quantities, using, for example, a multimodal recognition model (step S104). The position estimation device 10 also converts text-format information about landmarks acquired via the operation terminal 20, etc., into second feature quantities, which are vector-defined feature quantities, using, for example, a multimodal recognition model (step S104). This makes it possible to compare the information about objects detected by analyzing images with the information about objects acquired in text format.

[0040] Subsequently, the position estimation device 10 compares the object detected from the image captured by the camera 30 with the landmarks registered in the map database 40 to estimate whether or not the object detected from the image captured by the camera 30 corresponds to a landmark registered in the map database 40.

[0041] Specifically, the position estimation device 10 compares a first feature representing an object detected from the image captured by the camera 30 with a second feature representing a landmark registered in the map database 40 to estimate whether the object detected by the object estimation unit 14 corresponds to a landmark registered in the map database 40 (step S105).

[0042] Here, the second feature includes not only a feature representing information about the landmark obtained from the captured image, but also a feature representing text-formatted information about the landmark acquired via the operating terminal 20, etc. In other words, the second feature contains more detailed information about the landmark. Therefore, the position estimation device 10 can accurately estimate whether or not an object obtained from the captured image corresponds to a landmark registered in the map database 40. Furthermore, the landmark information may be represented by a feature representing at least one of text and an image.

[0043] Figure 4 shows the relationship between landmarks registered in the map database 40 and objects detected from images captured by the camera 30. In the example in Figure 4, information on eight landmarks M1 to M8 is registered in the map database 40, and three objects T1 to T3 are detected from images captured by the camera 30. Figure 4 also shows an example where no text-format information about landmarks is registered in the map database 40 as a comparative example.

[0044] First, in the comparison example shown in the upper part of Figure 4, no text-based information about the landmarks M1 to M8 registered in the map database 40 is provided. In this case, the position estimation device of the comparative example narrows down the landmarks M1, M2, and M7, which are most likely to be object T1, to two landmarks M4 and M5, which are most likely to be object T2, and to three landmarks M3, M6, and M8, which are most likely to be object T3, from the landmarks M1 to M8 registered in the map database 40. Therefore, there are 18 possible correspondences between the objects T1 to T3 detected from the images captured by the camera 30 and the landmarks M1 to M8 registered in the map database 40. Consequently, the position estimation device of the comparative example faces an increased computational burden in determining the correspondence between the objects T1 to T3 detected from the images captured by the camera 30 and the landmarks M1 to M8 registered in the map database 40.

[0045] In contrast, in the example shown in the lower part of Figure 4, landmarks M1 to M8 registered in the map database 40 are assigned text-based information about the landmarks. Specifically, landmark M1 is assigned the text-based information "Kid's table", landmark M2 is assigned the text-based information "Dusty table", landmark M3 is assigned the text-based information "Kitchen table", landmark M4 is assigned the text-based information "Digital clock", landmark M5 is assigned the text-based information "Big tall old clock", and landmark M6 is assigned the text-based information "Black short shelf".

[0046] In this case, the position estimation device 10 narrows down the landmarks M1 to M8 registered in the map database 40 to one landmark M1 that is most likely to be object T1, one landmark M4 that is most likely to be object T2, and one landmark M6 that is most likely to be object T3. Therefore, there is one possible correspondence between the objects T1 to T3 detected from the image captured by the camera 30 and the landmarks M1 to M8 registered in the map database 40. As a result, the position estimation device 10 can reduce the computational burden required to determine the correspondence between the objects T1 to T3 detected from the image captured by the camera 30 and the landmarks M1 to M8 registered in the map database 40.

[0047] Subsequently, the position estimation device 10 estimates the position of the position estimation device 10 (in other words, the autonomous mobile robot on which the position estimation device 10 is mounted, or the camera 30 attached to the position estimation device 10) based on the estimation results of the correspondence between objects T1 to T3 and landmarks M1 to M8 (step S106). The position estimation device 10 can also estimate the position of the camera using the correspondences included in the dataset 100 described above. More specifically, the data of highly reliable correspondence candidates included in the dataset 100 can be used as candidates for an inlier, which is a highly reliable dataset used for position estimation and viewpoint transformation (coordinate transformation) of the position estimation device 10, after estimating their likelihood. Alternatively, the final camera position and orientation may be calculated by minimizing the error between sets of correspondences (correspondence candidates) included in the dataset 100 using the weighted least squares method. The following two pieces of information (correspondence similarity and observational completeness) are considered as weighting coefficients. Correspondence similarity: Groups of correspondence candidates with high similarity among them are given more weight in the error calculation. Observational completeness: The smaller the size of the observed candidate correspondence relative to the size of the known map landmark, the more weight is given to it. The product of these values ​​is used as the final weighting coefficient. This reduces the impact of less likely (potentially incorrect) correspondences and the impact of partially observed objects that affect the accuracy of attitude calculations, resulting in robust attitude calculations.

[0048] Thus, the position estimation system 1 according to this disclosure can accurately estimate whether an object detected from the image captured by the camera 30 corresponds to a landmark registered in the map database 40 by acquiring and referring to not only information about landmarks obtained from the image captured by the camera 30, but also text-formatted information about landmarks obtained via the operation terminal 20, etc. Therefore, the position estimation system 1 according to this disclosure can narrow down the number of landmarks that are likely to be objects obtained from the image captured by the camera 30 from among multiple landmarks. As a result, even when multiple objects are detected from the image captured by the camera 30, the position estimation system 1 according to this disclosure can reduce the computational load required to determine the correspondence between the multiple objects and the multiple landmarks registered in the map database 40.

[0049] Furthermore, the position estimation system 1 may acquire text-based information about landmarks A1 to A4, such as facility names, building names, and room names, from floor maps MP1 installed in stations, shopping malls, public facilities, etc., as shown in Figure 5. This enables high-performance position estimation of the camera 30 on the floor represented by the floor map MP1.

[0050] Furthermore, this disclosure can be realized by having a CPU (Central Processing Unit) execute a computer program to perform part or all of the processing of the position estimation device 10 or the position estimation system 1 equipped therewith.

[0051] The program described above includes, when loaded into a computer, a set of instructions (or software code) for causing the computer to perform one or more of the functions described in the embodiments. The program may be stored in a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include RAM (Random-Access Memory), ROM (Read-Only Memory), flash memory, SSD (Solid-State Drive), or other memory technologies, CD-ROM, DVD (Digital Versatile Disc), Blu-ray® disc, or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically, or otherwise propagating signals. [Explanation of Symbols]

[0052] 1. Position estimation system 10 Position estimation device 11 Object detection unit 12 Acquisition Department 13 Conversion section 14 Object estimation part 15 Position estimation part 16 Calculation Section 20 Operating terminals 30 Cameras 40 Map Databases 50 Networks

Claims

1. An acquisition unit acquires text-format information related to a predetermined object and surrounding objects located around that object from a database in which text-format information related to a predetermined object is stored in association with the location information of that object. An object detection unit that detects an object and surrounding objects which are objects in the vicinity of that object, A conversion unit converts the object detected by the object detection unit into a first feature quantity which is a feature quantity defined as data representing the object, and converts the text-format information related to the predetermined object acquired by the acquisition unit into a second feature quantity which is a feature quantity defined as data representing the object. A calculation unit calculates first arrangement information that defines the arrangement relationship between the object detected by the object detection unit and the surrounding object, and calculates second arrangement information that defines the arrangement relationship between the object acquired by the acquisition unit and the surrounding object. An object estimation unit compares a first feature quantity representing the object detected by the object detection unit with a second feature quantity representing the predetermined object, and compares the first arrangement information calculated by the calculation unit with the second arrangement information to estimate whether the object detected by the object detection unit is the same as the predetermined object. An object estimation device equipped with the following features.

2. The object estimation unit performs the estimation for each of the multiple objects detected by the object detection unit. A dataset is maintained which associates the object detected by the object detection unit, which is estimated to be the same as the object estimation unit, with the predetermined object. The object estimation device according to claim 1.

3. The object estimation device according to claim 1, wherein the object estimation unit performs the estimation for each of the plurality of objects detected by the object detection unit, A position estimation unit estimates the position of the object estimation device based on the positional relationship of each of a plurality of pairs of objects detected by the object detection unit and the predetermined object, which are estimated to be the same by the object estimation unit. A position estimation system equipped with the following features.

4. Computers From a database in which text-formatted information related to a given object is stored in association with the object's location information, text-formatted information related to the given object and surrounding objects located around it is retrieved. The system detects an object and surrounding objects that are located in the vicinity of that object. The detected object is converted into a first feature, which is a feature defined as data representing the object, and the acquired text-format information related to the predetermined object is converted into a second feature, which is a feature defined as data representing the object. First arrangement information defining the arrangement relationship between the detected object and the surrounding objects is calculated, and second arrangement information defining the arrangement relationship between the acquired object and the surrounding objects is calculated. By comparing the first feature quantity representing the detected object with the second feature quantity representing the predetermined object, and by comparing the calculated first placement information with the second placement information, it is estimated whether the detected object is identical to the predetermined object. Object estimation method.

5. A process to retrieve text-format information related to a given object and surrounding objects located around it from a database in which text-format information related to a given object is stored in association with the object's location information, A process for detecting an object and surrounding objects that are located around that object, The process involves converting the detected object into a first feature, which is a feature defined as data representing the object, and converting the acquired text-format information related to the predetermined object into a second feature, which is a feature defined as data representing the object. A process to calculate first arrangement information that defines the arrangement relationship between the detected object and the surrounding objects, and to calculate second arrangement information that defines the arrangement relationship between the acquired object and the surrounding objects, A process to estimate whether the detected object is identical to the predetermined object by comparing the first feature quantity representing the detected object with the second feature quantity representing the predetermined object, and by comparing the calculated first placement information with the second placement information. A control program that instructs a computer to execute a command.

Citation Information

Patent Citations

  • JP1974085166A