Object estimation apparatus, position estimation system, object estimation method, and control program
The object estimation device uses vector-based feature amounts from a multimodal recognition model to accurately identify predetermined objects, addressing the computational burden of landmark matching in multiple landmark scenarios.
Patent Information
- Application Number
- JP2023215660
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-03
AI Technical Summary
Existing devices struggle to accurately estimate whether extracted landmarks correspond to predetermined landmarks, leading to increased computational load when multiple landmarks are detected.
An object estimation device that converts detected objects and text-formatted information into vector-based feature amounts using a multimodal recognition model, enabling accurate estimation and reducing computational load by narrowing down likely objects from a plurality of landmarks.
Accurately estimates whether detected objects are predetermined objects, reducing computational load by narrowing down candidates, even when multiple objects are detected.
Smart Images

Figure 2025099197000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object estimation device, a position estimation system, an object estimation method, and a control program.
Background Art
[0002] Patent Document 1 discloses a self-position estimation device capable of accurately recognizing different landmarks and estimating its own position with high accuracy when a plurality of landmarks are extracted.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Since the device disclosed in Patent Document 1 cannot accurately estimate whether the extracted landmark corresponds to a predetermined landmark stored in the storage means, it is impossible to narrow down a small number of predetermined landmarks that are highly likely to be the extracted landmarks from the plurality of predetermined landmarks stored in the storage means. As a result, in the device disclosed in Patent Document 1, when a plurality of landmarks are extracted, the computational load required to determine the correspondence between the plurality of extracted landmarks and the plurality of predetermined landmarks stored in the storage means becomes large.
[0005] The present disclosure has been made in view of the above background, and an object thereof is to provide an object estimation device, a position estimation system, an object estimation method, and a control program capable of accurately estimating whether an object detected from the surrounding environment is a predetermined object.
Means for Solving the Problems
[0006] The object estimation device according to the present disclosure includes an acquisition unit that acquires text-formatted information related to a predetermined object, an object detection unit that detects surrounding objects, and the object detected by the object detection unit is converted into a first feature amount, which is a feature amount defined by a vector, using a machine learning model such as a multimodal recognition model. At the same time, the text-formatted information related to the predetermined object acquired by the acquisition unit is converted into a second feature amount, which is a feature amount defined by a vector, using a machine learning model such as a multimodal recognition model. A conversion unit, and an object estimation unit that estimates whether the object detected by the object detection unit is the predetermined object by comparing the first feature amount representing the object detected by the object detection unit with a reference feature amount including the second feature amount representing the predetermined object. This object estimation device can accurately estimate whether an object detected from the surrounding environment is a predetermined object by acquiring and referring to text-formatted information related to the predetermined object. Therefore, this object estimation device can narrow down a small number of predetermined objects with a high possibility of being the object detected from the surrounding environment from a plurality of predetermined objects. As a result, even when a plurality of objects are detected from the surrounding environment, this object estimation device can reduce the computational load required to determine the correspondence between the plurality of objects and, for example, a plurality of predetermined objects registered in a map database.
[0007] The object estimation method according to the present disclosure is such that a computer acquires information in text form related to a predetermined object, detects surrounding objects, converts the detected objects into a first feature quantity which is a feature quantity defined by a vector, and converts the acquired information in text form related to the predetermined object into a second feature quantity which is a feature quantity defined by a vector, and estimates whether the detected object is the predetermined object by comparing the first feature quantity representing the detected object with a reference feature quantity including the second feature quantity representing the predetermined object. This object estimation method can accurately estimate whether an object detected from the surrounding environment is a predetermined object by acquiring and referring to information in text form related to the predetermined object. Therefore, this object estimation method can narrow down a small number of predetermined objects with a high possibility of being the object detected from the surrounding environment from a plurality of predetermined objects. As a result, this object estimation method can reduce the computational load required to determine the correspondence between the plurality of objects and, for example, a plurality of predetermined objects registered in a map database even when a plurality of objects are detected from the surrounding environment.
[0008] The control program according to the present disclosure causes a computer to execute a process of acquiring text-form information related to a predetermined object, a process of detecting surrounding objects, a process of converting the detected object into a first feature amount which is a feature amount defined by a vector, and converting the acquired text-form information related to the predetermined object into a second feature amount which is a feature amount defined by a vector, and a process of estimating whether or not the detected object is the predetermined object by comparing the first feature amount representing the detected object and a reference feature amount including the second feature amount representing the predetermined object. By acquiring and referring to text-form information related to a predetermined object, this control program can accurately estimate whether or not an object detected from the surrounding environment is the predetermined object. Therefore, this control program can narrow down a predetermined object with a high possibility of being an object detected from the surrounding environment from a plurality of predetermined objects. As a result, even when a plurality of objects are detected from the surrounding environment, this control program can reduce the computational load required to determine the correspondence relationship between the plurality of objects and, for example, a plurality of predetermined objects registered in a map database.
Effect of the Invention
[0009] According to the present disclosure, it is possible to provide an object estimation device, a position estimation system, an object estimation method, and a control program capable of accurately estimating whether or not an object detected from the surrounding environment is a predetermined object.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Embodiments for Carrying Out the Invention
[0011] Hereinafter, the present invention will be described through embodiments of the invention, but the invention according to the claims is not limited to the following embodiments. Also, not all of the configurations described in the embodiments are necessarily essential as means for solving the problems. For the sake of clarity of explanation, the following description and drawings have been appropriately omitted and simplified. In each drawing, the same elements are denoted by the same reference numerals, and redundant explanations are omitted as necessary.
[0012] <Embodiment 1> FIG. 1 is a diagram showing a configuration example of a position estimation system 1 according to Embodiment 1. The position estimation system 1 is applied to, for example, an autonomous mobile robot, and estimates the self-position of the autonomous mobile robot by collating an object detected from a captured image of a camera or the like with a landmark registered in a map database. Here, the position estimation system 1 refers not only to information on a landmark detected from a captured image of a camera or the like, but also to text-form information on a landmark acquired via an operation terminal or the like, so as to accurately estimate whether an object detected from a captured image of a camera or the like corresponds to a landmark registered in the map database. Therefore, the position estimation system 1 can narrow down a small number of landmarks that are likely to be the object detected from a captured image of a camera or the like from a plurality of landmarks. As a result, even when a plurality of objects are detected from a captured image of a camera or the like, the position estimation system 1 can reduce the computational load required to determine the correspondence between the plurality of objects and the plurality of landmarks registered in the map database. This will be specifically described below.
[0013] As shown in FIG. 1, the position estimation system 1 includes a position estimation device 10, an operation terminal 20, a camera 30, a map database 40, and a network 50. The position estimation device 10 can also be referred to as a position estimation system by itself. The position estimation device 10, the operation terminal 20, the camera 30, and the map database 40 are configured to be able to communicate with each other via a wired or wireless network 50. In this embodiment, a case where the position estimation system 1 is applied to an autonomous mobile robot will be described as an example.
[0014] The autonomous mobile robot moves from the current location to the destination while estimating its own position by collating the surrounding objects detected by the position estimation device 10 with the landmarks that are the objects registered in the map database 40.
[0015] The position estimation device 10 includes an object detection unit 11, an acquisition unit 12, a conversion unit 13, an object estimation unit 14, and a position estimation unit 15. Note that the object detection unit 11, the acquisition unit 12, the conversion unit 13, and the object estimation unit 14 constitute an object estimation device.
[0016] The object detection unit 11 detects the objects around the autonomous mobile robot by analyzing the captured image of the camera 30 attached to the autonomous mobile robot. Instead of the camera 30, a distance measurement sensor or a depth sensor may be used, but the camera 30 is suitable for use because it is lighter and less expensive compared to the distance measurement sensor and the depth sensor. For example, the object detection unit 11 detects a rectangular image portion surrounding the object included in the captured image of the camera 30 as the object. In this case, the center of the rectangular image surrounding the object becomes the representative point of the object. Note that the object detection unit 11 can improve the detection accuracy of the objects included in the captured image by using a learned model generated by machine learning using a plurality of captured images.
[0017] Here, the object detection unit 11 detects objects around the autonomous mobile robot arranged at the reference position, and registers information about the detected objects (or feature amounts representing them) in the map database 40 as information about landmarks (or feature amounts representing them). Note that the information about landmarks includes information such as the shape, color, and pattern of the objects that can be specified from the captured images. The information about landmarks also includes information such as the type of the objects that can be specified from the captured images. Furthermore, the information about landmarks includes the position information of the objects used to represent the reference position of the position estimation device 10 (autonomous mobile robot). Note that in the map database 40, additional landmarks may be registered as appropriate if the positional relationship with the already registered landmarks is clear.
[0018] The acquisition unit 12 acquires information in text format about the landmarks registered in the map database 40 separately from the information about the landmarks detected by the object detection unit 11. For example, when the landmark is a store, the acquisition unit 12 acquires detailed information about the landmark, such as the store name, in text format, and when the landmark is furniture or tableware, the specific type thereof. The information in text format about the landmarks (or feature amounts representing them) acquired by the acquisition unit 12 is registered in the map database 40 associated with the corresponding landmarks.
[0019] The information in text format about the landmarks acquired by the acquisition unit 12 is transmitted, for example, from the operation terminal 20.
[0020] The operation terminal 20 is a communicable terminal owned by the user or temporarily assigned to the user, and is, for example, a PC (Personal Computer) terminal, a mobile terminal such as a smartphone or a tablet terminal, or a dedicated communication terminal prepared for this system.
[0021] For example, the user can operate the monitor of the operation terminal 20 by touching it with a touch pen or a finger, or operate a mouse, keyboard, etc. of the operation terminal 20, thereby inputting text-formatted information regarding landmarks registered in the map database 40 into the operation terminal 20. The operation terminal 20 receives the text-formatted information regarding the landmark and transmits it to the position estimation device 10 via the network 50.
[0022] Note that the acquisition unit 12 may acquire text-formatted information regarding the landmark from an external device other than the operation terminal 20.
[0023] The conversion unit 13 converts information regarding an object (including an object registered as a landmark in the map database 40) detected by the object detection unit 11 into a feature amount defined by a vector using, for example, a multimodal recognition model. Also, the conversion unit 13 converts the text-formatted information regarding the landmark acquired by the acquisition unit 12 into a feature amount defined by a vector using, for example, a multimodal recognition model. Thereby, comparison between the information regarding the object detected by analyzing the captured image and the information regarding the object acquired in text form becomes possible.
[0024] The object estimation unit 14 collates the object detected by the object detection unit 11 with the landmarks registered in the map database 40, and estimates whether the object detected by the object detection unit 11 corresponds to a landmark registered in the map database 40.
[0025] Specifically, the object estimation unit 14 estimates whether the object detected by the object detection unit 14 corresponds to a landmark registered in the map database 40 by comparing the feature amount (first feature amount) representing the object detected by the object detection unit 11 with the reference feature amount representing the landmark registered in the map database 40.
[0026] Here, the reference feature amount includes, in addition to the feature amount representing information on the landmark obtained from the captured image, a feature amount (second feature amount) representing text-formatted information on the landmark acquired from the acquisition unit 12. That is, the reference feature amount includes more detailed information on the landmark. Therefore, the object estimation unit 14 can accurately estimate whether the object detected by the object detection unit 11 corresponds to a landmark registered in the map database 40.
[0027] Also, when a plurality of landmarks are registered in the map database 40, the object estimation unit 14 narrows down the landmarks that are likely to be the object detected by the object detection unit 11 from among the plurality of landmarks.
[0028] Specifically, the object estimation unit 14 compares the first feature amount representing the object detected by the object detection unit 11 with a plurality of reference feature amounts each representing one of the plurality of landmarks registered in the map database 40, and narrows down the landmarks with a reference feature amount having a degree of coincidence with the first feature amount equal to or higher than a predetermined threshold as the landmarks that are likely to be the object detected by the object detection unit 11. Alternatively, the object estimation unit 14 may compare the first feature amount representing the object detected by the object detection unit 11 with a plurality of reference feature amounts each representing one of the plurality of landmarks registered in the map database 40, and estimate the landmark with the reference feature amount having the highest degree of coincidence with the first feature amount as the landmark that is likely to be the object detected by the object detection unit 11.
[0029] Here, as already described, the reference feature amount includes, in addition to the feature amount representing information on the landmark obtained from the captured image, a second feature amount representing text-formatted information on the landmark acquired from the acquisition unit 12. That is, the reference feature amount includes more detailed information on the landmark. Therefore, the object estimation unit 14 can narrow down fewer landmarks that are likely to be the object detected by the object detection unit 11.
[0030] Furthermore, when a plurality of objects are detected by the object detection unit 11, the object estimation unit 14 narrows down the landmarks corresponding to each object detected by the object detection unit 11 from among the plurality of landmarks.
[0031] Then, the object estimation unit 14 evaluates a plurality of candidates for the correspondence between the plurality of objects detected by the object detection unit 11 and the plurality of landmarks registered in the map database 40 using a scoring function in which plausibility is scored. Then, the object estimation unit 14 officially adopts the candidate for the correspondence showing the highest score among the plurality of candidates for the correspondence between the plurality of objects detected by the object detection unit 11 and the plurality of landmarks registered in the map database 40.
[0032] Based on the estimation result by the object estimation unit 14, the position estimation unit 15 solves the Perspective-n-Point (PnP) problem to estimate the position of the position estimation device 10 (in other words, an autonomous mobile robot equipped with the position estimation device 10, or a camera 30 attached to the position estimation device 10). That is, the position estimation unit 15 estimates the position of the position estimation device 10 from the viewpoint deviation (the amount and direction of movement of the viewpoint) between the object detected by the object detection unit 11 and the landmark registered in the corresponding map database 40.
[0033] As described above, the position estimation system 1 according to the present disclosure can accurately estimate whether an object detected from a captured image of the camera 30 corresponds to a landmark registered in the map database 40 by obtaining and referring to not only information on landmarks obtained from the captured image of the camera 30 but also text-form information on landmarks obtained via the operation terminal 20 or the like. Therefore, the position estimation system 1 according to the present disclosure can narrow down a small number of landmarks that are highly likely to be objects obtained from the captured image of the camera 30 from a plurality of landmarks. As a result, even when a plurality of objects are detected from the captured image of the camera 30, the position estimation system 1 according to the present disclosure can reduce the computational load required to determine the correspondence between the plurality of objects and the plurality of landmarks registered in the map database 40.
[0034] Subsequently, the operation of the position estimation device 10 will be described with reference to FIG. 2. FIG. 2 is a flowchart showing the operation of the position estimation device 10.
[0035] First, the position estimation device 10 detects an object around the autonomous mobile robot arranged at the reference position from the captured image of the camera 30, and registers information (or a feature amount representing the same) on the detected object as information (or a feature amount representing the same) on a landmark in the map database 40 (step S101).
[0036] Thereafter, the position estimation device 10 acquires text-form information on the landmarks registered in the map database 40 via the operation terminal 20 (step S102). The acquired text-form information (or a feature amount representing the same) on the landmark is registered in the map database 40 in association with the corresponding landmark.
[0037] FIG. 3 is a diagram showing an example of the registration contents of the map database 40. In the example of FIG. 3, information regarding eight landmarks M1 to M8 is registered in the map database 40. Specifically, position information of each of the landmarks M1 to M8 and information in text format (language label) associated therewith are registered in the map database 40.
[0038] Thereafter, after moving along with the movement of the autonomous mobile robot, the position estimation device 10 detects an object around the position estimation device 10 (in other words, the autonomous mobile robot on which the position estimation device 10 is mounted) (step S103).
[0039] Here, the position estimation device 10 converts information regarding an object detected by analyzing the captured image of the camera 30 (including an object registered as a landmark in the map database 40) into a first feature amount which is a feature amount defined by a vector, for example, using a multimodal recognition model (step S104). Further, the position estimation device 10 converts information in text format regarding a landmark acquired via the operation terminal 20 or the like into a second feature amount which is a feature amount defined by a vector, for example, using a multimodal recognition model (step S104). Thereby, comparison between the information regarding the object detected by analyzing the captured image and the information regarding the object acquired in text format becomes possible.
[0040] Thereafter, the position estimation device 10 collates the object detected from the captured image of the camera 30 with the landmarks registered in the map database 40, and estimates whether or not the object detected from the captured image of the camera 30 corresponds to the landmarks registered in the map database 40.
[0041] Specifically, the position estimation device 10 estimates whether or not the object detected by the object estimation unit 14 corresponds to the landmarks registered in the map database 40 by comparing a first feature amount representing the object detected from the captured image of the camera 30 with a reference feature amount representing the landmarks registered in the map database 40 (step S105).
[0042] Here, in addition to the feature amount representing information on the landmark obtained from the captured image, the reference feature amount includes a feature amount (second feature amount) representing text-form information on the landmark acquired via the operation terminal 20 or the like. That is, the reference feature amount includes more detailed information on the landmark. Therefore, the position estimation device 10 can accurately estimate whether the object obtained from the captured image corresponds to the landmark registered in the map database 40.
[0043] FIG. 4 is a diagram showing the relationship between the landmarks registered in the map database 40 and the objects detected from the captured image of the camera 30. In the example of FIG. 4, information on eight landmarks M1 to M8 is registered in the map database 40, and three objects T1 to T3 are detected from the captured image of the camera 30. Note that FIG. 4 also shows an example in the case where text-form information on the landmark is not registered in the map database 40 as a comparative example.
[0044] First, in the comparative example of the upper diagram in FIG. 4, text-form information on the landmarks is not given to the landmarks M1 to M8 registered in the map database 40. In this case, the position estimation device of the comparative example narrows down three landmarks M1, M2, and M7 that are highly likely to be the object T1, two landmarks M4 and M5 that are highly likely to be the object T2, and three landmarks M3, M6, and M8 that are highly likely to be the object T3 from the landmarks M1 to M8 registered in the map database 40. Therefore, there are 18 candidates for the correspondence between the objects T1 to T3 detected from the captured image of the camera 30 and the landmarks M1 to M8 registered in the map database 40. Therefore, in the position estimation device of the comparative example, the computational load required to determine the correspondence between the objects T1 to T3 detected from the captured image of the camera 30 and the landmarks M1 to M8 registered in the map database 40 increases.
[0045] On the other hand, in the example of the lower diagram in FIG. 4, for the landmarks M1 to M8 registered in the map database 40, text-form information regarding the landmarks is provided. Specifically, text-form information of "Kid’s table" is provided for landmark M1, text-form information of "Dusty table" is provided for landmark M2, text-form information of "Kitchen table" is provided for landmark M3, text-form information of "Digital clock" is provided for landmark M4, text-form information of "Big tall old clock" is provided for landmark M5, and text-form information of "Black short shelf" is provided for landmark M6.
[0046] In this case, the position estimation device 10 narrows down one landmark M1 with a high possibility of being the object T1, one landmark M4 with a high possibility of being the object T2, and one landmark M6 with a high possibility of being the object T3 from the landmarks M1 to M8 registered in the map database 40. Therefore, there is one candidate for the correspondence relationship between the objects T1 to T3 detected from the captured image of the camera 30 and the landmarks M1 to M8 registered in the map database 40. Thus, the position estimation device 10 can reduce the computational burden required to determine the correspondence relationship between the objects T1 to T3 detected from the captured image of the camera 30 and the landmarks M1 to M8 registered in the map database 40.
[0047] Thereafter, based on the estimation result of the correspondence relationship between the objects T1 to T3 and the landmarks M1 to M8, the position of the position estimation device 10 (in other words, the autonomous mobile robot equipped with the position estimation device 10, or the camera 30 attached to the position estimation device 10) is estimated (step S106).
[0048] As described above, the position estimation system 1 according to the present disclosure can accurately estimate whether an object detected from a captured image of the camera 30 corresponds to a landmark registered in the map database 40 by obtaining and referring to not only information on landmarks obtained from the captured image of the camera 30 but also text-form information on landmarks obtained via the operation terminal 20 or the like. Therefore, the position estimation system 1 according to the present disclosure can narrow down, from a plurality of landmarks, the landmarks that are highly likely to be objects obtained from the captured image of the camera 30. As a result, even when a plurality of objects are detected from the captured image of the camera 30, the position estimation system 1 according to the present disclosure can reduce the computational load required to determine the correspondence between the plurality of objects and the plurality of landmarks registered in the map database 40.
[0049] Note that the position estimation system 1 may obtain text-form information on landmarks A1 to A4, such as facility names, building names, and room names, from a floor map MP1 installed in a station, a shopping mall, a public facility, etc., as shown in FIG. 5. Thereby, high-performance position estimation of the camera 30 on the floor represented by the floor map MP1 becomes possible.
[0050] Further, a part or all of the processing of the position estimation device 10 or the position estimation system 1 including the same according to the present disclosure can be realized by causing a CPU (Central Processing Unit) to execute a computer program.
[0051] When the above program is loaded into a computer, it includes a set of instructions (or software code) for causing the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, the computer-readable medium or tangible storage medium includes RAM (Random-Access Memory), ROM (Read-Only Memory), flash memory, SSD (Solid-State Drive) or other memory technologies, CD-ROM, DVD (Digital Versatile Disc), Blu-ray (registered trademark) disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted on a transitory computer-readable medium or a communication medium. By way of example and not limitation, the transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
Explanation of Signs
[0052] 1 Position Estimation System 10 Position Estimation Device 11 Object Detection Unit 12 Acquisition Unit 13 Conversion Unit 14 Object Estimation Unit 15 Position Estimation Unit 20 Operation Terminal 30 Camera 40 Map Database 50 Network
Claims
1. An acquisition unit that acquires text-formatted information related to a predetermined object; An object detection unit that detects surrounding objects; A conversion unit that converts the object detected by the object detection unit into a first feature amount that is a feature amount defined by a vector, and converts the text-formatted information related to the predetermined object acquired by the acquisition unit into a second feature amount that is a feature amount defined by a vector; An object estimation unit that estimates whether or not the object detected by the object detection unit is the predetermined object by comparing the first feature amount representing the object detected by the object detection unit with a reference feature amount including the second feature amount representing the predetermined object; An object estimation device comprising the above.
2. The object estimation unit estimates a predetermined object that is highly likely to be the object detected by the object detection unit from among a plurality of the predetermined objects. The object estimation device according to Claim 1.
3. The object estimation device according to Claim 1; A position estimation unit that estimates the position of the object estimation device based on the result estimated by the object estimation device; A position estimation system comprising the above.
4. A computer: Acquires text-formatted information related to a predetermined object; Detects surrounding objects; Converts the detected object into a first feature amount that is a feature amount defined by a vector, and converts the text-formatted information related to the predetermined object acquired into a second feature amount that is a feature amount defined by a vector; Estimates whether or not the detected object is the predetermined object by comparing the first feature amount representing the detected object with a reference feature amount including the second feature amount representing the predetermined object. An object estimation method.
5. A process of acquiring text-formatted information related to a predetermined object; A process of detecting surrounding objects; A process of converting the detected object into a first feature amount that is a feature amount defined by a vector, and converting the text-formatted information related to the predetermined object acquired into a second feature amount that is a feature amount defined by a vector; A process of estimating whether or not the detected object is the predetermined object by comparing the first feature amount representing the detected object with a reference feature amount including the second feature amount representing the predetermined object; A control program for causing a computer to execute the above.
Citation Information
Patent Citations
Self-location estimation device
JP2009020014A
Systems and methods for open vocabulary object detection
US20230154213A1
JP1974085166A