Object estimation device, position estimation system, object estimation method, and control program
By acquiring text format information related to a specified object and configuration information around the object, and by utilizing feature transformation and configuration information calculation, the problem of insufficient object inference accuracy in existing technologies is solved, and efficient object inference is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies cannot accurately determine whether an object detected from the surrounding environment is a specified object, resulting in excessive computational load when multiple objects are detected.
By acquiring text format information related to the specified object and configuration information of objects surrounding the object, and using feature transformation and configuration information calculation, it is possible to infer with good accuracy whether the object is the specified object.
This reduces the computational load in the case of multiple object detections, and improves the accuracy and efficiency of object inference.
Smart Images

Figure CN121783111A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an object inference device, a position inference system, an object inference method, and a control program. Background Technology
[0002] Patent document 1 discloses a self-position inference device that, when multiple landmarks are extracted, can accurately identify different landmarks from each other and infer its own position with high accuracy.
[0003] Patent Document 1: Japanese Patent No. 4985166 Summary of the Invention The apparatus disclosed in Patent Document 1 cannot accurately infer whether the extracted landmarks correspond to specified landmarks stored in the storage unit. Therefore, it is impossible to narrow down the number of specified landmarks stored in the storage unit to the specified landmarks with a high probability of being extracted. As a result, in the apparatus disclosed in Patent Document 1, when multiple landmarks are extracted, the computational load required to determine the correspondence between the extracted multiple landmarks and the multiple specified landmarks stored in the storage unit increases.
[0004] The present invention was made in view of the above background, and its purpose is to provide an object inference device, a position inference system, an object inference method and a control program that can accurately infer whether an object detected from the surrounding environment is a specified object.
[0005] The object inference apparatus of the present invention comprises: an acquisition unit that acquires text-formatted information relating to a specified object and surrounding objects located around the object from a database storing text-formatted information related to a specified object and the object's position information in a corresponding association; an object detection unit that detects the object and surrounding objects that are surrounding objects of the object; and a conversion unit that converts the object detected by the object detection unit into a first feature quantity defined as data representing the object, and converts the text-formatted information relating to the specified object acquired by the acquisition unit into a second feature quantity defined as data representing the object. The device comprises: a calculation unit that calculates first configuration information and second configuration information, wherein the first configuration information defines the configuration relationship between the object detected by the object detection unit and the surrounding objects, and the second configuration information defines the configuration relationship between the object acquired by the acquisition unit and the surrounding objects; and an object inference unit that infers whether the object detected by the object detection unit is the same as the specified object by comparing the first feature quantity representing the object detected by the object detection unit with the second feature quantity representing the specified object, and by comparing the first configuration information calculated by the calculation unit with the second configuration information. This object inference device, by acquiring and referencing text format information related to the specified object and configuration information of objects surrounding the specified object, can accurately infer whether an object detected from the surrounding environment is a specified object. Therefore, this object inference device can narrow down a plurality of specified objects to a specified object with a high probability of being detected from the surrounding environment. As a result, even when multiple objects are detected from the surrounding environment, the object inference device is able to reduce the computational load required to determine the correspondence between the multiple objects and, for example, multiple specified objects registered in a map database.
[0006] In the object inference method of this invention, a computer performs the following processing: From a database storing text-formatted information related to a specified object and its surrounding objects, corresponding to the object's location information; Detecting the object and its surrounding objects; Converting the detected object into a first feature quantity (defined as data representing the object), and converting the acquired text-formatted information related to the specified object into a second feature quantity (defined as data representing the object); Calculating first configuration information and second configuration information, wherein the first configuration information defines the configuration relationship between the detected object and the surrounding objects, and the second configuration information defines the configuration relationship between the acquired object and the surrounding objects; and By comparing the first feature representing the detected object with the second feature representing the specified object, and by comparing the calculated first configuration information with the second configuration information, it is inferred whether the detected object is the same as the specified object.
[0007] This object inference method, by acquiring and referencing textual information related to a specified object and the configuration information of objects surrounding the specified object, can accurately infer whether an object detected from the surrounding environment is a specified object. Therefore, this object inference method can narrow down a plurality of specified objects to those with a high probability of being detected from the surrounding environment. Consequently, even when multiple objects are detected from the surrounding environment, this object inference method can reduce the computational load required to determine the correspondence between these multiple objects and, for example, multiple specified objects registered in a map database.
[0008] The control program involved in this invention causes a computer to perform the following processing: obtaining text-formatted information related to a specified object and surrounding objects located around the object from a database that stores information in a corresponding association between text-formatted information related to a specified object and the object's location information; detecting the object and surrounding objects that are surrounding the object; converting the detected object into a first feature quantity defined as data representing the object, and converting the obtained text-formatted information related to the specified object into a second feature quantity defined as data representing the object; calculating first configuration information and second configuration information, wherein the first configuration information defines the configuration relationship between the detected object and the surrounding objects, and the second configuration information defines the configuration relationship between the obtained object and the surrounding objects; and inferring whether the detected object is the same as the specified object by comparing the first feature quantity representing the detected object with the second feature quantity representing the specified object, and by comparing the calculated first configuration information with the second configuration information.
[0009] This control program, by acquiring and referencing textual information related to a specified object and the configuration information of objects surrounding the specified object, can accurately infer whether an object detected from the surrounding environment is a specified object. Therefore, the control program can narrow down a plurality of specified objects to those with a high probability of being detected from the surrounding environment. Consequently, even when multiple objects are detected from the surrounding environment, the control program can reduce the computational load required to determine the correspondence between these multiple objects and, for example, multiple specified objects registered in a map database.
[0010] Invention Effects This invention provides an object inference device, a position inference system, an object inference method, and a control program that can accurately infer whether an object detected from the surrounding environment is a specified object. Attached Figure Description
[0011] Figure 1 This is a diagram illustrating a configuration example of the position inference system involved in Implementation 1.
[0012] Figure 2 This is a flowchart illustrating the operation of the position inference device involved in Embodiment 1.
[0013] Figure 3 This is a diagram representing an example of the registered content in a map database.
[0014] Figure 4 It is a graph that shows the relationship between landmarks registered in a map database and objects detected from camera images.
[0015] Figure 5 This is a diagram used to illustrate an applicable example of the position inference system involved in Embodiment 1. Detailed Implementation
[0016] The present invention will now be described through embodiments thereof, but the invention is not limited to these embodiments. Furthermore, not all configurations described in the embodiments are necessarily necessary to solve the problem. For clarity, the following descriptions and drawings have been appropriately omitted and simplified. In the drawings, the same symbols are used to denote the same elements, and repeated descriptions are omitted as necessary.
[0017] <Implementation Method 1> Figure 1 This diagram illustrates a configuration example of the location inference system 1 according to Embodiment 1. The location inference system 1 is applicable, for example, to an autonomous mobile robot, and infers the robot's own position by comparing objects detected from camera images, etc., with landmarks registered in a map database. Here, the location inference system 1 references not only information related to landmarks detected from camera images, etc., but also text-formatted information related to landmarks obtained via an operating terminal, etc., enabling it to accurately infer whether objects detected from camera images, etc., correspond to landmarks registered in the map database. Therefore, the location inference system 1 can narrow down from multiple landmarks to those with a high probability of being objects detected from camera images, etc. As a result, even when multiple objects are detected from camera images, etc., the location inference system 1 can reduce the computational load required to determine the correspondence between these multiple objects and multiple landmarks registered in the map database. A detailed explanation follows.
[0018] like Figure 1 As shown, the location inference system 1 includes a location inference device 10, an operating terminal 20, a camera 30, a map database 40, and a network 50. The location inference device 10 can also function as a standalone location inference system. The location inference device 10, the operating terminal 20, the camera 30, and the map database 40 are configured to communicate with each other via a wired or wireless network 50. In this embodiment, an example of the location inference system 1 being applicable to an autonomous mobile robot will be described.
[0019] The autonomous mobile robot moves from its current location to its destination by comparing its own location with the surrounding objects detected by the location inference device 10 and the landmarks of the objects registered in the map database 40.
[0020] The position inference device 10 includes an object detection unit 11, an acquisition unit 12, a conversion unit 13, an object inference unit 14, and a position inference unit 15. Furthermore, the object detection unit 11, the acquisition unit 12, the conversion unit 13, and the object inference unit 14 constitute the object inference device. The object inference unit 14 includes a calculation unit 16.
[0021] The object detection unit 11 detects objects around the autonomous mobile robot by analyzing images captured by the camera 30 mounted on the robot. While a range sensor or depth sensor can be used instead of the camera 30, the camera 30 is lighter and less expensive, making it more suitable. For example, the object detection unit 11 detects a rectangular portion of the image surrounding the camera 30 as the object. In this case, the center of the rectangular image surrounding the object becomes the representative point of that object. Furthermore, by using a pre-trained model generated through machine learning from multiple images, the object detection unit 11 can improve the detection accuracy of objects contained in the images.
[0022] Here, the object detection unit 11 detects at least one object surrounding the autonomous mobile robot positioned at the reference position, and registers information related to the detected object (or feature quantities representing that information) in the map database 40 (database) as landmark-related information (or feature quantities representing that information). Furthermore, the landmark-related information includes information such as the shape, color, and pattern of the object that can be determined from the photographic image. It also includes information such as the type of object that can be determined from the photographic image. Moreover, the landmark-related information includes position information of the object used to indicate the reference position of the position inference device 10 (autonomous mobile robot). Additionally, if the positional relationship with the registered landmark is clear, additional landmarks can be appropriately registered in the map database 40.
[0023] The acquisition unit 12 acquires landmark-related information in text format, separately from the landmark-related information detected by the object detection unit 11, and registers it in the map database 40. For example, the acquisition unit 12 acquires detailed information related to the landmark in text format, such as the shop name if the landmark is a shop, or the specific type if the landmark is furniture or tableware. The landmark-related text format information (or the feature quantity representing the information) acquired by the acquisition unit 12 is associated with the corresponding landmark and registered in the map database 40.
[0024] Information in text format related to the landmark, acquired by the acquisition unit 12, is sent from the operation terminal 20, for example.
[0025] The operating terminal 20 is a communication terminal owned by the user or temporarily assigned to the user, such as a PC (Personal Computer) terminal, a smartphone or tablet terminal, or a dedicated communication terminal prepared for this system.
[0026] For example, a user can operate the terminal 20 by touching its display with a stylus or finger, or by using the terminal 20's mouse or keyboard to input text-formatted information related to landmarks registered in the map database 40. The terminal 20 receives the landmark-related text-formatted information and transmits it to the location inference device 10 via the network 50.
[0027] In addition, the acquisition unit 12 can also acquire text-formatted information related to landmarks from external devices other than the operation terminal 20.
[0028] The conversion unit 13 uses a multimodal recognition model, such as an encoder capable of inputting both language and images, to convert information related to objects detected by the object detection unit 11 (including objects registered as landmarks in the map database 40) into vector-defined feature quantities. Furthermore, the conversion unit 13 uses, for example, a multimodal recognition model to convert text-formatted information related to landmarks acquired by the acquisition unit 12 into vector-defined feature quantities. This allows for comparison of information related to objects detected by analyzing photographic images and information related to objects acquired in text format. Moreover, the calculation unit 16 calculates first configuration information, which defines the configuration relationships between objects detected by the object detection unit 11 and surrounding objects. More specifically, the calculation unit 16 represents the objects detected by the object detection unit 11 and multiple surrounding objects as multiple nodes, representing them as a node-edge graph connected by edges.
[0029] Then, a semantic graph is constructed by connecting objects that are closer than a certain distance threshold to each other with edges. Furthermore, for each node representing an object, starting from that object's node, all possible paths for moving a fixed number of steps (e.g., 3 steps) to nodes representing adjacent surrounding objects connected by edges are investigated, and a histogram (semantic histogram) is constructed by counting the patterns of the categories (types) of the objects traversed. The resulting histogram is calculated using the results obtained through modeling methods such as L2 regularization as descriptors, and the configuration information (first configuration information) obtained by modeling or quantifying the configuration relationships of multiple surrounding objects within a specified object is calculated. Similarly, the calculation unit 16 applies this method to the configuration relationships between an object and its surrounding objects, acquired by the acquisition unit 12, to calculate second configuration information, which is information obtained by modeling or quantifying the configuration relationships between the object and its surrounding objects. This configuration information (first configuration information and second configuration information) can be represented as vectors capable of numerical computation.
[0030] The object inference unit 14 compares the object detected by the object detection unit 11 with the landmarks registered in the map database 40 to infer whether the object detected by the object detection unit 11 corresponds to a landmark registered in the map database 40.
[0031] Specifically, the object inference unit 14 infers whether the object detected by the object inference unit 14 corresponds to a landmark registered in the map database 40 by comparing the feature quantity (first feature quantity) representing the object detected by the object detection unit 11 with the feature quantity (second feature quantity) representing the landmark registered in the map database 40.
[0032] More specifically, firstly, the object inference unit 14 infers whether the object detected by the object detection unit 11 corresponds to a landmark registered in the map database 40 by comparing a feature quantity (first feature quantity) representing an object detected by the object detection unit 11 with a second feature quantity representing a landmark registered in the map database 40. Here, the second feature quantity includes a feature quantity representing text-formatted information related to the landmark acquired by the acquisition unit 12. Furthermore, the second feature quantity may also include a feature quantity representing information related to the landmark obtained from the photographic image. In this case, the second feature quantity contains more detailed information about the landmark. Therefore, the object inference unit 14 can accurately infer whether the object detected by the object detection unit 11 corresponds to a landmark registered in the map database 40. Furthermore, as landmark-related information, feature quantities representing at least one of the text and the image can be used.
[0033] Furthermore, when multiple landmarks are registered in the map database 40, the object inference unit 14 narrows down the multiple landmarks to the landmarks with the highest probability of being detected by the object detection unit 11.
[0034] Specifically, the object inference unit 14 compares a first feature representing an object detected by the object detection unit 11 with multiple baseline features representing multiple landmarks registered in the map database 40, and narrows it down to landmarks with a second feature whose consistency with the first feature is above a predetermined threshold, as landmarks with a high probability of being detected by the object detection unit 11. More specifically, the consistency is calculated by using the inner product (cosine similarity) of the first feature and the second feature as the similarity. Alternatively, the object inference unit 14 may also compare a first feature representing an object detected by the object detection unit 11 with multiple second features representing multiple landmarks registered in the map database 40, and infer the landmark with the second feature that has the highest consistency with the first feature as a landmark with a high probability of being detected by the object detection unit 11.
[0035] Here, as described above, the second feature quantity includes not only features representing information related to landmarks obtained from the photographic image, but also features representing text-format information related to the landmarks acquired by the acquisition unit 12. That is, the second feature quantity contains more detailed information about the landmark. Therefore, the object inference unit 14 can narrow down the landmarks from multiple landmarks to those with a high probability of being detected by the object detection unit 11. Furthermore, the object inference unit 14 can narrow down the landmarks based on both the comparison result of the first and second feature quantities described above and the comparison result of the first and second configuration information calculated by the calculation unit 16. In this case, the object inference unit 14 calculates the similarity as the inner product (cosine similarity) of the first and second configuration information calculated by the calculation unit 16. Moreover, by calculating the weighted sum of the similarities of the separately calculated first and second feature quantities, a similarity based on the comparison of both feature quantities and configuration information is calculated, thereby calculating the consistency score. Then, landmarks with a consistency score of a predetermined threshold or higher are narrowed down to those with a high probability of being detected by the object detection unit 11.
[0036] Furthermore, when multiple objects are detected by the object detection unit 11, the object inference unit 14 narrows down the multiple landmarks to the landmarks corresponding to each object detected by the object detection unit 11.
[0037] Furthermore, the object inference unit 14 evaluates multiple candidate correspondences between multiple objects detected by the object detection unit 11 and multiple landmarks registered in the map database 40 using a score function that scores the likelihood. The object inference unit 14 then selects the candidate with the highest score from among the multiple candidate correspondences between multiple objects detected by the object detection unit 11 and multiple landmarks registered in the map database 40. Moreover, the object inference unit 14 can maintain a dataset 100 that establishes correspondence associations between objects detected by the object detection unit 11 and specified objects acquired by the acquisition unit 12, where the object inference unit 14 infers them as corresponding candidates (the same). Therefore, even in situations where the object detection unit 11 moves in real-time within the environment, a highly reliable dataset can be readily utilized based on the situation. More specifically, the dataset is represented and maintained using a compatibility / consistency graph, which represents each corresponding candidate in the dataset as a node and connects them with edges. In this scenario, when two candidate correspondences corresponding to two nodes satisfy a predefined spatial consistency condition, edges are opened between these nodes. Here, the condition that the distances between objects contained in the two correspondences are approximately equal on both the observation and the map can be used. The set of mutually consistent correspondences is efficiently obtained as a maximal clique of the consistency graph. A maximal clique is defined as a complete subgraph of a graph in which adding any adjacent node will not result in a complete graph. The set of multiple candidate correspondences obtained as maximal cliques is fractionalized based on the sum of the similarities between the candidate correspondences contained in that set. Thus, especially when the number of obtained correspondences is equal, a set of correspondences with higher overall similarity can be determined. Therefore, the data of these highly reliable candidate correspondences can be used as candidates for inliners—a highly reliable dataset—for position inference and viewpoint transformation (coordinate transformation) of the position inference device 10, after estimating their likelihood.
[0038] The position inference unit 15 infers the position of the position inference device 10 (in other words, the autonomous mobile robot equipped with the position inference device 10 or the camera 30 equipped with the position inference device 10) by solving the Perspective-n-Point (PnP) problem based on the inference results of the object-based inference unit 14. That is, the position inference unit 15 infers the position of the position inference device 10 based on the viewpoint deviation (the amount and direction of viewpoint movement) between the object detected by the object detection unit 11 and the corresponding landmark registered in the map database 40.
[0039] Thus, the location inference system 1 of this invention not only acquires and references information related to landmarks obtained from the photographic images of the camera 30, but also acquires and references text-formatted information related to landmarks obtained via the operation terminal 20, etc., enabling it to accurately infer whether an object detected in the photographic images of the camera 30 corresponds to a landmark registered in the map database 40. Therefore, the location inference system 1 of this invention can narrow down from multiple landmarks to those with a high probability of being an object obtained from the photographic images of the camera 30. As a result, even when multiple objects are detected in the photographic images of the camera 30, the location inference system 1 of this invention can reduce the computational load required to determine the correspondence between these multiple objects and multiple landmarks registered in the map database 40.
[0040] Next, use Figure 2 The operation of the position inference device 10 will be explained. Figure 2 This is a flowchart illustrating the operation of the position inference device 10.
[0041] First, the location inference device 10 detects objects around the autonomous mobile robot positioned at the reference position from the photographed image of the camera 30, and registers the information related to the detected objects (or the feature quantity representing the information) in the map database 40 as landmark-related information (or the feature quantity representing the information) (step S101).
[0042] Then, the location inference device 10 acquires text-formatted information related to landmarks registered in the map database 40 via the operation terminal 20 (step S102). The acquired text-formatted information related to the landmarks (or the feature quantities representing the information) are associated with the corresponding landmarks and registered in the map database 40.
[0043] Figure 3 This is a diagram representing an example of the registered content in map database 40. In Figure 3 In the example, map database 40 contains information related to eight landmarks M1 to M8. Specifically, map database 40 contains the location information of each landmark M1 to M8 and the associated text-format information (language tags).
[0044] Then, after the position inference device 10 moves along with the autonomous mobile robot, it detects the objects around the position inference device 10 (in other words, the autonomous mobile robot equipped with the position inference device 10) (step S103).
[0045] Here, the location inference device 10 uses, for example, a multimodal recognition model to convert the relevant information of objects (including objects registered as landmarks in the map database 40) detected by analyzing the photographic images of the camera 30 into a first feature quantity defined by vectors (step S104). Furthermore, the location inference device 10 uses, for example, a multimodal recognition model to convert the landmark-related text-format information obtained via the operation terminal 20, etc., into a second feature quantity defined by vectors (step S104). Thus, it is possible to compare the relevant information of objects detected by analyzing photographic images with the relevant information of objects obtained in text format.
[0046] Then, the location inference device 10 compares the object detected in the photographed image from the camera 30 with the landmark registered in the map database 40 to infer whether the object detected in the photographed image from the camera 30 corresponds to the landmark registered in the map database 40.
[0047] Specifically, the location inference device 10 infers whether the object detected by the object inference unit 14 corresponds to a landmark registered in the map database 40 by comparing a first feature quantity representing an object detected from the photographed image of the camera 30 with a second feature quantity representing a landmark registered in the map database 40 (step S105).
[0048] Here, the second feature quantity includes not only features representing information related to landmarks obtained from the photographic image, but also features representing text-formatted information related to landmarks obtained via the operating terminal 20, etc. That is, the second feature quantity contains more detailed information about the landmark. Therefore, the location inference device 10 can accurately infer whether an object obtained from the photographic image corresponds to a landmark registered in the map database 40. Furthermore, as landmark-related information, features representing at least one of the text and the image can be used.
[0049] Figure 4 This is a graph showing the relationship between landmarks registered in map database 40 and objects detected from photographic images taken by camera 30. Figure 4 In the example, information related to eight landmarks M1 to M8 is registered in the map database 40, and three objects T1 to T3 are detected from the images captured by camera 30. Additionally, in Figure 4 As a comparative example, an example is also shown where information in text format related to landmarks is not registered in map database 40.
[0050] First of all, Figure 4In the comparative example shown in the figure above, no text-formatted information related to the landmarks M1 to M8 registered in the map database 40 is assigned. In this case, the location inference device of the comparative example narrows down the landmarks M1 to M8 registered in the map database 40 to three landmarks M1, M2, and M7 that are highly likely to be objects T1, to two landmarks M4 and M5 that are highly likely to be objects T2, and to three landmarks M3, M6, and M8 that are highly likely to be objects T3. Therefore, there are 18 candidate correspondences between objects T1 to T3 detected from the photographed images of camera 30 and landmarks M1 to M8 registered in the map database 40. Therefore, in the location inference device of the comparative example, the computational burden required to determine the correspondence between objects T1 to T3 detected from the photographed images of camera 30 and landmarks M1 to M8 registered in the map database 40 is increased.
[0051] In contrast, Figure 4 In the example below, landmarks M1 to M8 registered in map database 40 are assigned text-formatted information related to the landmark. Specifically, landmark M1 is assigned the text format "Kid's table", landmark M2 is assigned the text format "Dusty table", landmark M3 is assigned the text format "Kitchen table", landmark M4 is assigned the text format "Digital clock", landmark M5 is assigned the text format "Big tall old clock", and landmark M6 is assigned the text format "Black short shelf".
[0052] In this case, the location inference device 10 narrows down the landmarks M1 to M8 registered in the map database 40 to a landmark M1 that is highly likely to be object T1, to a landmark M4 that is highly likely to be object T2, and to a landmark M6 that is highly likely to be object T3. Therefore, there are candidates for a correspondence between objects T1 to T3 detected in the photographed images from camera 30 and landmarks M1 to M8 registered in the map database 40. Thus, the location inference device 10 can reduce the computational burden required to determine the correspondence between objects T1 to T3 detected in the photographed images from camera 30 and landmarks M1 to M8 registered in the map database 40.
[0053] Then, the position inference device 10 infers the position of the autonomous mobile robot equipped with the position inference device 10 (in other words, the camera 30 equipped with the position inference device 10) based on the inference results of the correspondence between objects T1-T3 and landmarks M1-M8 (step S106). Furthermore, the position inference device 10 can infer the camera's position using the correspondence contained in the dataset 100. More specifically, the data of highly reliable candidate correspondences contained in the dataset 100 can be used as candidates for highly reliable datasets, i.e., inliners, for position inference and viewpoint transformation (coordinate transformation) of the position inference device 10, after estimating their likelihood. Alternatively, the final camera position pose can be calculated by minimizing the error between the sets of correspondences (candidate correspondences) contained in the dataset 100 using weighted least squares. As weighting coefficients, the following two pieces of information are considered: correspondence similarity and observation completeness. Correspondence similarity: In error calculation, groups of candidate correspondences with high similarity between candidate correspondences are given greater weight. Observation completeness: Greater emphasis is placed on observing candidate sizes that are smaller than the sizes of known map landmarks. The product of these values is used as the final weighting coefficient. This mitigates the impact of low-likelihood (potentially erroneous) correspondences and the influence of only partially observed objects on the accuracy of pose calculation, enabling robust pose calculation.
[0054] Thus, the location inference system 1 of this invention not only acquires and references information related to landmarks obtained from the photographic images of the camera 30, but also acquires and references text-formatted information related to landmarks obtained via the operation terminal 20, etc., enabling it to accurately infer whether an object detected in the photographic images of the camera 30 corresponds to a landmark registered in the map database 40. Therefore, the location inference system 1 of this invention can narrow down from multiple landmarks to those with a high probability of being an object obtained from the photographic images of the camera 30. As a result, even when multiple objects are detected in the photographic images of the camera 30, the location inference system 1 of this invention can reduce the computational load required to determine the correspondence between these multiple objects and multiple landmarks registered in the map database 40.
[0055] In addition, location inference system 1 can also be derived from, for example... Figure 5 The floor plan MP1, which is set up in stations, shopping malls, public facilities, etc., obtains text-formatted information related to landmarks A1 to A4, such as facility names, building names, and room names. This enables high-performance position inference of the camera 30 within the floors represented by the floor plan MP1.
[0056] Furthermore, the present invention can realize part or all of the processing of the position inference device 10 or the position inference system 1 having the position inference device 10 by having the central processing unit (CPU) execute a computer program.
[0057] When the above-described program is read into a computer, it includes a set of commands (or software code) for causing the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of limitation and not example, a computer-readable medium or a tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray disc (registered trademark) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices.
[0058] The program may be transmitted on a temporary computer-readable medium or a communication medium. As a limiting but not restrictive example, a temporary computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
[0059] Symbol Explanation 1-Location inference system, 10-Location inference device, 11-Object detection unit, 12-Acquisition unit, 13-Conversion unit, 14-Object inference unit, 15-Location inference unit, 16-Calculation unit, 20-Operation terminal, 30-Camera, 40-Map database, 50-Network.
Claims
1. An object inference device, characterized in that, have: The acquisition unit acquires text-formatted information related to the specified object and surrounding objects located around the object from a database that stores text-formatted information related to the specified object and the object's location information in a corresponding association. An object detection unit that detects an object and surrounding objects that are objects surrounding the object. The conversion unit converts the object detected by the object detection unit into a first feature quantity, which is defined as data representing the object, and converts the text format information related to the specified object acquired by the acquisition unit into a second feature quantity, which is defined as data representing the object. The computing unit calculates first configuration information and second configuration information, wherein the first configuration information defines the configuration relationship between the object and the surrounding objects detected by the object detection unit, and the second configuration information defines the configuration relationship between the object and the surrounding objects acquired by the acquisition unit. and The object inference unit infers whether the object detected by the object detection unit is the same as the specified object by comparing a first feature quantity representing the object detected by the object detection unit with a second feature quantity representing the specified object, and by comparing the first configuration information calculated by the calculation unit with the second configuration information.
2. The object inference device according to claim 1, characterized in that, The object inference unit performs the inference for each of the plurality of objects detected by the object detection unit. Furthermore, the dataset that establishes a corresponding association between the objects inferred as the same by the object inference unit and the objects detected by the object detection unit and the specified objects is maintained.
3. A location inference system, characterized in that, have: The object inference apparatus of claim 1, wherein the object inference unit performs the inference on each of a plurality of objects detected by the object detection unit; and The position inference unit infers the position of the object inference device based on the positional relationships between the object and the specified object detected by the object detection unit and the object inferred as the same by the object inference unit.
4. A method for inferring an object, characterized in that, The computer performs the following processing: Retrieve text-formatted information related to the specified object and surrounding objects located around it from a database that stores text-formatted information associated with the specified object and the object's location information. Detect an object and the objects surrounding that are objects surrounding the object; The detected object is converted into a feature quantity, i.e., the first feature quantity, which is defined as the data representing the object; and the acquired text format information related to the specified object is converted into a feature quantity, i.e., the second feature quantity, which is defined as the data representing the object. Calculate the first configuration information and the second configuration information, wherein the first configuration information defines the configuration relationship between the detected object and the surrounding objects, and the second configuration information defines the configuration relationship between the acquired object and the surrounding objects; and By comparing the first feature representing the detected object with the second feature representing the specified object, and by comparing the calculated first configuration information with the second configuration information, it is inferred whether the detected object is the same as the specified object.
5. A control program, characterized in that, The computer will perform the following processing: Retrieve text-formatted information related to the specified object and surrounding objects located around it from a database that stores text-formatted information associated with the specified object and the object's location information. Detect an object and the objects surrounding that are objects surrounding the object; The detected object is converted into a feature quantity, i.e., the first feature quantity, which is defined as the data representing the object; and the acquired text format information related to the specified object is converted into a feature quantity, i.e., the second feature quantity, which is defined as the data representing the object. Calculate the first configuration information and the second configuration information, wherein the first configuration information defines the configuration relationship between the detected object and the surrounding objects, and the second configuration information defines the configuration relationship between the acquired object and the surrounding objects; and By comparing the first feature representing the detected object with the second feature representing the specified object, and by comparing the calculated first configuration information with the second configuration information, it is inferred whether the detected object is the same as the specified object.
Citation Information
Patent Citations
JP1974085166A