Information processing device, information processing method, and program
The information processing device estimates the position of moving objects by analyzing images and using placement information to identify a viewpoint, addressing the time-consuming map preparation in conventional methods and enhancing estimation accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CYBER AGENT
- Filing Date
- 2025-09-10
- Publication Date
- 2026-05-20
AI Technical Summary
Conventional methods for estimating the position of a moving object require the preparation of a detailed map for geometric matching with point cloud data, which is time-consuming.
An information processing device estimates the position of a moving object by analyzing target images captured by imaging devices, detecting observed objects, and using placement information to identify a viewpoint position, eliminating the need for a detailed map and leveraging a simple arrangement of objects in the environment.
Enables accurate and efficient position estimation of moving objects by simplifying the map preparation process and utilizing a large-scale generative model for object detection and candidate location extraction, reducing implementation effort.
Smart Images

Figure 0007863243000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] In recent years, the development of technologies for estimating the position (self-position) of a moving object has advanced. For example, in Non-Patent Document 1, a method for estimating the position of a moving object is proposed by geometrically matching point cloud data measured by sensors such as LiDAR (Light Detection and Ranging) and radar with a map. A method for estimating the position of a moving object by geometrically matching point cloud data measured by sensors such as LiDAR (Light Detection and Ranging) and radar with a map has been proposed.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] According to the conventional method, the position of a moving object can be estimated. However, the inventor of the present invention has found that the conventional method has the following problems. That is, in the conventional estimation method, as a precondition, a detailed map for geometrically matching with point cloud data needs to be prepared. Preparing this map is time-consuming.
[0005] In one respect, this disclosure has been made in consideration of these circumstances. One of the purposes of this disclosure is to provide a technique for estimating the position of a moving object in a simple manner. [Means for solving the problem]
[0006] This disclosure adopts the following configuration to solve the aforementioned problems. Note that the following configurations can be combined as appropriate.
[0007] An information processing device relating to one aspect of this disclosure includes a control unit. The control unit is configured to acquire one or more target images captured by one or more imaging devices attached to a moving body, to analyze the acquired one or more target images to detect an observation object that appears in at least one of the acquired one or more target images, to estimate the position of the moving body by searching for a viewpoint position that matches the detection result of the observation object in the environment in which the moving body is moving, based on placement information, and to output information regarding the position estimation result. The placement information is configured to indicate the arrangement of objects present in the environment.
[0008] In this configuration, the position of the moving object is estimated using the positions of objects observed from the moving object as clues. If the arrangement of objects present in the environment is identified, the position of the moving object can be estimated. In other words, the only preparation required for position estimation is to prepare a simple map (arrangement information) showing the arrangement of objects. Therefore, with this configuration, the position of the moving object can be estimated in a simple manner.
[0009] In the information processing device relating to the above aspect, one or more imaging devices may be composed of multiple imaging devices facing different directions from each other. One or more target images may be composed of multiple target images captured by multiple imaging devices. In a position estimation method that uses the position of an observed object as a clue, the more variations there are in the arrangement of the detected observed objects, The viewpoint position of the imaging device can be precisely determined. With this configuration, multiple imaging devices face different directions from each other, allowing for the acquisition of target images over a wide field of view (the range in which observed objects are captured). Therefore, by reflecting the capture results of observed objects present in each direction within the wide field of view, an improvement in the accuracy of position estimation can be expected.
[0010] In the information processing device relating to the above aspect, the multiple imaging devices may consist of three or more imaging devices. With this configuration, by using three or more imaging devices, it is possible to capture observed objects in three or more directions. This makes it possible to expect appropriate position estimation even in situations where there are three or more unknowns (for example, XY coordinates and orientation are the targets for estimation).
[0011] In the information processing device relating to the above aspect, the search for the position of a viewpoint in the environment may be comprised of obtaining simulation results that show objects to be imaged from each of the multiple candidate positions in the environment, generated by simulating imaging by one or more imaging devices from each of the multiple candidate positions in the environment based on placement information; comparing the detection results of observed objects with the obtained simulation results; and, according to the results of the comparison, extracting candidate positions from the multiple candidate positions as viewpoint positions that match the detection results of observed objects, where objects matching the detected observed objects are imaged. With this configuration, the convergence of the position estimation process can be ensured by limiting the search range of positions to multiple candidate positions.
[0012] In the information processing device relating to the above aspect, detecting observed objects and extracting candidate locations may be performed by providing a large-scale generative model with a prompt that includes one or more acquired target images, acquired simulation results, a first instruction commanding the detection of observed objects from one or more target images, and a second instruction commanding the extraction of candidate locations from a plurality of candidate locations in which an object matching the observed object is imaged, thereby causing the large-scale generative model to detect observed objects and extract candidate locations. With this configuration, by using a large-scale generative model, the effort of preparing a dedicated program for performing the detection of observed objects and the extraction of candidate locations can be eliminated. This is expected to reduce the effort required for implementation.
[0013] In the information processing device relating to the above aspect, matching and extraction may be performed by providing a large-scale generative model with a prompt that includes instructions to extract candidate locations from a plurality of candidate locations in which an object matching the detection result of the observed object and the simulation result is imaged, thereby causing the large-scale generative model to perform matching and extraction. With this configuration, by using a large-scale generative model, the effort of preparing a dedicated program for performing candidate location extraction can be eliminated. This is expected to reduce the effort required for implementation.
[0014] In the information processing device relating to the above aspect, detecting an observed object may be performed by providing a large-scale generative model with a prompt that includes one or more acquired target images and an instruction to detect an observed object from the one or more target images, thereby causing the large-scale generative model to detect the observed object. With this configuration, by using a large-scale generative model, the effort of preparing a dedicated program for performing the detection of observed objects can be eliminated. This can be expected to reduce the effort required for implementation.
[0015] In the information processing device relating to the above aspect, the control unit may be configured to further estimate the position of the moving object by estimation methods other than the estimation method of searching for the position of a viewpoint that matches the detection result of the observed object. In the position estimation method used, it may be difficult to estimate the position if the observed object is not detected. With this configuration, position estimation can be supplemented by using other estimation methods in combination.
[0016] In the information processing device relating to the above aspect, estimating the position of a moving object using another estimation method may consist of estimating the position of the moving object by acquiring sensing data from sensors deployed on the moving object to observe its movement, and calculating the relative amount of movement of the moving object from the acquired sensing data. A position estimation method based on relative amount of movement may be faster than a position estimation method that includes image analysis (detection of observed objects). With this configuration, the frequency of position estimation can be supplemented by using a position estimation method based on relative amount of movement in combination.
[0017] In the information processing device relating to the above aspect, the moving object may be a living organism. With this configuration, when the moving object is a living organism, the position of the moving object can be estimated by a simple method.
[0018] In the information processing device relating to the above aspect, the organism may be an organism moving within the store. The environment may be the space within the store. With this configuration, when a moving object (organism) moves within the store, the position of the moving object can be estimated by a simple method.
[0019] In the information processing device relating to the above aspect, the moving object may be a moving device. With this configuration, when the moving object is a device, the position of the moving object can be estimated by a simple method.
[0020] In the information processing device relating to the above aspect, the device may be a device used within a store. The environment may be the space within the store. With this configuration, when a mobile object (device) moves within the store, the position of the mobile object can be estimated by a simple method.
[0021] The form of the information processing device described herein is not limited to the above-described information processing device. As another form of the information processing device relating to each of the above aspects, one aspect of the disclosure may be an information processing method that implements all or part of the above-described configurations, a program, or a machine-readable storage medium that stores such a program. Here, a machine-readable storage medium may be a non-temporary medium that stores information such as programs by electrical, magnetic, optical, mechanical, or chemical action. A non-temporary storage medium may include storage media (CDs, DVDs, semiconductor memory, etc.), auxiliary storage devices of a computer, external storage devices connected to a computer, etc.
[0022] For example, an information processing method relating to one aspect of this disclosure may be performed by a computer. The information processing method may include acquiring one or more target images captured by one or more imaging devices attached to a moving object, detecting an observed object in at least one of the acquired target images by analyzing the acquired one or more target images, estimating the position of the moving object by searching for a viewpoint position that matches the detection result of the observed object in the environment in which the moving object is moving, based on placement information, and outputting information regarding the position estimation result. The placement information may be configured to indicate the arrangement of objects present in the environment.
[0023] Also, for example, a program according to one aspect of the present disclosure may be a program for causing a computer to execute an information processing method. The information processing method includes obtaining one or more target images captured by one or more imaging devices attached to a moving object, analyzing the obtained one or more target images to detect an observation object appearing in at least any one of the obtained one or more target images, and based on the placement information, searching for a position of a viewpoint that conforms to the detection result of the observation object in the environment where the moving object moves, thereby estimating the position of the moving object, and outputting information regarding the estimation result of the position. The placement information may be configured to indicate the placement of objects existing in the environment. and may include outputting information regarding the estimation result of the position. The placement information may be configured to indicate the placement of objects existing in the environment.
Effects of the Invention
[0024] According to one aspect of the present disclosure, it is possible to provide a technique for estimating the position of a moving object in a simple manner.
Brief Description of the Drawings
[0025] [[ID=I5]] [Figure 1] FIG. 1 schematically shows an example of a scene to which the present disclosure is applied. [Figure 2] FIG. 2 schematically shows a specific example of position estimation. [Figure 3] FIG. 3 schematically shows an example of a moving object. [Figure 4] FIG. 4 schematically shows an example of a method for detecting an observation object. [Figure 5] FIG. 5 schematically shows an example of a method for searching for a position of a viewpoint that conforms to the detection result of an observation object. [Figure 6] FIG. 6 schematically shows an example of a method for collating the detection result of an observation object with a simulation result. [Figure 7] FIG, 7 schematically shows an example of a method for collating the detection result of an observation object with a simulation result. [Figure 8] FIG. 8 schematically shows an example of a scene for estimating a position by another estimation method. [Figure 9]Figure 9 schematically shows an example of the hardware configuration of an information processing device. [Figure 10] Figure 10 schematically shows an example of the software configuration of an information processing device. [Figure 11] Figure 11 is a flowchart showing an example of a processing procedure for an information processing device. [Figure 12] Figure 12 is a flowchart showing an example of a processing procedure (subroutine) for position estimation using other estimation methods. [Figure 13] Figure 13 shows a simulation example of imaging from candidate positions during viewpoint search in the first experimental example. [Figure 14] Figure 14 shows the positional measurement errors in the comparative example and the experimental example. [Modes for carrying out the invention]
[0026] Hereinafter, embodiments relating to one aspect of this disclosure will be described with reference to the drawings. However, the embodiments described below are merely illustrative in all respects of this disclosure. Various improvements or modifications may be made without departing from the scope of this disclosure. In implementing this disclosure, specific configurations may be adopted as appropriate depending on the embodiment. In this embodiment, the data appearing is described in natural language, but more specifically, it is specified in pseudo-language, commands, parameters, machine code, electrical signals, etc., that can be recognized by machines such as computers.
[0027] §1 Examples of Application Figure 1 schematically shows an example of a scenario to which this disclosure applies. The information processing device 1 according to this embodiment is one or more computers configured to estimate the position of a moving object MB.
[0028] In this embodiment, the position estimation of the mobile body MB may be performed at any time while the mobile body MB is moving within the environment 4. One or more imaging devices 2 are attached to the mobile body MB. The environment 4 is provided with arrangement information 40. The arrangement information 40 is configured to indicate the arrangement of objects 31 present in the environment 4.
[0029] The information processing device 1 acquires one or more target images 20 captured by one or more imaging devices 2 attached to the mobile body MB. The information processing device 1 analyzes the acquired one or more target images 20 to detect an observed object 33 that appears in at least one of the acquired one or more target images 20. The detection result 35 of the observed object 33 is obtained. The observed object 33 is an object 31 present in the environment 4 that is captured by the target image 20 of the imaging device 2 (i.e., appears in one or more target images 20) and detected by image analysis. Based on the placement information 40, the information processing device 1 estimates the position of the mobile body MB by searching for a viewpoint position that matches the detection result 35 of the observed object 33 in the environment 4 in which the mobile body MB is moving. As a result, the information processing device 1 obtains the position estimation result 55. The information processing device 1 outputs information related to the position estimation result 55.
[0030] Figure 2 schematically shows a specific example of position estimation according to this embodiment. In the example in Figure 2, objects C1 to C14 exist in environment 4, three imaging devices 2 are attached to the mobile body MB, and it is assumed that objects (C3, C4, C10) are captured in the target image 20 of each imaging device 2. In this scenario, each object (C3, C4, C10) is an example of an observed object 33. The position of each object (C3, C4, C10) can be obtained from the arrangement information 40. From the positional relationship of each object C1 to C14 shown by the arrangement information 40, a position 5 can be searched for as a position where each object (C3, C4, C10) can be captured by each imaging device 2. This position 5 is an example of a viewpoint position that matches the detection result 35. In this scenario, the information processing device 1 can estimate that the mobile body MB is located at the searched position 5 (i.e., position 5 becomes the estimation result 55).
[0031] In this embodiment, as shown in the example in Figure 2, the position of the mobile body MB can be estimated by detecting the observation object 33 in the target image 20 and using the position of the detected observation object 33 indicated by the placement information 40 as a clue. Therefore, if the arrangement of objects 31 present in the environment 4 is identified, the position of the mobile body MB can be estimated. That is, although a detailed map may be used, the preliminary preparation for performing position estimation is sufficient with a simplified map (placement information 40) showing the arrangement of objects 31, as shown in the example in Figure 2. Thus, according to this embodiment, the position of the mobile body MB can be estimated in a simplified manner.
[0032] [Placement information] The configuration of the placement information 40 is not particularly limited and can be determined as appropriate depending on the embodiment, as long as the placement of the objects 31 present in environment 4 can be identified. For example, the placement information 40 may be configured to geometrically define the relative positional relationship of each object 31 present in environment 4. For example, the placement information 40 may consist of a set (list) of combinations of identification information and position of the objects 31 present in environment 4. The placement information 40 may also be called a placement map. The placement information 40 may be configured to illustrate the placement of each object 31 as shown in Figure 2, or it may not be configured in that way.
[0033] The identification information may consist of any information that can identify object 31, such as an identifier, name (proper noun, product name, etc.), type (category), etc. The identification information may also be called a label. The identification information may be defined to have a correspondence with the detection result 35 of the observed object 33 so that the position indicated by the placement information 40 can be referenced from the detection result 35 of the observed object 33. The correspondence may be defined arbitrarily.
[0034] For example, if the detection result 35 of the observed object 33 indicates the name of the observed object 33, the identification information may include the name of object 31. If the detection result 35 indicates the type of the observed object 33, the identification information may include the type of object 31. In this way, the identification information may be defined to correspond one-to-one with the detection result 35. However, the correspondence between the detection result 35 and the identification information is not limited to such a one-to-one relationship. For example, consider a scenario where the identification information consists of the type of instant food, etc., while the detection result 35 is obtained as the name of cup noodles, etc. In this scenario, detection Although the result 35 and the identification information do not have a one-to-one relationship, it is possible to compare the detection result 35 with the identification information, for example, cup noodles are classified as instant food. Therefore, as long as it is possible to compare the detection result 35 with the placement information 40 (identification information), the correspondence between the detection result 35 and the identification information does not necessarily have to be a one-to-one relationship and may be defined as appropriate depending on the embodiment.
[0035] The position of object 31 may be expressed in any format. For example, the position of object 31 may be expressed in two-dimensional or three-dimensional coordinates. The coordinates may be absolute coordinates or relative coordinates, as long as they are similar to the positional relationships in real space. That is, the spatial scale in the placement information 40 may or may not match the scale of real space. The position (coordinates) may be expressed as a point or as a range.
[0036] The location information 40 may be generated by any method. For example, the location information 40 may be generated manually. For example, if the environment 4 is a store, the store's information map may be used as is as the location information 40. The information map may include, for example, a floor guide, a floor map, a map of the product shelves (a map of the displayed items), etc. The location information 40 may be generated automatically or manually from the information map. Also, for example, the location information 40 may be generated by sensing the environment 4 using sensors such as an imaging device and a positioning module. For example, the location information 40 may be generated by imaging the environment 4 with an imaging device while positioning with a positioning module, detecting objects 31 in the captured image, and estimating the position of the detected objects 31 from the positioning results throughout the entire environment 4. The type of positioning module may be arbitrarily selected as long as the position can be measured. The positioning method may be appropriately selected from known methods such as satellite positioning and positioning using wireless communication with a wireless access point. The wireless access point may include, for example, a beacon and a base station. In one example, the placement information 40 may also include any information about objects other than object 31 that exist in the environment 4 (for example, the location of obstacles).
[0037] [environment] Environment 4 may refer to the physical space (domain) in which the mobile body MB moves. Environment 4 may include, for example, indoor spaces, outdoor spaces, etc. Indoor spaces may include spaces inside stores. Stores may include commercial facilities such as supermarkets, convenience stores, drugstores, and restaurants. Stores may also include mixed commercial facilities such as shopping malls. Outdoor spaces may include, for example, roads, plazas, parks, amusement parks, etc. Any type of object 31 may exist in Environment 4. The arrangement of the object 31 is described by the arrangement information 40. Environment 4 may define the range of movement of the mobile body MB, and the arrangement information 40 may provide the spatial context that is the subject of position estimation.
[0038] [Object] Object 31 may be any entity present in the environment 4 that is visually detectable by the imaging device 2. Visually detectable means detectable by image analysis. The entity may include, for example, objects, other features, etc. Object 31 may also be called a landmark. Object 31 may be detected as a single object or as a group of objects. Object 31 may be detectable as a component of an object, such as symbolic information or graphic information attached to an object. Symbolic information may include letters, numbers, symbols, etc. Graphic information may include drawings such as characters. Symbolic information and graphic information may be attached to the ground, such as road signs or parking space numbers. Symbolic information and graphic information may be attached to structures such as pillars or walls.
[0039] The type of object 31 is not particularly limited and can be appropriately selected depending on the embodiment. Object 31 may include, for example, products / merchandise, tools, equipment, building structures, monuments, works of art and crafts, signs, advertisements, other displays, and parts thereof. Products / merchandise may include, for example, foodstuffs, beverages, daily necessities, clothing, electrical appliances, etc. Tools may include, for example, household furniture such as desks, chairs, and shelves. Equipment may include, for example, kitchen equipment, refrigeration equipment, cash registers, etc. Building structures may include, for example, entrances, doors, windows, columns, etc. Works of art and crafts may include, for example, statues, paintings, etc. Signs may include, for example, signs inside stores, road signs, etc. Other displays may include, for example, parking space numbers, road markings, etc.
[0040] The object 31 to be detected may be appropriately selected depending on the embodiment, such as the operational scenario. For example, the mobile body MB may be an entity that moves within the store. Accordingly, the object 31 may be an entity that exists within the store. For example, the object 31 may include goods sold in the store, signs, advertisements, combinations thereof, and at least a part of the symbolic information or at least a part of the graphic information attached to them.
[0041] [Imaging device] The type of imaging device 2 is not particularly limited, and can be appropriately selected depending on the embodiment, as long as it can generate an image in which object 31 (observed object 33) can be detected. The imaging device 2 may include any sensor that acquires data in the form of an image or image representation, such as an RGB camera, a hemispherical camera, a 360-degree camera, a depth sensor, an infrared sensor, radar, or LiDAR.
[0042] The number of imaging devices 2 arranged on the mobile body MB can be arbitrarily selected. In one example, there may be multiple imaging devices 2. In this way, one or more target images 20 may consist of multiple target images 20 captured by multiple imaging devices 2. Furthermore, when multiple imaging devices 2 are arranged on the mobile body MB, the direction each imaging device 2 faces can be arbitrarily determined. In one example, each imaging device 2 may be attached to the mobile body MB so that they face different directions from each other. That is, in one example, one or more imaging devices 2 may consist of multiple imaging devices 2 facing different directions from each other.
[0043] As shown in Figure 2, in the method for estimating the position of a moving object MB using the position of the observed object 33 as a clue, the more variations there are in the arrangement of the detected observed object 33, the more clues there are, and therefore the more accurately the position of the viewpoint of the imaging device 2 can be determined. In one example of this embodiment, by having multiple imaging devices 2 facing in different directions from each other, the target image 20 can be obtained over a wide field of view (the acquisition range of the observed object 33). By reflecting the acquisition results of the observed object 33 present in each direction of this wide field of view, an improvement in the accuracy of position estimation can be expected.
[0044] In one example, when using multiple imaging devices 2, the number of imaging devices 2 may be three or more, as illustrated in Figures 1 and 2. That is, multiple imaging devices 2 may consist of three or more imaging devices 2. Multiple target images 20 may consist of three or more target images 20 captured by three or more imaging devices 2. According to this example, by using three or more imaging devices 2, the observed object 33 can be captured in three or more directions. This makes it possible to expect appropriate position estimation even in situations where the number of unknowns to be estimated is three or more, such as when the XY coordinates (2D coordinates) and orientation are the targets of estimation.
[0045] Note that the number of imaging devices 2 is not limited to these examples. In one example, the number of imaging devices 2 arranged on the moving body MB may be two or one. Also, facing in different directions means that at least one of the vertical and horizontal directions is not the same. When adopting a configuration in which each imaging device 2 faces in different directions, each imaging device The orientation of the imaging devices 2 may be determined as appropriate depending on the embodiment, provided that they do not coincide with each other. In one example, as illustrated in Figures 1 and 2, each of the multiple imaging devices 2 may be arranged to face directions that are equally spaced in the circumferential direction in the horizontal plane. For example, when three imaging devices 2 are placed on a moving body MB, each imaging device 2 may be arranged to face directions that are 120 degrees apart in the circumferential direction. This effectively ensures a wide field of view, and as a result, accurate position estimation can be expected.
[0046] [Target image] The target image 20 may be any captured image generated by the imaging device 2 and subject to the process of detecting the observed object 33.
[0047] When using multiple imaging devices 2, multiple target images 20 may be obtained by acquiring one or more captured images (target images 20) from each imaging device 2. In one example, the multiple imaging devices 2 may be controlled to capture images synchronously. As a result, the multiple target images 20 may consist of images captured synchronously by each imaging device 2. For example, when using three imaging devices 2, three target images 20 may be acquired by capturing images synchronously by the three imaging devices 2. The three acquired target images 20 may be used to estimate the same target position. By having each imaging device 2 capture images synchronously, it is possible to obtain the field of view at the same time in the direction that each imaging device 2 is facing (i.e., the field of view at the same position). The imaging of the imaging devices 2 may be controlled by any computer. The imaging devices 2 may be controlled by the information processing device 1, or by another computer (controller for the imaging device 2, controller for the mobile device MB, etc.). However, the method of acquiring the target images 20 is not limited to this example. For example, if the moving object MB remains in the same position, or if other images can be used to estimate the same target position, images with different acquisition times may be used as target images 20 to estimate the same target position.
[0048] Furthermore, in both cases where there is one imaging device 2 and where there are multiple imaging devices 2, the direction in which the imaging device 2 faces may be changeable. For example, the imaging device 2 may be configured to change its direction of orientation by mechanical control using a drive device. The type of drive device may be arbitrarily selected. The drive device may be composed of a known device such as an electric pan / tilt head. Also, for example, the imaging device 2 may be held in the hand of a moving person (an example of a moving body MB), and the direction in which the imaging device 2 faces may be changed manually, such as by changing the orientation of the hand holding the imaging device 2. When the direction in which the imaging device 2 faces is changeable, multiple images observing different directions may be obtained from a single imaging device 2 by changing the direction in which it faces and taking images. The multiple acquired images may be used as target images 20 to estimate the same target position. In particular, when the moving body MB is stationary, a wide field of view can be secured by changing the orientation of the imaging device 2 and obtaining multiple target images 20. Furthermore, if the observed object 33 is not captured in the image captured by the imaging device 2 while the mobile body MB is moving, the direction the imaging device 2 is facing may be changed until the observed object 33 is captured. For example, if the drive device is connected to the information processing device 1, the information processing device 1 may change the direction the imaging device 2 is facing using the drive device until the observed object 33 is captured in the target image 20. This ensures the detection of the observed object 33.
[0049] [Information Processing Device] The information processing device 1 may consist of one or more arbitrary computers. For example, the information processing device 1 may be a controller (control device) for the mobile device MB or imaging device 2. The information processing device 1 may be an arbitrary computer separate from the controller for the mobile device MB or imaging device 2. The information processing device 1 may be, for example, a general-purpose server device, a general-purpose PC (Personal Computers, laptops, terminal devices, etc. Terminal devices include smartphones, etc. This may include tablet devices, store terminals, etc.
[0050] The information processing device 1 may be connected directly or indirectly to the imaging device 2. Indirect connection may be via another computer (such as a controller). The information processing device 1 and the imaging device 2 may be configured as a single unit or as separate units. When the mobile unit MB is a device, the information processing device 1, the imaging device 2, and the mobile unit MB may be configured as a single unit. At least a part of the information processing device 1, the imaging device 2, and the mobile unit MB may be configured separately. For example, the information processing device 1 and the imaging device 2 may be configured as a single unit by a terminal device equipped with a camera. By deploying one or more terminal devices on the mobile unit MB, an imaging system consisting of the information processing device 1, one or more imaging devices 2, and the mobile unit MB may be obtained.
[0051] [Mobile] Figure 3 schematically shows an example of a mobile body MB according to this embodiment. The mobile body MB may be any moving object. The type of mobile body MB may be appropriately selected depending on the embodiment. The object may include at least one of a device (machine) and a living organism.
[0052] (Device) In one example, the mobile body MB may be a moving device MB1. The method of moving the device MB1 is not particularly limited and may be appropriately selected depending on the embodiment. In one example, the device MB1 may be appropriately configured to be movable by operation or autonomously. The operation may be direct or remote manual operation. According to one example of this embodiment, in a scenario where the mobile body MB is a moving device MB1, the position of the mobile body MB can be estimated by a simple method.
[0053] The type of device MB1 can be arbitrarily selected. For example, device MB1 may include at least one of an autonomously moving robotic device MB11 and a device MB12 that moves by direct or remote manual operation. The robotic device MB11 may include, for example, an autonomously driven vehicle, an autonomously flying aircraft, or other robotic devices configured to move autonomously. The aircraft may include a drone.
[0054] Direct manual operation may include, for example, controlling the movement of the device MB12 using an operating unit provided on the device MB12, or physically interacting with the device MB12. Physical interaction may include, for example, pushing or towing. Remote manual operation may include, for example, controlling the movement of the device MB12 remotely using a controller or computer. The device MB12 may include, for example, a manually operated vehicle, a remotely operated aircraft, a cart that is pushed or operated, or other devices configured to be movable by manual operation.
[0055] In one example, device MB1 may be device MB13 used within the store. Environment 4 may be the space within the store. Device MB13 may include at least one of robotic devices MB11 and MB12 used within the store. Robotic devices MB11 used within the store may include, for example, a cleaning robot, a serving robot, a transport robot, a guidance robot, a security robot, etc. Device MB12 used within the store may include, for example, a shopping cart. The shopping cart may include a smart cart, a checkout cart, etc. According to one example of this embodiment, when the mobile body MB (device MB13) moves within the store, the position of the mobile body MB can be estimated by a simple method.
[0056] (biological) In one example, the mobile body MB may be a living organism MB2. The type of living organism MB2 is not particularly limited and may be appropriately selected depending on the embodiment. Living organism MB2 may include at least one of a human MB21 and another living organism MB22 other than a human MB21. The other living organism MB22 may include pets such as dogs and cats. According to one example of this embodiment, the mobile body MB is a living organism In a scenario involving MB2, the position of the moving object MB can be estimated using a simple method.
[0057] In one example, organism MB2 may be organism MB23 moving within the store. Environment 4 may be the space within the store. Organism MB23 may include at least one of human MB21 and other organism MB22 moving within the store. Human MB21 moving within the store may be store clerks, patrol officers, customers, etc. Patrol officers may include, for example, security guards, police officers, etc. Store clerks may include not only persons employed by the store but also any persons working within the store (cleaners, etc.). According to one example of this embodiment, when a moving body MB (organism MB23) moves within the store, the position of the moving body MB can be estimated by a simple method.
[0058] [Attached to a mobile device] One or more imaging devices 2 may be attached to the mobile body MB in any way. The method of attaching the imaging devices 2 to the mobile body MB is not particularly limited, as long as it is possible to image the field of view from the mobile body MB, and may be appropriately selected depending on the embodiment. For example, if the mobile body MB is device MB1, at least one of the one or more imaging devices 2 may be mounted on the mobile body MB (i.e., equipped on the mobile body MB). Also, whether the mobile body MB is device MB1 or biological MB2, at least one of the one or more imaging devices 2 may be attached separately to the mobile body MB.
[0059] The method for separately attaching the imaging device 2 to the mobile body MB may be determined as appropriate depending on the embodiment. For example, if the mobile body MB is device MB1, separately attaching it to the mobile body MB may include attaching it to the mobile body MB using a separate control system from device MB1 which constitutes the mobile body MB. Attaching it to the mobile body MB using a separate control system may include external attachment by any method, such as physically fixing it to the housing of the mobile body MB. Known methods such as adhesive, screw fastening, fitting, and the use of fasteners may be used for attachment. Fasteners may include clamps, pegs, clips, etc. If there is a power supply (battery, etc.), the mobile body MB (device MB1) may be equipped with an external interface capable of external power supply, such as USB (Universal Serial Bus). In this case, attaching it to the mobile body MB using a separate control system may include being connected to the external interface of the mobile body MB and operating while receiving power from the power supply of the mobile body MB.
[0060] Furthermore, for example, if the mobile body MB is a biological organism MB2, separate attachment to the mobile body MB may include being grasped by the mobile body MB, or being attached to the mobile body MB's clothing or equipment. Clothing may include anything worn on the body. Clothing may include, for example, clothes, accessories, etc. Clothes may include pet clothing. Accessories may include hats (including helmets), belts, bracelets, etc. Accessories may also include pet collars, harnesses, leashes, etc. Equipment may be any object that can be carried on the mobile body MB by any means such as grasping or wearing. Equipment may include, for example, bags, baskets, etc. When the imaging device 2 is operated in a store, the equipment on the mobile body MB may include, for example, any equipment that can be used in the store, such as shopping baskets, shopping bag grips, etc. The method of attaching the imaging device 2 to the clothing or equipment is not particularly limited and may be determined as appropriate depending on the embodiment. Known methods may be used for attachment. For example, attaching it separately to the mobile body MB may include attaching it to a pocket on the mobile body MB's clothing.
[0061] [Analysis of target images] One or more target images 20 may be analyzed in any way that can detect the observed object 33. The method for analyzing the target images 20 may be selected from known methods such as general image analysis methods (edge extraction, pattern matching, etc.) or methods using a pre-trained machine learning model. The pre-trained machine learning model may include large-scale generative models such as large-scale visual language models (VLMs) and multimodal models.
[0062] Figure 4 schematically shows an example of a method for detecting an observed object 33 according to this embodiment. In one example, detecting an observed object 33 may be configured by providing a large-scale generative model M1 with a prompt P1 that includes one or more acquired target images 20 and an instruction I11 that commands the large-scale generative model M1 to detect an observed object 33 from the one or more target images 20, thereby causing the large-scale generative model M1 to detect the observed object 33. The large-scale generative model M1 may consist of a model capable of performing image analysis, such as a large-scale visual language model (VLM) or a multimodal model. Known models may be used for the large-scale generative model M1. Instruction I11 may be configured to instruct the detection of an observed object 33. The configuration of instruction I11 is not particularly limited as long as it is possible to command the detection of an object that could become an observed object 33, and may be appropriately determined according to the embodiment. In one example, instruction I11 may consist of natural language text such as "Please detect an object". Instruction I11 may consist of command data other than text (tokens, etc.) that can be interpreted by the large-scale generative model M1. Instruction I11 may include a list of candidate objects to be extracted. According to one example of this embodiment, by using the large-scale generative model M1, the effort of preparing a dedicated program for detecting the observed object 33 can be eliminated. This is expected to reduce the effort required for implementation. Note that the method for detecting the observed object 33 is not limited to this example. In one example, the observed object 33 may be detected by a general image analysis method. Alternatively, the observed object 33 may be detected by a trained machine learning model other than a large-scale generative model.
[0063] (Implementing body) Image analysis of one or more target images 20 may be performed by any computer. For example, image analysis of one or more target images 20 may be performed on the information processing device 1, or on a computer other than the information processing device 1. Analyzing one or more target images 20 may include directly analyzing one or more target images 20 by the information processing device 1, and having another computer analyze one or more target images 20. Detecting the observed object 33 may include detecting the observed object 33 by performing image analysis by the information processing device 1, and directly or indirectly obtaining the detection result 35 by the other computer. Indirect acquisition may involve obtaining it via another computer, etc. The other computer may obtain the one or more target images 20 to be analyzed from the information processing device 1, or from a route other than the information processing device 1. The information processing device 1 may appropriately acquire one or more target images 20 captured by each of the one or more imaging devices 2. When analysis is performed by another computer, acquiring one or more target images 20 may include having another computer acquire one or more target images 20 in order to perform image analysis.
[0064] For example, the large-scale generative model M1 may be deployed on the information processing device 1. The information processing device 1 may obtain the detection result 35 of the observed object 33 by giving the large-scale generative model M1 a prompt P1 and executing the calculation process of the large-scale generative model M1. Alternatively, for example, the large-scale generative model M1 may be deployed on another computer. The other computer may obtain the detection result 35 of the observed object 33 by giving the large-scale generative model M1 a prompt P1 and executing the calculation process of the large-scale generative model M1. The other computer may execute the calculation process of the large-scale generative model M1 in response to a command from the information processing device 1, or it may execute the calculation process of the large-scale generative model M1 autonomously. The information processing device 1 may obtain the detection result 35 of the observed object 33 directly or indirectly from the other computer.
[0065] (Data format) The data format of the detection result 35 may be appropriately selected depending on the embodiment. For example, Detecting the measurement object 33 may include identifying (classifying) the observation object 33. The detection result 35 may include the identification result of the observation object 33. The identification result may be configured to show identifiers, names, types, etc., in text format. For example, the detection result 35 may include a list of the identification results of the observation object 33. As mentioned above, in one example, the level of identification may match the level of identification information included in the placement information 40. The level of identification does not have to match the level of identification information included in the placement information 40, as long as the correspondence between the detection result 35 and the identification information can be identified. In addition, the observation object 33 may be identified in multiple stages (multiple levels), such as two levels of name and type. As a result, the detection result 35 (identification result) may be obtained in multiple stages. In one example, the detection result 35 may be configured to show the detection result of the observation object 33 in image format, such as a bounding box.
[0066] [Search for viewpoint locations that match the detection results] In one example, the estimated position may be two-dimensional or three-dimensional. Furthermore, if the orientation of the imaging device 2 relative to the moving object MB can be determined, for example, by fixing the imaging device 2 to the moving object MB or by measuring the orientation of the imaging device 2 with an arbitrary sensor, then the direction of the observed object 33 captured by the target image 20 can be determined. When the direction of the observed object 33 is determined, the orientation of the moving object MB can be estimated along with the position of the moving object MB by fitting the orientation of the moving object MB to the determined direction of the observed object 33, as illustrated in Figure 2. Therefore, in one example, the orientation of the moving object MB may be estimated together with its position. Estimating the position may also mean estimating the attitude (position and orientation).
[0067] Estimating the position based on the placement information 40 may involve directly or indirectly referencing the placement information 40. Indirect referencing may involve not directly using the placement information 40, but instead using information derived from the placement information 40 (such as the simulation results 50 described later).
[0068] Furthermore, the calculation process for searching for a viewpoint position that matches the detection result 35 may be performed by any computer. For example, the calculation process for searching may be performed on the information processing device 1, or on a computer other than the information processing device 1. Searching for a viewpoint position that matches the detection result 35 may include searching by the information processing device 1 and having the other computer perform the search. Estimating the position may include estimating the position by the information processing device 1 and obtaining the position estimation result 55 directly or indirectly from the other computer. For example, the information processing device 1 may obtain the position estimation result 55 from the other computer by having the other computer (external computer) perform the analysis of one or more target images 20 and the search for a viewpoint position that matches the detection result 35.
[0069] The viewpoint position searched for as matching the detection result 35 may be the position estimation result 55 (estimated position of the mobile body MB). The viewpoint may be the imaging position of the imaging device 2. The relationship between the position of the mobile body MB and the imaging position of the imaging device 2 may be arbitrarily defined. In one example, the position of the mobile body MB and the imaging position of the imaging device 2 may be considered to be the same. The relationship between the position of the mobile body MB and the imaging position of the imaging device 2 may be determined from information such as the mounting position of the imaging device 2 relative to the mobile body MB, and the determined relationship may be reflected in the position estimation. Furthermore, as long as the position can be estimated in a manner such as that illustrated in Figure 2, the method for searching for the viewpoint position matching the detection result 35 is not particularly limited and may be determined appropriately depending on the embodiment.
[0070] (Search method) Figure 5 schematically shows an example of a method for searching for a viewpoint position that matches the detection result 35 of the observed object 33 according to this embodiment. In one example, the information processing device 1 uses the placement information 40 to search for a viewpoint position that matches the detection result 35 of the observed object 33. Subsequently, simulation results 50 may be obtained by simulating imaging by one or more imaging devices 2 from each of the multiple candidate positions V1 in the environment 4. The simulation results 50 may be configured to show the objects 31 that are imaged from each of the multiple candidate positions V1 among the objects 31 present in the environment 4. The information processing device 1 may compare the detection result 35 of the observed object 33 with the obtained simulation results 50. Depending on the result of the comparison, the information processing device 1 may extract candidate positions V5 from the multiple candidate positions V1 where an object 31 matching the detected observed object 33 is imaged, as the viewpoint position that fits the detection result 35 of the observed object 33.
[0071] In other words, for example, searching for the viewpoint position in environment 4 may consist of obtaining simulation results 50 of imaging at each candidate position V1, comparing the detection results 35 of the observed object 33 with the simulation results 50, and extracting a suitable candidate position V5 from among the multiple candidate positions V1. Using these simulation results 50 is an example of indirectly referring to the placement information 40. According to this example, by limiting the search range for the position to multiple candidate positions V1, the convergence of the position estimation process can be ensured. Note that the series of information processing from obtaining the simulation results 50 to extracting the suitable candidate position V5 may be performed by a computer other than the information processing device 1. The information processing device 1 may obtain the extraction results of the candidate positions V5 from the other computer.
[0072] In one example, multiple candidate positions V1 may be specified in a predetermined manner. Candidate positions V1 may be specified by known methods for estimating positions in space. For example, candidate positions V1 may be specified by random means, by dividing the entire space of environment 4 into n parts to specify n positions (where n is a natural number), or by using a dense-sparse search method. If the position estimation of the mobile body MB is being performed continuously, candidate positions V1 may be specified within a range near past position estimation results (estimated positions), such as the previous estimation result. Also, when estimating the attitude of the mobile body MB, the orientation may be specified along with the candidate positions V1. That is, candidate positions V1 may also be candidate attitudes. The number of candidate positions V1 may be determined arbitrarily.
[0073] Simulating imaging by the imaging device 2 from candidate position V1 may involve identifying an object 31 located in the imaging direction of the imaging device 2 from candidate position V1, based on the arrangement information 40. The simulation method is not particularly limited and may be determined appropriately depending on the embodiment, as long as the object 31 to be imaged from candidate position V1 can be identified. If the arrangement information 40 includes arbitrary information about objects other than the object 31 (such as obstacles), the existence of these objects may be reflected in the simulation result 50. At least some of the multiple simulation results 50 may be generated each time or generated in advance. The timing of generating the simulation result 50 (executing the simulation) is not particularly limited and may be determined appropriately depending on the embodiment.
[0074] The simulation result 50 may be generated by the information processing device 1, or by another computer other than the information processing device 1. Multiple generated simulation results 50 may be stored in any storage area. Any storage area may consist of, for example, the memory resources of the information processing device 1, the memory resources of another computer, or an external storage device (including a storage medium). The external storage device may include a data server such as a NAS (Network Attached Storage). In one example, when comparing with the detection result 35, the target simulation result 50 may be appropriately obtained from any storage area. Alternatively, the simulation result 50 may be obtained by executing a simulation (the process of generating the simulation result 50) at any timing before the comparison.
[0075] For example, if the simulation result 50 is generated in advance, all of the pre-generated results The simulation results 50 may be used to compare with the detection results 35. Alternatively, simulation results 50 in which the candidate position V1 belongs to a specific range may be selectively used to compare with the detection results 35. The specific range may be specified in any way, such as the range near the previously estimated position or the range where the provisionally identified moving object MB exists.
[0076] Extracting a suitable candidate position V5 from among multiple candidate positions V1 may involve extracting a simulation result 500 from among multiple simulation results 50 in which the object 31 imaged from candidate position V1 matches the observed object 33 shown in the detection result 35. The candidate position V1 of the extracted simulation result 500 may be the suitable candidate position V5. In the example in Figure 5, a scenario is assumed in which a detection result 35 has been obtained indicating that objects (C3, C4, C10) have been detected as observed objects 33. In this scenario, by comparing the detection result 35 with each simulation result 50, the simulation result 500 for candidate position V1 in which the object 31 imaged from candidate position V1 is the object (C3, C4, C10) can be extracted from among the multiple simulation results 50. As a result, the candidate position V1 of this simulation result 500 can be identified as the suitable candidate position V5.
[0077] The data format of the simulation results 50 may be appropriately selected depending on the embodiment. For example, the simulation results 50 for each candidate position V1 may be configured to show the identification information of the object 31 imaged from each candidate position V1 in text format. When the detection results 35 (identification results) and simulation results 50 are given in text format, the detection results 35 and simulation results 50 may be compared on a text basis. For example, the detection results 35 may include a list of identification information of the detected observed object 33. The simulation results 50 may include a list of identification information of the object 31 imaged from the candidate position V1. Comparing the detection results 35 and simulation results 50 may be done by comparing these lists on a text basis. If the detection results 35 (identification results) are obtained in multiple stages, such as two levels of name and type, the level to be compared may be arbitrarily selected. For example, when the simulation results 50 show the object 31 by name, the level to be compared may be selected according to the simulation results 50, such as selecting the identification result of the name. Furthermore, the simulation results 50 for each candidate position V1 may be configured to show the object 31 captured from each candidate position V1 in image format. When the detection results 35 and simulation results 50 are given in image format, the detection results 35 and simulation results 50 may be matched on an image basis. Image-based matching may be performed by any method. Known methods such as image matching may be used for image-based matching. As long as the candidate position V5 can be identified, other configurations of the method for matching the detection results 35 and simulation results 50 are not particularly limited and may be determined as appropriate depending on the embodiment.
[0078] Figure 6 schematically shows an example of a method for comparing the detection result 35 of the observed object 33 with the simulation result 50 according to this embodiment. In one example, the information processing device 1 may provide a prompt P2 to the large-scale generative model M2, causing the large-scale generative model M2 to extract candidate positions V5. That is, the matching and extraction may be configured by providing a prompt P2 to the large-scale generative model M2, causing the large-scale generative model M2 to perform the matching and extraction.
[0079] Prompt P2 may include the detection result 35 of the observed object 33, the simulation result 50, and instruction I21. Instruction I21 may be configured to command the extraction of a candidate position V5 from among multiple candidate positions V1 in which an object 31 matching the observed object 33 is imaged. Extraction of a suitable candidate position V5 (simulation result 500) The configuration of instruction I21 is not particularly limited and may be determined as appropriate depending on the embodiment, as long as it is possible to issue such instructions. For example, instruction I21 may consist of natural language text such as "Please extract the simulation results that match the detection results from the given simulation results." Instruction I21 may also consist of instruction data other than text (such as tokens) that can be interpreted by the large-scale generative model M2.
[0080] The large-scale generative model M2 may be appropriately selected according to the embodiment, such as the data format of the detection results 35 and simulation results 50 to be given. The large-scale generative model M2 may be composed of models such as a large-scale language model (LLM), a large-scale visual language model (VLM), or a multimodal model. Known models may be used for the large-scale generative model M2. The large-scale generative model M2 may be the same as or different from the large-scale generative model M1.
[0081] The computational processing of the large-scale generative model M2 may be performed on any computer. For example, the large-scale generative model M2 may be deployed on the information processing device 1. The information processing device 1 may obtain the extraction result of suitable candidate positions V5 by giving the large-scale generative model M2 a prompt P2 and executing the computational processing of the large-scale generative model M2. Alternatively, for example, the large-scale generative model M2 may be deployed on another computer. The other computer may obtain the extraction result of suitable candidate positions V5 by giving the large-scale generative model M2 a prompt P2 and executing the computational processing of the large-scale generative model M2. The other computer may execute the computational processing of the large-scale generative model M2 in response to a command from the information processing device 1, or it may execute the computational processing of the large-scale generative model M2 autonomously. The information processing device 1 may obtain the extraction result of suitable candidate positions V5 directly or indirectly from the other computer. In other words, having the large-scale generative model M2 extract candidate positions V5 may include having the information processing device 1 perform calculations on the large-scale generative model M2, and having another computer perform calculations on the large-scale generative model M2, and obtaining the calculation results of the large-scale generative model M2 (the extraction results of candidate positions V5) directly or indirectly from the other computer.
[0082] In one example of this embodiment, by using the large-scale generative model M2, the effort of preparing a dedicated program for matching the detection result 35 and the simulation result 50 can be eliminated. This is expected to reduce the effort required for implementation. In one example, matching between different levels is permitted, such as matching the detection result 35, which shows the observed object 33 by type, with the simulation result 50, which shows the object 31 by name. However, when matching between different levels is permitted, the range of matching expands, as in concept retrieval, but the effort required to prepare reference data (corpus, etc.) to identify the correspondence may increase. In contrast, it is known that large-scale generative models can acquire common sense by learning from a large amount of data. In one example of this embodiment, reference data may be provided, but even without providing reference data, matching can be performed between different levels within the range of common sense acquired by the large-scale generative model M2. From this perspective as well, a reduction in the effort required for implementation can be expected.
[0083] Figure 7 schematically shows another example of a method for comparing the detection result 35 of the observed object 33 with the simulation result 50 according to this embodiment. The comparison method in Figure 7 is a method in which, in addition to comparing and extracting as in the method in Figure 6, the large-scale generative model is made to perform the detection of the observed object 33 as well. In one example, the information processing device 1 may give the large-scale generative model M3 a prompt P3 to perform the detection of the observed object 33 and the extraction of candidate positions V5. That is, the detection of the observed object 33 and the extraction of candidate positions V5 may be performed by giving the large-scale generative model M3 a prompt P3 to cause the large-scale generative model M3 to detect the observed object 33 and extract candidate positions V5.
[0084] The prompt P3 may include one or more acquired target images 20, acquired simulation results 50, a first instruction I31, and a second instruction I32. The first instruction I31 may be configured to command the detection of an observed object 33 from one or more target images 20. The first instruction I31 may be configured similarly to instruction I11. The first instruction I31 may consist of natural language text, or it may consist of non-text instruction data that can be interpreted by the large-scale generative model M3. The detection result 35 of the observed object 33 may be obtained as a result of the inference processing corresponding to instruction I11 being performed in the large-scale generative model M3. The second instruction I32 may be configured similarly to instruction I21, except that the detection result 35 is obtained as a result of the inference processing corresponding to the first instruction I31. The second instruction I32 may be configured as appropriate to perform an operation (inference processing according to the second instruction I32) to extract candidate positions V5 after the detection result 35 is obtained, such as by giving an explicit command described after the first instruction I31. The second instruction I32 may consist of natural language text, or it may consist of instruction data other than text that can be interpreted by the large-scale generative model M3.
[0085] The prompt P3 may be given to the large-scale generative model M3 all at once or in stages. For example, the simulation result 50 and the second instruction I32 may be given to the large-scale generative model M3 after obtaining the detection result 35 of the observed object 33 by providing one or more target images 20 and the first instruction I31 to the large-scale generative model M3 and executing the calculation process of the large-scale generative model M3.
[0086] The large-scale generative model M3 may be appropriately selected according to the embodiment, such as the data format of prompt P3. The large-scale generative model M3 may consist of a model capable of performing image analysis, such as a large-scale visual language model (VLM) or a multimodal model. A known model may be used for the large-scale generative model M3. The large-scale generative model M3 may be identical to or different from at least one of the large-scale generative models (M1, M2).
[0087] The computational processing of the large-scale generative model M3 may be performed on any computer. For example, the large-scale generative model M3 may be deployed on the information processing device 1. The information processing device 1 may obtain the extraction result of suitable candidate positions V5 by giving the large-scale generative model M3 a prompt P3 and executing the computational processing of the large-scale generative model M3. Alternatively, for example, the large-scale generative model M3 may be deployed on another computer. The other computer may obtain the extraction result of suitable candidate positions V5 by giving the large-scale generative model M3 a prompt P3 and executing the computational processing of the large-scale generative model M3. The other computer may execute the computational processing of the large-scale generative model M3 in response to a command from the information processing device 1, or it may execute the computational processing of the large-scale generative model M3 autonomously. The information processing device 1 may obtain the extraction result of suitable candidate positions V5 directly or indirectly from the other computer. In other words, having the large-scale generative model M3 extract candidate positions V5 may include having the information processing device 1 perform calculations on the large-scale generative model M3, and having another computer perform calculations on the large-scale generative model M3, and obtaining the calculation results of the large-scale generative model M3 (the extraction results of candidate positions V5) directly or indirectly from the other computer.
[0088] According to one example of this embodiment, by using the large-scale generative model M3, it is possible to eliminate the need to prepare a dedicated program for detecting the observed object 33 and extracting candidate positions V5. This is expected to reduce the effort required for implementation.
[0089] The matching method is not limited to the examples in Figures 6 and 7, and may be modified as appropriate depending on the embodiment. For example, the matching of the detection result 35 and the simulation result 50 may be performed on a large scale. This may be done using a pre-trained machine learning model other than a model-generating model. The detection results 35 and simulation results 50 may be matched using a rule-based method. The matching rules may be defined as appropriate depending on the embodiment.
[0090] Furthermore, the method for searching for a viewpoint position that matches the detection result 35 is not limited to the example in Figure 4 above, and may be appropriately modified depending on the embodiment. In one example, when multiple observation objects 33 are detected, the viewpoint position (position of the moving body MB) that matches the detection result 35 may be searched (estimated) by mathematical analysis, such as identifying the position of the intersection of lines drawn from the position of each observation object 33, which is specified by the arrangement information 40, in the direction of the imaging device 2.
[0091] [If no observed object is detected] If no observation object 33 is detected at all, the position of the observation object 33 cannot be used as a clue for position estimation. For this reason, in one example, the information processing device 1 may determine that it is difficult to estimate the position in the process of searching for the position of a viewpoint that matches the detection result 35, and may omit the search process. If the points where no observation object 33 is detected are limited to specific points, the information processing device 1 may omit the search process and estimate that the mobile body MB is located at a specific point (i.e., the position of the mobile body MB is a specific point) based on the fact that no observation object 33 has been detected.
[0092] [Combination with other estimation methods] In one example, the information processing device 1 may further estimate the position of the moving object MB using an estimation method other than the above estimation method, which involves searching for a viewpoint position that matches the detection result 35 of the observed object 33. That is, the information processing device 1 may use the above position estimation method, which uses the position of the observed object 33 as a clue, in combination with other estimation methods. As described above, in the position estimation method that uses the position of the observed object 33 as a clue, it may become difficult to estimate the position if the observed object 33 is not detected. In contrast, according to this example, the position estimation can be complemented by using other estimation methods in combination. That is, fail-safe measures can be ensured.
[0093] Other estimation methods are not limited to the above estimation method, which involves the process from detecting the observed object 33 to searching for a viewpoint position that matches the detection result 35 of the observed object 33, and may be appropriately selected depending on the embodiment. Other estimation methods may include known methods, including the method described in Non-Patent Document 1. For the sake of explanation, the above estimation method, which involves searching for a viewpoint position that matches the detection result 35 of the observed object 33, will also be referred to as the "first estimation method," and other estimation methods will also be referred to as the "second estimation method."
[0094] Figure 8 schematically shows an example of a scenario in which the position of the mobile body MB is estimated by another estimation method according to this embodiment. In one example, the information processing device 1 may acquire sensing data 60 from a sensor 6 deployed on the mobile body MB to observe the movement of the mobile body MB as a second estimation method. The information processing device 1 may be directly or indirectly connected to the sensor 6. The information processing device 1 may estimate the position of the mobile body MB by calculating the relative amount of movement of the mobile body MB from the acquired sensing data 60. That is, estimating the position of the mobile body MB by the second estimation method (other estimation method) may consist of acquiring sensing data 60 from the sensor 6 and estimating the position of the mobile body MB by calculating the relative amount of movement of the mobile body MB from the acquired sensing data 60. As a result, the information processing device 1 may acquire the position estimation result 56 by the second estimation method.
[0095] In this position estimation method based on relative movement, the total amount of movement from the initial position can be calculated by continuously calculating and accumulating the calculated relative movement. The position of the moving object MB can then be estimated from the calculated total amount of movement. This is possible. The initial position may be given as appropriate. As long as the position can be estimated by such calculations, the type of sensor 6 is not particularly limited and may be selected as appropriate depending on the embodiment.
[0096] For example, if the mobile body MB is equipped with wheels, sensor 6 may include an odometry sensor (such as a wheel speed sensor). Sensing data 60 may include odometry data (such as the amount of wheel rotation). The information processing device 1 may calculate the relative amount of movement of the mobile body MB from the odometry data using any method. A known method for self-position estimation using odometry may be used for calculating the amount of movement. Sensor 6 may further include other sensors besides the odometry sensor, such as an acceleration sensor. Sensing data 60 may further include other sensing data such as acceleration data. As a result, the information processing device 1 may extend the odometry and calculate the relative amount of movement of the mobile body MB based on the extended odometry.
[0097] In one example, sensor 6 may include a velocity sensor. Sensing data 60 may include velocity data. The information processing device 1 may calculate the relative amount of movement of the moving object MB from the velocity data using any method. A known method may be used to calculate the amount of movement.
[0098] In one example, sensor 6 may include an inertial sensor. The type of inertial sensor may be appropriately selected depending on the embodiment. For example, the inertial sensor may be configured to measure inertial motion (translational motion and rotational motion) with 6 degrees of freedom by including an acceleration sensor and a gyro sensor (angular velocity sensor). The inertial sensor may be configured to measure inertial motion with 9 degrees of freedom by further including a geomagnetic sensor. The inertial sensor may also be called an inertial measurement unit (IMU) or a motion sensor. The degrees of freedom for measurement by the sensor are not limited to these examples and may be set to degrees other than 6 and 9. The sensing data 60 may include measurement data of inertial motion. The information processing device 1 may calculate the relative amount of movement of the moving body MB from the measurement data of inertial motion using any method. Known methods such as inertial navigation and pedestrian dead reckoning (PDR) may be used to calculate the amount of movement.
[0099] A position estimation method based on relative movement can sometimes be executed faster than a first estimation method that includes image analysis (detection of the observed object 33). Therefore, according to one example of this embodiment, the frequency of position estimation can be supplemented by using a position estimation method based on relative movement in combination with the first estimation method. The information processing device 1 may execute the position estimation method based on relative movement more frequently than the first estimation method. In one example, when the position estimation method based on relative movement is executed more frequently than the first estimation method, the information processing device 1 may use the position estimation result 55 (estimated position) obtained by the first estimation method as a reference position. The reference position is an example of an initial position. The information processing device 1 may estimate the position of the moving object MB by accumulating the relative movement amount to the position estimation result 55 (reference position) obtained by the first estimation method at the previous (current) timing until the next timing to execute position estimation by the first estimation method (i.e., it may obtain an estimation result 56). Then, the information processing device 1 may obtain a position estimation result 55 by performing position estimation using the first estimation method at the next timing, and may update the reference position with the obtained position estimation result 55 (estimated position). The information processing device 1 may estimate the position of the moving object MB by repeatedly updating the reference position with the estimation result 55 of the first estimation method and estimating the position from the reference position using the estimation method based on the relative amount of movement, thereby using each estimation method in combination.
[0100] The second estimation method (other estimation method) is not limited to a method of estimating position based on relative movement and may be modified as appropriate depending on the embodiment. For example, the second estimation method may include methods such as satellite positioning, positioning using wireless communication with a wireless access point, and positioning by geometric matching with point cloud data. Sensors deployed on the mobile MB may include, for example, a satellite positioning module, a communication module, LiDAR, radar, etc. The system may include GPS (Global Positioning System) sensors, GNSS (Global Navigation Satellite System) sensors, etc. The information processing device 1 combines three or more estimation methods. The position of the moving object MB may be estimated using this method. In one example, the first estimation method may be used to complement position estimation by a conventional estimation method.
[0101] In one example, the process of estimating the position using the second estimation method (obtaining the position estimation result 56) may be performed on the information processing device 1 as described above, or it may be performed by a computer other than the information processing device 1. Estimating the position of the mobile object MB using the second estimation method may include estimating the position of the mobile object MB by the information processing device 1 and having the other computer estimate the position of the mobile object MB using another estimation method. In the latter case, the information processing device 1 may obtain the position estimation result 56 using the second estimation method directly or indirectly from the other computer.
[0102] Furthermore, in one example, outputting information regarding the position estimation result 55 may include outputting information regarding the position estimation result 56. The frequency at which position estimation is performed using the first estimation method and the second estimation method may be the same or different. If only position estimation using the first estimation method is performed within a nearby time range, the information processing device 1 may output information regarding the estimation result 55. If only position estimation using the second estimation method is performed within a nearby time range, the information processing device 1 may output information regarding the estimation result 56. The nearby time range may be arbitrarily defined.
[0103] When position estimation by the first estimation method and the second estimation method is performed within a nearby time range, the information processing device 1 can obtain the position estimation results (55, 56) from each method. When the position estimation results (55, 56) from the first estimation method and the second estimation method are obtained within a nearby time range, the obtained estimation results (55, 56) may be handled as appropriate depending on the embodiment. In one example, the position estimation results (55, 56) from the first estimation method and the second estimation method may be used as they are. Outputting information about the position estimation result 55 may include outputting information about each of the position estimation results (55, 56). In one example, one of the position estimation results (55, 56) from the first estimation method and the second estimation method may be adopted, and the output of the other may be omitted. For example, the other may be discarded. Outputting information about the position estimation result 55 may include outputting information about one of the position estimation results (55, 56) from the first estimation method and the second estimation method. In addition, in one example, the location estimation results (55, 56) from the first estimation method and the second estimation method may be integrated. The integration of the estimation results (55, 56) may be a simple integration such as a simple average, or a weighted integration such as a weighted average. The weight (priority) of each estimation result (55, 56) may be arbitrarily defined. Outputting information about the location estimation result 55 may include outputting information about the integrated result of the location estimation results (55, 56) from the first estimation method and the second estimation method.
[0104] §2 Example Configuration [Hardware configuration] Figure 9 schematically shows an example of the hardware configuration of the information processing device 1 according to this embodiment. In one example, the information processing device 1 may be configured as a computer in which a control unit 11, a storage unit 12, an external interface 13, an input device 14, and an output device 15 are electrically connected.
[0105] The control unit 11 is configured to perform information processing based on the program and various data. For example, the control unit 11 includes a hardware processor such as a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory). Good. The control unit 11 (CPU) is an example of a processor resource.
[0106] The storage unit 12 is configured to hold arbitrary data. For example, the storage unit 12 may include a hard disk drive, a solid-state drive, semiconductor memory, etc. The storage unit 12, RAM, and ROM are examples of memory resources. In one example of this embodiment, the storage unit 12 may store various information such as program 81. Program 81 is a program that causes the information processing device 1 to execute information processing related to the position estimation of the moving object MB (described later in Figures 11 and 12). Program 81 includes a series of instructions for said information processing.
[0107] In one example, the placement information 40 may be stored in the storage unit 12. The storage unit 12 may store the simulation results 50 for each candidate position V1 instead of or together with the placement information 40. Holding the simulation results 50 may indirectly mean holding the placement information 40. The placement information 40 and the simulation results 50 may be omitted from the storage unit 12. In one example, when a large-scale generation model M1 is deployed in the information processing device 1, the storage unit 12 may store model data M10 representing the large-scale generation model M1. When a large-scale generation model M2 is deployed in the information processing device 1, the storage unit 12 may store model data M20 representing the large-scale generation model M2. When a large-scale generation model M3 is deployed in the information processing device 1, the storage unit 12 may store model data M30 representing the large-scale generation model M3. The configuration of the model data (M10, M20, M30) is not particularly limited and may be determined as appropriate depending on the embodiment, as long as the calculations of the large-scale generative models (M1, M2, M3) can be reproduced. For example, the model data (M10, M20, M30) may include values of calculation parameters adjusted by machine learning, the structure of the model (e.g., the structure of a neural network), etc. The model data (M10, M20, M30) may be incorporated into the program 81. The model data (M10, M20, M30) may be omitted from the storage unit 12.
[0108] In one example, program 81 may be stored in a storage medium 91 instead of, or together with, the storage unit 12. The storage medium 91 is configured to store various types of information (stored programs, etc.) by electrical, magnetic, optical, mechanical, or chemical means so that a machine such as a computer can read the information. The storage unit 12 and the storage medium 91 are examples of non-temporary storage media. The information processing device 1 may retrieve program 81 from the storage medium 91. The storage medium 91 may be a disk-type storage medium (CD, DVD, etc.) or a non-disk-type storage medium such as semiconductor memory (flash memory, etc.). Any drive device may be used to read the information stored in the storage medium 91. The type of drive device may be selected according to the storage medium 91. The drive device may be connected to the information processing device 1 by any method. The storage medium 91 may include an external storage device. In one example, at least one of the arrangement information 40, simulation results 50, and model data (M10, M20, M30) may be stored in the storage medium 91 instead of or together with the storage unit 12.
[0109] The external interface 13 is configured to connect to an external device by wire or wireless connection. The external interface 13 may include, for example, a USB (Universal Serial Bus) port, a dedicated port, a communication port (communication module), etc. The type and number of external interfaces 13 may be determined as appropriate depending on the embodiment. If the external interface 13 includes a communication port (communication module), the communication network standard may be arbitrarily selected. The communication standard may be appropriately selected from, for example, the Internet, wireless communication network, mobile communication network, telephone network, dedicated network, etc. In one example, the information processing device 1 may be connected to at least one of the mobile device MB and the imaging device 2 via the external interface 13. The information processing device 1 may also be connected to the sensor 6 via the external interface 13.
[0110] The input device 14 is configured to accept information input. The input device 14 is composed of, for example, an imaging device, a microphone, a mouse, a keyboard, a touch panel, an operator, etc. This may be done. The imaging device of the input device 14 may be the same as or different from that of the imaging device 2. The output device 15 is configured to output information. The output device 15 may consist of, for example, a display, a speaker, etc. The information processing device 1 may be operated using the input device 14 and the output device 15. The input device 14 and the output device 15 may be directly connected to the information processing device 1, or they may be indirectly connected via the external interface 13. The input device 14 and the output device 15 may be integrated in at least part by a touch panel display, etc.
[0111] Regarding the specific hardware configuration of the information processing device 1, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. Hardware processors include microprocessors, FPGAs (field-programmable gate arrays), DSPs (digital signal processors), and GPs. It may consist of a U (Graphics Processing Unit), an ASIC (application-specific integrated circuit), etc. External interface 13, input device 14 and output device 15 At least one of the following may be omitted. At least one of the program 81, the placement information 40, the simulation results 50, and the model data (M10, M20, M30) may be stored in an external storage device such as a NAS. An external storage device is also an example of a non-temporary storage medium. The information processing device 1 may consist of multiple computers. In this case, the hardware configuration of each computer may or may not be the same. The information processing device 1 may be a computer designed specifically for the services provided, as well as a general-purpose server device, a general-purpose PC, a notebook PC, a terminal device, etc.
[0112] [Software Configuration] Figure 10 schematically shows an example of the software configuration of the information processing device 1 according to this embodiment. The control unit 11 of the information processing device 1 executes instructions contained in the program 81 stored in the storage unit 12 using the CPU. As a result, the information processing device 1 operates as a computer equipped with an image acquisition unit 111, an analysis unit 112, a first estimation unit 113, a second estimation unit 114, and an output processing unit 115 as software modules. In other words, in one example, each software module of the information processing device 1 may be implemented by the control unit 11 (CPU).
[0113] The image acquisition unit 111 is configured to acquire one or more target images 20 captured by one or more imaging devices 2 attached to the mobile body MB. The analysis unit 112 is configured to detect an observation object 33 that appears in at least one of the acquired target images 20 by analyzing the acquired one or more target images 20. The first estimation unit 113 is configured to estimate the position of the mobile body MB by searching for a viewpoint position that matches the detection result 35 of the observation object 33 in the environment 4 in which the mobile body MB is moving, based on the placement information 40. The second estimation unit 114 is configured to estimate the position of the mobile body MB by a second estimation method other than the first estimation method, which involves searching for a viewpoint position that matches the detection result 35 of the observation object 33. The output processing unit 115 is configured to output information regarding the position estimation result 55.
[0114] In this example, each software module of the information processing device 1 is implemented by a general-purpose CPU. However, the method of implementing each of the above modules is not limited to this example and may be modified as appropriate depending on the embodiment. Some or all of the above software modules may be implemented by one or more dedicated processors or chipsets. Each of the above modules may be implemented as a hardware module. Regarding the software configuration of the information processing device 1, modules may be omitted, replaced, and added as appropriate depending on the embodiment.
[0115] §3 Example of Operation Figure 11 is a flowchart showing an example of a processing procedure for estimating the position of a moving object MB by the information processing device 1 according to this embodiment. The information processing device 1 (control unit 11) is configured to execute the following steps in accordance with the instructions included in the program 81. The following processing procedure is an example of an information processing method (control method) executed by a computer. The following processing procedure is merely an example, and each step may be modified as much as possible. Furthermore, steps in the following processing procedure can be omitted, replaced, and added as appropriate, depending on the embodiment.
[0116] (Step S101) In step S101, the control unit 11 operates as an image acquisition unit 111. That is, the control unit 11 acquires one or more target images 20 captured by one or more imaging devices 2 attached to the mobile body MB.
[0117] In one example, multiple imaging devices 2 may be mounted on a mobile body MB so that each faces a different direction from the others. The control unit 11 may acquire multiple target images 20 captured by the multiple imaging devices 2. In another example, three or more imaging devices 2 may be mounted on a mobile body MB so that each faces a different direction from the others. The control unit 11 may acquire three or more target images 20 captured by three or more imaging devices 2.
[0118] In one example, the moving object MB may be a living organism MB2. In another example, the living organism MB2 may be a living organism MB23 that moves around inside the store, and the environment 4 may be the space inside the store. In yet another example, the moving object MB may be a moving device MB1. In another example, the device MB1 may be a device MB13 used inside the store, and the environment 4 may be the space inside the store. Once one or more target images 20 are acquired, the control unit 11 proceeds to the next step S102.
[0119] (Step S102) In step S102, the control unit 11 operates as an analysis unit 112. That is, the control unit 11 analyzes one or more acquired target images 20 to detect an observed object 33 that is captured in at least one of the acquired target images 20. As a result, the control unit 11 obtains a detection result 35 of the observed object 33.
[0120] In one example, the control unit 11 may cause the large-scale generative model M1 to detect the observed object 33 by providing the large-scale generative model M1 with a prompt P1 that includes an instruction I11 that commands the large-scale generative model M1 to detect the observed object 33 from one or more acquired target images 20. As a result, the control unit 11 may obtain the detection result 35 of the observed object 33 as the calculation result of the large-scale generative model M1. Once the detection result 35 of the observed object 33 is obtained, the control unit 11 proceeds to the next step S103.
[0121] (Step S103) In step S103, the control unit 11 determines the branch destination of the process according to the detection result 35 from step S102. If the observation object 33 is detected, the control unit 11 proceeds to step S104. On the other hand, in one example, if the observation object 33 is not detected, the control unit 11 may determine that it is difficult to estimate the position of the moving body MB and omit step 104. In this case, the control unit 11 may terminate the processing procedure related to this example. Also, in one example, if the locations where the observation object 33 is not detected are limited to specific locations, the control unit 11 may estimate that the moving body MB is located at those specific locations. In this case, the control unit 11 may proceed to step S105.
[0122] (Step S104) In step S104, the control unit 11 operates as the first estimation unit 113. That is, Based on the placement information 40, the control unit 11 estimates the position of the mobile body MB by searching for a viewpoint that matches the detection result 35 of the observed object 33 in the environment 4 in which the mobile body MB is moving. As a result, the control unit 11 obtains the position estimation result 55.
[0123] The method for searching for a viewpoint that matches the detection result 35 is not particularly limited, and can be determined appropriately depending on the embodiment, as long as the position can be estimated in the manner illustrated in Figure 2. In one example, the control unit 11 may obtain a simulation result 50 generated by simulating imaging by one or more imaging devices 2 from each of a plurality of candidate positions V1 in the environment 4 based on the arrangement information 40. The simulation result 50 may be configured to show the objects 31 that are imaged from each of the plurality of candidate positions V1 among the objects 31 present in the environment 4. The control unit 11 may compare the detection result 35 of the observed object 33 with the obtained simulation result 50. Depending on the result of the comparison, the control unit 11 may extract from the plurality of candidate positions V1 a candidate position V5 in which an object 31 matching the detected observed object 33 is imaged, as the viewpoint that matches the detection result 35 of the observed object 33 (i.e., the estimated position result 55).
[0124] In one example, the control unit 11 may use a large-scale generative model M2 to compare with the simulation result 50 and extract candidate positions V5. That is, in one example, the control unit 11 may give a prompt P2 to the large-scale generative model M2 to extract candidate positions V5. The prompt P2 may include the detection result 35 of the observed object 33, the simulation result 50, and instruction I21. Instruction I21 may be appropriately configured to command the extraction of candidate positions V5 from among a plurality of candidate positions V1 in which an object 31 matching the observed object 33 is imaged.
[0125] In one example, the control unit 11 may use the large-scale generative model M3 for processing steps S102 and S104. That is, in one example, the control unit 11 may give the large-scale generative model M3 a prompt P3 to cause the large-scale generative model M3 to perform detection of the observed object 33 and extraction of candidate positions V5. The prompt P3 may include one or more acquired target images 20, acquired simulation results 50, a first instruction I31, and a second instruction I32. The first instruction I31 may be configured to command the detection of the observed object 33 from one or more target images 20. The second instruction I32 may be configured to command the extraction of a candidate position V5 from among a plurality of candidate positions V1 in which an object 31 matching the observed object 33 is imaged. Once the position estimation result 55 is obtained, the control unit 11 proceeds to the next step S105.
[0126] (Step S111) In one example, the control unit 11 may operate as a second estimation unit 114 and perform the processing of step S111 at least partially in parallel with the processing of steps S101 to S104. In step S111, the control unit 11 estimates the position of the moving object MB using a second estimation method other than the first estimation method performed in steps S102 and S104.
[0127] Figure 12 is a flowchart showing an example of a position estimation processing procedure (subroutine) in the second estimation method (other estimation method) according to this embodiment. Depending on the embodiment, steps in the processing procedure in Figure 12 can be omitted, replaced, and added as appropriate. The processing in step S111 may include the processing in steps S1111 and S1112 below.
[0128] In step S1111, the control unit 11 acquires sensing data 60 from sensors 6 deployed on the mobile body MB to observe the movement of the mobile body MB. In step S1112, the control unit 11 estimates the position of the mobile body MB by calculating the relative amount of movement of the mobile body MB from the acquired sensing data 60. As a result, the control unit 11 calculates the position estimation result. The control unit 11 obtains 56. Once it obtains the position estimation result 56, it proceeds to the next step 105.
[0129] Note that the second estimation method used in step S111 is not limited to the method shown in Figure 12. For example, the second estimation method may include methods such as satellite positioning, positioning using wireless communication with a wireless access point, or positioning by geometric matching with point cloud data.
[0130] Furthermore, step S111 may be omitted. In this case, the second estimation unit 114 may be omitted from the software configuration of the information processing device 1.
[0131] Furthermore, the timing of executing the process in step S111 is not limited to the example in Figure 11, and may be appropriately changed depending on the embodiment. The process in step S111 may be executed before step S101. The process in step S111 may be executed between any of the steps S101 to S104. The process in step S111 may be executed after the process in step S104.
[0132] Furthermore, the frequency at which the process in step S111 is executed may be the same as, or different from, the frequency at which the series of processes in steps S101 to S104 are executed. For example, the control unit 11 may execute the process in step S111 at a higher frequency than the series of processes in steps S101 to S104.
[0133] (Step S105) Returning to Figure 11, in step S105, the control unit 11 operates as an output processing unit 115. That is, the control unit 11 outputs information regarding the position estimation results (55, 56).
[0134] The content and destination of the information to be output may be appropriately selected depending on the embodiment. In one example, the control unit 11 may output the position estimation results (55, 56) as they are. In another example, the control unit 11 may perform predetermined information processing according to the acquired position estimation results (55, 56). The control unit 11 may output the result of performing that information processing as information related to the estimation results (55, 56). For example, the control unit 11 may accept input of a destination and perform navigation from the position of the estimation result (55, 56) to the destination. Navigation may be performed by a known method. The content of the output navigation is an example of the result of performing predetermined information processing. If the scale of the placement information 40 matches real space, the control unit 11 can perform navigation that conforms to the scale of real space. Even if the scale of the placement information 40 does not match real space, the control unit 11 can perform navigation that does not depend on real space, such as the direction of travel. In one example, the output destination may be RAM, memory unit 12, output device 15, storage medium 91, another computer, external storage device, etc.
[0135] Once the information is output, the control unit 11 terminates the processing procedure related to this example. The control unit 11 may repeatedly execute the processes of steps S101 to S105 and step S111 at any time. The control unit 11 may also repeatedly execute the processes of steps S101 to S105 and step S111 in real time.
[0136] (Features) In this embodiment, by processing steps S102 and S104, observation objects 33 that appear in the target image 20 are detected, and the position of the mobile body MB can be estimated using the position of the detected observation objects 33 indicated by the placement information 40 as a clue. Therefore, if the placement of objects 31 present in the environment 4 is identified, the position of the mobile body MB can be estimated. That is, a detailed map may be used, but position estimation is practical. The only preparation required for implementation is to prepare a simplified map (placement information 40) showing the arrangement of object 31. Therefore, according to this embodiment, the position of the moving body MB can be estimated in a simplified manner.
[0137] Furthermore, in one example, the arrangement of products within a store is frequently changed, and the condition of the shelves changes as products are sold, which can easily cause changes in environment 4, potentially leading to frequent map updates. When using a detailed map, updating this map is also time-consuming. In contrast, in this embodiment, a simplified map (arrangement information 40) is sufficient, so even in situations where environment 4 changes frequently, a reduction in effort can be expected.
[0138] Furthermore, in environments with many objects not registered on the map, such as crowded environments or store environments during stocking, the measured point cloud data may deviate significantly from the map due to the influence of these objects. As a result, conventional methods of geometrically matching point cloud data may become difficult to use for position estimation. In contrast, in this example, if the observed object 33 is visible in the target image 20 to a detectable degree, its position can be estimated in steps S102 and S104. Therefore, according to this example, robust position estimation can be expected.
[0139] Furthermore, in vast environments such as shopping malls and amusement parks, the presence of objects registered on the map within the LiDAR measurement range may be small, making it difficult to collect point cloud data sufficient for geometric matching. Even in this case, conventional methods of geometric matching of point cloud data may become difficult to use for position estimation. In contrast, in this example, if the observed object 33 is captured in the target image 20 to a detectable degree, its position can be estimated by processing in steps S102 and S104. Therefore, according to this example, robust position estimation can be expected from this viewpoint as well. Accordingly, in one example, the first estimation method may be used to complement position estimation by conventional estimation methods. This makes it possible to expect stable position estimation in various environments.
[0140] §4 Variant While embodiments of this disclosure have been described in detail above, the above description is merely illustrative in all respects. The processes and means described in this disclosure can be freely combined and implemented as long as no technical inconsistencies arise. Furthermore, various improvements or modifications may be made to the above embodiments as appropriate.
[0141] §5 Experimental Examples To verify the effectiveness of this disclosure, the following experiments were conducted. However, this disclosure is not limited to the following experimental examples.
[0142] (Comparative example) In the comparative example, a commercially available robotic device (KachakaPro) equipped with 2D LiDAR was moved. It was adopted as the main body. While taking measurements with the mounted 2D LiDAR, the robotic device was small The robot was operated inside the store. By comparing the point cloud map of the retail store, which was generated in advance, with the measurement results (point cloud data) of the 2D LiDAR using the method described in Non-Patent Document 1, the position of the robot device was determined. The orientation (including direction) was estimated.
[0143] (Example of experiment) In the first experimental example, the same robotic apparatus as in the comparative example was used as the mobile body. The robotic apparatus was equipped with three commercially available imaging devices (RealSense Depth Camera D455). As in Figure 1, the three imaging devices were arranged so that they faced in directions at 120-degree intervals in the circumferential direction. The method shown in Figure 4 was used to detect the observed object from the target image obtained from each imaging device. It was adopted. For the large-scale generative model, GPT-4.1 was adopted. The method shown in Figure 5 was used to search for a suitable viewpoint.
[0144] Figure 13 shows a simulation example of imaging from candidate positions in viewpoint search in the first experimental example. In viewpoint search using the method in Figure 5, 1 million candidate positions (including orientation) were randomly selected. As illustrated in Figure 13, a retail store floor map (a diagram of the arrangement of objects in the store) was used as layout information, and the simulation results (a list of the names of objects imaged from the candidate position) for each specified candidate position were obtained in text format. The obtained simulation results and the detection results of observed objects were compared in a text-based manner. Based on the comparison results, candidate positions in which objects matching the detected observed objects were imaged were extracted from the 1 million candidate positions, and the extracted candidate positions were adopted as the estimated position of the robot device.
[0145] In the second experimental example, the estimated position of the robot device was obtained by integrating the estimation results from the comparative example and the first experimental example. A simple average was used for the integration method.
[0146] Using the methods of the comparative example, the first experimental example, and the second experimental example, position estimation results were obtained at 67 locations. In addition, true values were obtained using the method described in Non-Patent Document 1, which further utilizes odometry data obtained from the robotic device. The position estimation results and true values were compared using each method, and the average values of translational error (m) and rotational error (rad) were calculated.
[0147] Figure 14 shows the position measurement errors (translational error and rotational error) in the comparative example, the first experimental example, and the second experimental example. As shown in Figure 14, in the first experimental example, both the translational error and rotational error were smaller compared to the comparative example. From this result, it was found that the position of the moving object can be appropriately estimated by using the position of the observed object captured in the image (target image) of the imaging device as a clue. Furthermore, in the second experimental example, both the translational error and rotational error were smaller compared to the first experimental example and the comparative example. From this result, it was found that by using this method in combination with other estimation methods, it is possible to complement position estimation and improve estimation accuracy. [Explanation of Symbols]
[0148] 1...Information processing device, 11...Control unit, 12...Storage unit, 2...imaging device, 20...target image, 31...Object, 33...Observed object, 35...Detection results, 40...Placement information, 55...Estimated results, 4...Environment, MB...Moving object
Claims
1. An information processing device comprising a control unit, The control unit, To acquire one or more target images captured by one or more imaging devices attached to a moving object. By analyzing the acquired one or more target images, the observed object is detected in at least one of the acquired one or more target images. Based on the placement information, the position of the moving body is estimated by searching for a viewpoint that matches the detection result of the observed object in the environment in which the moving body is moving, and Outputting information regarding the estimation result of the aforementioned position, It is configured to perform, The aforementioned placement information is configured to indicate the placement of objects present in the environment, Exploring the location of the viewpoint in the aforementioned environment is Based on the aforementioned arrangement information, simulation results are obtained that show the objects present in the environment that are imaged from each of the multiple candidate locations by simulating imaging by one or more imaging devices from each of the multiple candidate locations in the environment. Comparing the detection results of the observed object with the acquired simulation results, Based on the matching results, candidate positions in which an object matching the detected observed object is imaged are extracted from among the multiple candidate positions as viewpoint positions that match the detection results of the observed object. Composed of, Information processing device.
2. The detection of the observed object and the extraction of the candidate positions are performed by: the acquisition of one or more target images, the acquisition of the simulation results, a first instruction commanding the detection of the observed object from the one or more target images, and the extraction of the candidate positions from the plurality of candidate positions in which an object matching the observed object is captured. This is achieved by providing a large-scale generative model with a prompt that includes a second instruction to detect the observed object and extract the candidate position, thereby causing the large-scale generative model to detect the observed object and extract the candidate position. The information processing apparatus according to claim 1.
3. The matching and extraction are performed by giving the large-scale generative model a prompt that includes an instruction to extract from the plurality of candidate positions the candidate position in which an object matching the detection result of the observed object, the simulation result, and the observed object is imaged, thereby causing the large-scale generative model to perform the matching and extraction. The information processing apparatus according to claim 1.
4. An information processing device comprising a control unit, The control unit, To acquire one or more target images captured by one or more imaging devices attached to a moving object. By analyzing the acquired one or more target images, the observed object is detected in at least one of the acquired one or more target images. Based on the placement information, the position of the moving body is estimated by searching for a viewpoint that matches the detection result of the observed object in the environment in which the moving body is moving, and Outputting information regarding the estimation result of the aforementioned position, It is configured to perform, The aforementioned placement information is configured to indicate the placement of objects present in the environment using labels. The detection result of the observed object includes the identification result of the observed object in text format. Searching for a viewpoint that matches the detection result of the observed object based on the aforementioned placement information is performed by searching for a viewpoint that matches the detection result of the observed object based on the result of matching the label of the placement information with the identification result of the observed object in a text-based manner. Information processing device.
5. The one or more imaging devices described above are composed of multiple imaging devices facing in different directions from each other. The one or more target images described above are composed of multiple target images captured by the multiple imaging devices, The information processing apparatus according to any one of claims 1 to 4.
6. The aforementioned plurality of imaging devices are composed of three or more imaging devices. The information processing apparatus according to claim 5.
7. The detection of the observed object is performed by providing a large-scale generative model with a prompt that includes one or more acquired target images and an instruction to detect the observed object from the one or more target images, thereby causing the large-scale generative model to detect the observed object. The information processing apparatus according to claim 1 or 4.
8. The control unit is To acquire sensing data from sensors deployed on the moving object to observe its movement, and The position of the moving object is estimated by calculating the relative amount of movement of the moving object from the acquired sensing data. It is configured to perform further actions. The information processing apparatus according to any one of claims 1 to 4.
9. The aforementioned moving body is a living organism. The information processing apparatus according to any one of claims 1 to 4.
10. The aforementioned organism is an organism that moves around inside the store. The aforementioned environment is the space inside the store. The information processing apparatus according to claim 9.
11. The aforementioned moving body is a moving device. The information processing apparatus according to any one of claims 1 to 4.
12. The aforementioned device is a device used within a store, The aforementioned environment is the space inside the store. The information processing apparatus according to claim 11.
13. A method of information processing performed by a computer, To acquire one or more target images captured by one or more imaging devices attached to a moving object. By analyzing the acquired one or more target images, the observed object is detected in at least one of the acquired one or more target images. Based on the placement information, the position of the moving body is estimated by searching for a viewpoint that matches the detection result of the observed object in the environment in which the moving body is moving, and Outputting information regarding the estimation result of the aforementioned position, Includes, The aforementioned placement information is configured to indicate the placement of objects present in the environment, Exploring the location of the viewpoint in the aforementioned environment is Based on the aforementioned arrangement information, simulation results are obtained that show the objects present in the environment that are imaged from each of the multiple candidate locations by simulating imaging by one or more imaging devices from each of the multiple candidate locations in the environment. Comparing the detection results of the observed object with the acquired simulation results, Based on the matching results, candidate positions in which an object matching the detected observed object is imaged are extracted from among the multiple candidate positions as viewpoint positions that match the detection results of the observed object. Composed of, Information processing methods.
14. An information processing method performed by a computer, To acquire one or more target images captured by one or more imaging devices attached to a moving object. By analyzing the acquired one or more target images, the observed object is detected in at least one of the acquired one or more target images. Based on the placement information, the position of the moving body is estimated by searching for a viewpoint that matches the detection result of the observed object in the environment in which the moving body is moving, and Outputting information regarding the estimation result of the aforementioned position, Includes, The aforementioned placement information is configured to indicate the placement of objects present in the environment using labels. The detection result of the observed object includes the identification result of the observed object in text format. Searching for a viewpoint that matches the detection result of the observed object based on the aforementioned placement information is performed by searching for a viewpoint that matches the detection result of the observed object based on the result of matching the label of the placement information with the identification result of the observed object in a text-based manner. Information processing methods.
15. A program that causes a computer to execute an information processing method, The aforementioned information processing method is To acquire one or more target images captured by one or more imaging devices attached to a moving object. By analyzing the acquired one or more target images, the observed object is detected in at least one of the acquired one or more target images. Based on the placement information, the position of the moving body is estimated by searching for a viewpoint that matches the detection result of the observed object in the environment in which the moving body is moving, and Outputting information regarding the estimation result of the aforementioned position, Includes, The aforementioned placement information is configured to indicate the placement of objects present in the environment, Exploring the location of the viewpoint in the aforementioned environment is Based on the aforementioned arrangement information, simulation results are obtained that show the objects present in the environment that are imaged from each of the multiple candidate locations by simulating imaging by one or more imaging devices from each of the multiple candidate locations in the environment. Comparing the detection results of the observed object with the acquired simulation results, Based on the matching results, candidate positions in which an object matching the detected observed object is imaged are extracted from among the multiple candidate positions as viewpoint positions that match the detection results of the observed object. Composed of, program.
16. A program for causing a computer to execute an information processing method, The aforementioned information processing method is To acquire one or more target images captured by one or more imaging devices attached to a moving object. By analyzing the acquired one or more target images, the observed object is detected in at least one of the acquired one or more target images. Based on the placement information, the position of the moving body is estimated by searching for a viewpoint that matches the detection result of the observed object in the environment in which the moving body is moving, and Outputting information regarding the estimation result of the aforementioned position, Includes, The aforementioned placement information is configured to indicate the placement of objects present in the environment using labels. The detection result of the observed object is the identification result of the observed object in text format. Includes fruit, Searching for a viewpoint that matches the detection result of the observed object based on the aforementioned placement information is performed by searching for a viewpoint that matches the detection result of the observed object based on the result of matching the label of the placement information with the identification result of the observed object in a text-based manner. program.