Information processing method and information processing system
Patent Information
- Application Number
- PCT/JP2026/003611
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-02-02
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026003611_01102026_PF_FP_ABST
Abstract
Description
Information Processing Method and Information Processing System
[0001] The present disclosure relates to an information processing method and an information processing system.
[0002] In recent years, a technique for estimating self-position by referring to a map of the surrounding environment (hereinafter, also simply referred to as "map") is known. As an example of such a technique, a technique called SLAM (Simultaneous Localization and Mapping), which performs self-position estimation while creating a map, is also known. However, there may be a deviation between the scale of the created map and the scale of the real world. Therefore, it is required to correct the scale of the map. Patent Document 1 discloses a technique for correcting the scale of a map.
[0003] Also, as a technique for correcting the scale of a map, a technique is known in which a user measures the distance between a plurality of positions in the real world and inputs the measured distance. According to such a technique, the scale of the map can be corrected based on the distance between the plurality of positions input by the user and the change in self-position between the plurality of positions.
[0004] International Publication No. 2022 / 134475
[0005] However, the work of measuring the distance and the work of inputting the distance may impose a burden on the user. Furthermore, when a user measures the distance between a plurality of positions existing in a large-scale environment, it may be difficult for the user to measure the distance between the plurality of positions while moving between the plurality of positions.
[0006] Therefore, it is desired to provide a technique capable of correcting the scale of a map while reducing the burden on the user.
[0007] According to the present disclosure, there is provided an information processing method executed by a processor, comprising: acquiring a map created based on image data; recognizing a recognition target based on image data obtained by a camera; and outputting a map with a corrected scale based on first distance data relating to the recognition target and the map.
[0008] Furthermore, according to this disclosure, an information processing system is provided which includes a processor that acquires a map created based on image data, recognizes a recognition target based on image data obtained by a camera, and outputs a scaled map based on first distance data relating to the recognition target and the map.
[0009] This figure illustrates an example of the effects of map scale misalignment. This figure shows an example of the configuration of the information processing system 1 according to the embodiment of this disclosure. This figure explains the function of the self-position estimation unit 110. This figure illustrates the case where the recognition target includes an object. This figure illustrates the case where the recognition target includes marker A2. This figure illustrates the case where the recognition target includes a normalized area. This figure shows an example of a combination of image data obtained by camera C1 and the position of camera C1. This figure shows an example of a method for obtaining measured distance data. This figure illustrates an example of scale correction of map M1. This is a flowchart showing the processing flow of the information processing system 1 according to the embodiment of this disclosure. This is a hardware configuration diagram showing an example of a computer 1000 that realizes the functions of the information processing system 1.
[0010] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0011] The explanation will be given in the following order: 0. Overview of the Embodiment 1. Details of the Embodiment 1.1. Configuration of the Information Processing System 1.2. Processing Flow of the Information Processing System 1.3. Effects of the Embodiment 2. Hardware Configuration Example 3. Modification
[0012] <0. Overview of Embodiments> First, an overview of the embodiments of this disclosure will be described.
[0013] In recent years, a technique called SLAM (Simulation-Landing Map Modeling) has become known, which performs self-localization while creating a map. This technique allows for the estimation of one's own position within a map. However, there is a possibility of a discrepancy between the scale of the created map and the scale of the real world (hereinafter simply referred to as "map scale discrepancy"). An example of the effects of map scale discrepancy will be explained with reference to Figure 1.
[0014] Figure 1 illustrates an example of the effect of map scaling. Referring to Figure 1, state T8 is shown as having no scaling. Let's assume that in state T8, with no scaling, a camera in the real world moves 1m. In this case, since there is no scaling between positions F81 and F84, it can be correctly estimated that the camera moved 1m from position F81 to position F82.
[0015] Referring to Figure 1, the state T9 with a scale shift is also shown. Even in the state T9 with a scale shift, we assume that the camera in the real world has moved 1m. In this case, because there is a scale shift between positions F91 and F94, it may be incorrectly estimated that the camera has only moved 0.5m from position F91 to position F92.
[0016] When there is a scale discrepancy in the map like this, it is possible that the estimated camera position from position F91 to position F92 may be inaccurate.
[0017] For example, ICVFX (In_Camera visual effects) is a known technique for rendering the background according to the camera's position and orientation. In such a technique, if there is a scale misalignment, the background may be rendered in a position deviating from its normal position, potentially resulting in a discrepancy in the rendered background's location. Alternatively, Augmented Reality (AR) technology is known for placing superimposed objects according to the camera's position and orientation. In such a technique, if there is a scale misalignment, the superimposed objects may be placed in a position deviating from their normal position, potentially resulting in a discrepancy in the placement of the superimposed objects.
[0018] Therefore, it is necessary to adjust the map scale. In particular, this specification mainly describes techniques that can adjust the map scale while reducing the burden on the user.
[0019] Furthermore, a technique could be conceivable in which markers of known size are placed in the real world, a map is created based on image data obtained by a camera, the size of the markers appearing in the image data is measured, and the map scale is corrected based on the measured size of the markers and the known size of the markers.
[0020] However, in techniques that correct the map scale while creating a map, it is conceivable that it may be difficult for the camera to simultaneously capture images for map creation and images of markers. Furthermore, in techniques that correct the map scale while creating a map, if the measurement of the marker size fails, it becomes necessary to re-measure the marker size while creating the map again, which increases the processing cost of map creation.
[0021] On the other hand, the technology according to the embodiments of this disclosure is a technology for correcting the scale of a map based on a completed map.
[0022] This technology eliminates the need to simultaneously photograph the map and the markers. Furthermore, even if the measurement of the marker size fails, this technology eliminates the need to re-measure the marker size while recreating the map, thus reducing the processing cost of map creation. Moreover, this technology eliminates the need to capture markers in the image data during map creation, making it possible to exclude unnecessary markers from the map.
[0023] The embodiments of this disclosure have been described above.
[0024] <1. Details of the Embodiments> Next, the details of the embodiments of the present disclosure will be described.
[0025] [1.1. Configuration of the Information Processing System] The configuration of the information processing system 1 according to the embodiment of this disclosure will be described with reference to Figures 2 to 9.
[0026] Figure 2 is a diagram showing an example configuration of an information processing system 1 according to an embodiment of the present disclosure. As shown in Figure 2, the information processing system 1 according to an embodiment of the present disclosure may include a camera C1, a self-position estimation unit 110, a map storage unit 120, a distance measurement determination unit 130, a prior knowledge storage unit 140, a distance measurement unit 150, a scale estimation unit 160, and a scale correction unit 170. The information processing system 1 may be implemented by a computer. The information processing system 1 may be implemented by multiple devices or by a single device.
[0027] Here, the map storage unit 120 and the prior knowledge storage unit 140 may be included in a storage unit (not shown). The storage unit (not shown) is a recording medium that includes memory and stores programs executed by a control unit (not shown), and stores data necessary for the execution of these programs. The storage unit (not shown) also temporarily stores data for calculations performed by the control unit (not shown). The storage unit (not shown) is composed of a magnetic storage device, a semiconductor storage device, an optical storage device, or a magneto-optical storage device, etc.
[0028] Furthermore, the self-position estimation unit 110, the distance measurement and determination unit 130, the distance measurement unit 150, the scale estimation unit 160, and the scale correction unit 170 may be included in a control unit (not shown). The control unit (not shown) may be composed of one or more processors. If the control unit (not shown) is composed of a processor, such a processor may be composed of electronic circuits. The control unit (not shown) can be realized by the execution of a program by such a processor.
[0029] (Camera C1) Camera C1 is configured to include an image sensor and obtains image data by capturing the real world with the image sensor. Camera C1 continuously acquires image data by continuously capturing the real world in a time series. Camera C1 outputs the acquired image data to the self-position estimation unit 110, the distance measurement determination unit 130, and the distance measurement unit 150, respectively. The type of camera C1 is not limited. For example, camera C1 may be a monocular camera, a compound camera, or a depth camera.
[0030] The image data is an example of sensing data used by the self-position estimation unit 110. That is, camera C1 is an example of a predetermined sensor that obtains sensing data used by the self-position estimation unit 110. Therefore, the sensor that obtains sensing data used by the self-position estimation unit 110 (hereinafter also referred to as "sensor used for self-position estimation") may include sensors other than camera C1.
[0031] For example, the sensors used for self-position estimation may include a camera C1 and an inertial measurement device. The inertial measurement device is also called an IMU (Inertial Measurement Unit). Therefore, in the following explanation, the inertial measurement device will also be referred to as "IMU". Alternatively, the sensors used for self-position estimation may include a LiDAR (Light Detection and Ranging), or a LiDAR and a camera C1.
[0032] (Self-position estimation unit 110) The self-position estimation unit 110 performs self-position estimation while creating a map M1 based on image data obtained by the camera C1. The self-position estimation unit 110 obtains the position and orientation of the camera C1 on the map M1 through self-position estimation. Typically, the position and orientation of the camera C1 may each be represented by three-dimensional information. However, the dimensions of the information representing the position and orientation of the camera C1 are not limited. The function of the self-position estimation unit 110 will be explained with reference to Figure 3.
[0033] Figure 3 illustrates the function of the self-position estimation unit 110. As shown in Figure 3, the self-position estimation unit 110 includes a self-position estimater 112. The self-position estimater 112 takes image data as input and outputs a map M1 corresponding to the image data. The self-position estimater 112 also outputs the position and orientation of camera C1 as the self-position corresponding to the image data. By using this self-position estimater 112, the self-position estimation unit 110 obtains the map M1 and the position and orientation of camera C1.
[0034] Typically, the self-localization estimator 112 may be implemented using a model generated by machine learning. However, the self-localization estimator 112 does not have to be a model generated by machine learning. For example, the self-localization estimator 112 may perform self-localization while creating a map M1 using a rule-based method. The self-localization unit 110 stores the created map M1 in the map storage unit 120.
[0035] Furthermore, when the creation of map M1 is complete, the self-position estimation unit 110 outputs to the distance measurement determination unit 130 that the creation of map M1 is complete, and also outputs the position and orientation of camera C1 in map M1 to the distance measurement unit 150. Note that the determination of whether or not the creation of map M1 is complete may be made in any way.
[0036] For example, the self-position estimation unit 110 may determine whether the creation of map M1 is complete based on whether or not a map creation completion button (not shown) has been pressed by the user. In this case, in order to make it easier for the user to understand the progress of map creation, the map M1 being created may be displayed by a display unit (not shown). After the creation of map M1 is complete, the user may move camera C1 or the object to be recognized so that the object to be recognized is within the shooting range of camera C1.
[0037] Furthermore, since the scale of map M1 can be corrected by the scale correction unit 170, there may be a discrepancy between the scale of map M1 created by the self-position estimation unit 110 and the scale of the real world. However, if the sensor used for self-position estimation includes an IMU or LiDAR, a map M1 corresponding to the scale of the real world will be created, and the position and orientation of camera C1 will be estimated according to the scale of the real world.
[0038] (Map storage unit 120) The map storage unit 120 stores the map M1 created by the self-position estimation unit 110. The map storage unit 120 also outputs the map M1 to the scale correction unit 170 according to the control of the scale correction unit 170.
[0039] (Prior Knowledge Storage Unit 140) The prior knowledge storage unit 140 stores a pre-prepared machine learning model or pre-accumulated data as prior knowledge. As will be explained later, the distance measurement and determination unit 130 recognizes the object to be recognized. From such machine learning models or accumulated data, data indicating the distance between multiple locations corresponding to the object to be recognized is acquired as second distance data (hereinafter also referred to as "known distance data") in accordance with the control of the distance measurement and determination unit 130, and output to the distance measurement and determination unit 130.
[0040] Furthermore, the type of machine learning model is not limited. For example, the machine learning model may be a language model. More specifically, the language model may be a Large Language Model (LLM) constructed using a large amount of data and deep learning techniques. With such a large language model, the likelihood of known distance data being responded to according to the control of the distance measurement and determination unit 130 increases.
[0041] The accumulated data may be a three-dimensional model of a recognition target. For example, the three-dimensional model may be generated by photogrammetry that forms a three-dimensional model of a recognition target based on image data of the recognition target captured from various angles. Alternatively, the accumulated data may be drawing data in which distances in the real world are associated with a drawing. For example, the drawing data may be a floor plan in which distances between a plurality of positions on a floor (for example, a distance between walls or a height of a ceiling) are associated.
[0042] (Distance measurement determination unit 130) When the creation of the map M1 is completed, the distance measurement determination unit 130 recognizes a recognition target based on image data obtained by the camera C1. For example, the distance measurement determination unit 130 may recognize the recognition target using a model generated by machine learning, or may recognize the recognition target on a rule basis.
[0043] Then, based on recognizing the recognition target, the distance measurement determination unit 130 determines whether the recognition target is available for correcting the scale of the map M1. To this end, the distance measurement determination unit 130 attempts to acquire known distance data indicating distances between a plurality of positions corresponding to the recognition target.
[0044] More specifically, the distance measurement determination unit 130 attempts to acquire the known distance data from a machine learning model stored by the prior knowledge storage unit 140 or from the accumulated data. Then, the distance measurement determination unit 130 may determine whether the recognition target is available for correcting the scale of the map M1 based on whether the acquisition of the known distance data has succeeded.
[0045] When it is determined that the recognition target cannot be used for correcting the scale of the map M1, the recognition target may be re-recognized. On the other hand, when it is determined that the recognition target is available for correcting the scale of the map M1, the scale of the map M1 can be corrected based on measured distance data, the known distance data, and the map M1. Herein, examples of the recognition target and examples of acquiring the known distance data in each example will be described with reference to FIGS. 4 to 6.
[0046] (When the recognition target includes an object) Figure 4 is a diagram for explaining the case where the recognition target includes an object. As shown in FIG. 4, the recognition target that can be recognized based on image data obtained by the camera C1 may include the object A1. For example, the object A1 may be a mark symbolizing a specific location (i.e., a landmark). As shown in FIG. 4, the landmark may be an image shaped to imitate a person. Alternatively, the landmark may be a building (e.g., Tokyo Tower) or the like.
[0047] The distance measurement determination unit 130 attempts to acquire known distance data indicating distances between a plurality of positions on the object A1. In the example shown in FIG. 4, the upper end and the lower end of the object A1 are the plurality of positions on the object A1. FIG. 4 shows known distance data a1 [m] indicating the distance between the plurality of positions. However, the plurality of positions need not be limited to the upper end and the lower end of the object A1.
[0048] For example, the distance measurement determination unit 130 generates a question sentence asking about the size of the object A1, and acquires a response D1 "The size of the object A1 is a1 [m]." output from the language model based on inputting the generated question sentence to the language model stored in the prior knowledge storage unit 140. Then, the distance measurement determination unit 130 may acquire the distance between the plurality of positions from the response D1 as the known distance data a1 [m]. Alternatively, the distance measurement determination unit 130 may acquire the known distance data from a three-dimensional model.
[0049] (When the recognition target includes a marker) Figure 5 is a diagram for explaining the case where the recognition target includes the marker A2. As shown in FIG. 5, the recognition target that can be recognized based on image data may include the marker A2. For example, the marker A2 may be a marker enclosed with a product received by a user. Alternatively, the marker A2 may be a marker printed on paper by a user.
[0050] The distance measurement and determination unit 130 attempts to acquire known distance data indicating the distance between multiple feature points on marker A2. In the example shown in Figure 5, among the 6x4 grid of marker A2, the leftmost and rightmost points of the leftmost and rightmost points of any four consecutive horizontal grids are the multiple positions on marker A2. Figure 5 shows the known distance data a2 [m] indicating the distance between these multiple positions. However, the multiple positions are not limited to the leftmost and rightmost points of four horizontally aligned grids. For example, the multiple positions may be the leftmost and rightmost points of three horizontally aligned grids.
[0051] For example, the distance measurement and determination unit 130 may acquire known distance data from stored data stored by the prior knowledge storage unit 140.
[0052] (Including a Normalized Area) Figure 6 is a diagram illustrating the case where the recognition target includes a normalized area. As shown in Figure 6, a recognition target that can be recognized based on image data may include a normalized area A3. The size of the normalized area A3 is pre-normalized. For example, the normalized area A3 may be a field where a sports match (e.g., soccer, American football, or baseball) is played, or a studio where some kind of event is held. Figure 6 shows the case where the normalized area A3 is a field where a soccer match is played.
[0053] The distance measurement and determination unit 130 attempts to acquire known distance data indicating the distance between multiple locations in the standardized area A3. In the example shown in Figure 6, the two ends of the touchline in the standardized area A3 are the multiple locations in the standardized area A3. Figure 6 shows the known distance data a3 [m] indicating the distance between these multiple locations. However, the multiple locations are not limited to the two ends of the touchline. For example, the multiple locations may be the two ends of the goal line.
[0054] For example, if the standardized area A3 is a field where a sports match is held, the distance measurement and determination unit 130 may generate a question asking for the size of the standardized area A3, and then obtain a response output from the language model stored in the prior knowledge storage unit 140 based on the input of the generated question to the language model. The distance measurement and determination unit 130 may then obtain the distances between multiple locations as known distance data a3 [m] from the obtained response.
[0055] As another example, if standardized area A3 is a studio, the distance measurement and determination unit 130 may acquire known distance data from a three-dimensional model stored by the prior knowledge storage unit 140. Alternatively, if standardized area A3 is a studio, the distance measurement and determination unit 130 may acquire known distance data from a floor plan stored by the prior knowledge storage unit 140.
[0056] (Notification to the user) The determination that the recognition target is usable for scaling map M1 may be indicated by a display unit (not shown). This allows the user to understand that the recognition target is usable for scaling map M1. The user may also change the position and orientation of camera C1 so that the recognition target is photographed by camera C1 from multiple different locations. This allows multiple locations on the recognition target to be recognized by the distance measuring unit 150.
[0057] (Supplement regarding the object of recognition) In the following explanation, we will mainly assume that the object of recognition is object A1. However, the object of recognition may not be object A1, but marker A2 or normalization area A3. In other words, object A1 in the following explanation may be replaced with marker A2 or normalization area A3.
[0058] (Distance Measurement Unit 150) When the distance measurement unit 150 determines that object A1 can be used for scaling correction of map M1, it measures the distances between multiple locations on object A1 based on the image data and obtains data indicating the measured distances as first distance data (hereinafter also referred to as "measured distance data"). Figure 4 shows an example in which measured distance data y [m] indicating the distances between multiple locations on object A1 is obtained.
[0059] More specifically, the distance measurement unit 150 acquires multiple image data obtained by the camera C1. The distance measurement unit 150 also obtains measured distance data based on the acquired multiple image data and the position of the camera C1 on the map M1 at the time the camera C1 obtained the multiple image data. Below, an example of a method for obtaining measured distance data will be specifically described with reference to Figures 7 and 8.
[0060] Figure 7 shows an example of a combination of image data obtained by camera C1 and the position of camera C1. Referring to Figure 7, because the user changed the position of camera C1 while changing the orientation of camera C1, image data G1 to G3, which captures object A1 from different angles, are output from camera C1 in a continuous time series.
[0061] Furthermore, the self-position estimater 112 outputs the position P1 and orientation Q1 of camera C1 in map M1 when image data G1 is obtained, the position P2 and orientation Q2 of camera C1 in map M1 when image data G2 is obtained, and the position P3 and orientation Q3 of camera C1 in map M1 when image data G3 is obtained. In this way, multiple combinations of image data and the position and orientation of camera C1 can be obtained.
[0062] Of the multiple combinations obtained in this way, at least two combinations should be used to obtain the measured distance data. In the explanation using Figure 8, an example is shown in which image data G1 and G2, camera positions P1 and P2, and camera orientations Q1 and Q2 are used to obtain the measured distance data. However, to improve the accuracy of the measured distance data, three or more combinations may be used to obtain the measured distance data.
[0063] Figure 8 shows an example of a method for obtaining measured distance data. Figure 8 shows an example in which the distance measurement unit 150 recognizes positions F11 and F12 on object A1 from image data G1. Similarly, Figure 8 shows an example in which the distance measurement unit 150 recognizes positions F21 and F22 on object A1 from image data G2.
[0064] In the example shown in Figure 8, positions F11 and F21 correspond to the upper end of object A1, and positions F12 and F22 correspond to the lower end of object A1. However, positions F11 and F21 are not limited to positions corresponding to the upper end of object A1, and may be changed according to the positions corresponding to the known distance data described above. Similarly, positions F12 and F22 are not limited to positions corresponding to the lower end of object A1, and may be changed according to the positions corresponding to the known distance data described above.
[0065] The distance measurement unit 150 can estimate a straight line from the position P1 of camera C1 in map M1 to the upper end of object A1, based on the position P1 and orientation Q1 of camera C1 in map M1 when image data G1 is obtained, and the position F11 in image data G1. Similarly, the distance measurement unit 150 can estimate a straight line from the position P2 of camera C1 in map M1 to the upper end of object A1, based on the position P2 and orientation Q2 of camera C1 in map M1 when image data G2 is obtained, and the position F21 in image data G2.
[0066] The distance measuring unit 150 measures the upper end of object A1 using the principle of triangulation, based on a straight line from the position P1 of camera C1 on map M1 to the upper end of object A1, and a straight line from the position P2 of camera C1 on map M1 to the upper end of object A1.
[0067] Furthermore, the distance measurement unit 150 can estimate a straight line from the position P1 of camera C1 in map M1 to the lower end of object A1, based on the position P1 and orientation Q1 of camera C1 in map M1 when image data G1 is obtained, and the position F12 in image data G1. Similarly, the distance measurement unit 150 can estimate a straight line from the position P2 of camera C1 in map M1 to the lower end of object A1, based on the position P2 and orientation Q2 of camera C1 in map M1 when image data G2 is obtained, and the position F22 in image data G2.
[0068] The distance measuring unit 150 measures the lower end of object A1 using the principle of triangulation, based on a straight line from the position P1 of camera C1 on map M1 to the lower end of object A1, and a straight line from the position P2 of camera C1 on map M1 to the lower end of object A1.
[0069] The distance measuring unit 150 calculates the distance between the upper end and lower end of object A1 measured in this manner, and obtains data indicating the calculated distance as measured distance data y [m].
[0070] (Scale Estimation Unit 160) The scale estimation unit 160 calculates a scale value based on the ratio of known distance data to measured distance data. In the example above, the known distance data is a1 [m] and the measured distance data is y [m]. In this case, if the scale value is s, the scale estimation unit 160 may calculate the scale value s as a1 / y by dividing the measured distance data y [m] by the known distance data a1 [m].
[0071] (Scale Correction Unit 170) The scale correction unit 170 corrects the scale of map M1 based on the measured distance data and map M1. More specifically, the scale correction unit 170 corrects the scale of map M1 based on the scale value s estimated by the scale estimation unit 160 and map M1. An example of scale correction of map M1 will be explained with reference to Figure 9.
[0072] Figure 9 illustrates an example of scale correction for map M1. Referring to Figure 9, known distance data a1 [m] and measured distance data y [m] are shown. In this case, the scale value s is a1 / y. The scale correction unit 170 corrects the scale of map M1 by multiplying the scale of map M1 by the scale value s = a1 / y. The scale correction unit 170 outputs the scaled map M1 as map M2.
[0073] The scale correction unit 170 may unconditionally correct the scale of map M1. However, there may be cases where the scale value s becomes an abnormal value. For example, if a small model (miniature) made to resemble the real thing is used as the recognition target instead of the real thing, the scale value s may become an abnormal value. In addition, the scale value s may become an abnormal value if the recognition target is significantly deformed or if there is an error in the printing size of the marker used as an example of the recognition target.
[0074] Therefore, the scale correction unit 170 may determine whether the scale value s is appropriate by determining whether the scale value s falls within a predetermined range. If the scale correction unit 170 determines that the scale value s is appropriate, it may correct the scale of the map M1. On the other hand, if the scale correction unit 170 determines that the scale value s is not appropriate, it may reacquire the measured distance data without correcting the scale of the map M1.
[0075] As mentioned above, if the sensors used for self-localization include an IMU or LiDAR, a map M1 corresponding to the scale of the real world is created, and the position and orientation of camera C1 corresponding to the scale of the real world are estimated. Alternatively, the position and orientation of camera C1 corresponding to the scale of the real world can also be estimated by monocular depth estimation. When a map M1 corresponding to the scale of the real world is created in this way, the scale of map M1 corresponding to the scale of the real world is obtained, so it is considered meaningful to determine the validity of the scale value s.
[0076] The configuration of the information processing system 1 according to the embodiment of this disclosure has been described above.
[0077] [1.2. Processing Flow of the Information Processing System] The processing flow of the information processing system 1 according to the embodiment of this disclosure will be described with reference to Figure 10 (and to Figures 2 to 9 as appropriate). Figure 10 is a flowchart showing the processing flow of the information processing system 1 according to the embodiment of this disclosure.
[0078] The self-position estimation unit 110 creates map M1 (S11). Then, the self-position estimation unit 110 determines whether the creation of map M1 is complete or not (S12). If the self-position estimation unit 110 determines that the creation of map M1 is not complete (NO in S12), it returns to S11. On the other hand, if the self-position estimation unit 110 determines that the creation of map M1 is complete (YES in S12), it proceeds to S13.
[0079] The distance measurement and determination unit 130 recognizes an object based on image data obtained by the camera C1 (S13). Then, based on the recognition of the object, the distance measurement and determination unit 130 determines whether or not the object can be used to correct the scale of the map M1 (S14). To this end, the distance measurement and determination unit 130 attempts to acquire known distance data indicating the distances between multiple locations corresponding to the recognized object. The distance measurement and determination unit 130 can determine whether or not the recognized object can be used to correct the scale of the map M1 based on whether or not it succeeds in acquiring the known distance data.
[0080] If the distance measurement determination unit 130 determines that the object cannot be used to correct the scale of map M1 (NO in S14), it returns to S13. On the other hand, if the distance measurement determination unit 130 determines that the object can be used to correct the scale of map M1 (YES in S14), it proceeds to S15.
[0081] The distance measurement unit 150 measures the distance between multiple positions on an object based on the image data (S15). The distance measurement unit 150 then obtains data indicating the measured distance as measured distance data. The scale estimation unit 160 calculates a scale value based on the ratio of known distance data to measured distance data (S16).
[0082] The scale correction unit 170 determines whether the scale value is valid or not (S17). If the scale correction unit 170 determines that the scale value is not valid (NO in S17), it may return to S15. On the other hand, if the scale correction unit 170 determines that the scale value is valid (YES in S17), it corrects the scale of map M1 based on the scale value and map M1 (S18). The scale correction unit 170 outputs the scaled map M1 as map M2.
[0083] The processing flow of the information processing system 1 according to the embodiment of this disclosure has been described above.
[0084] [1.3. Effects of the Embodiment] According to the information processing system 1 of the embodiment of this disclosure, a technology is provided that makes it possible to modify the scale of the map while reducing the burden on the user.
[0085] Furthermore, according to the information processing system 1 of the embodiment of this disclosure, it becomes unnecessary to take photographs for map creation and photograph markers simultaneously. Furthermore, according to the information processing system 1 of the embodiment of this disclosure, even if the measurement of the marker size fails, it is not necessary to remeasure the marker size while creating the map again, thus reducing the processing cost required for map creation. Furthermore, according to the information processing system 1 of the embodiment of this disclosure, it is not necessary to include markers in the image data when creating the map, so it is possible to exclude unnecessary markers from the map.
[0086] The effects of the information processing system 1 according to the embodiment of this disclosure have been described above.
[0087] <2. Hardware Configuration Example> The information processing described above is realized through the cooperation of software and hardware. Below, an example of the hardware configuration of a computer 1000 that can be applied to the information processing system 1 according to the embodiment of this disclosure will be described.
[0088] Figure 11 is a hardware configuration diagram showing an example of a computer 1000 that implements the functions of the information processing system 1. The computer 1000 includes a processing circuitry 1100, RAM 1200, ROM 1300, secondary storage device 1400, communication interface 1500, input / output interface 1600, display unit 1700, camera unit 1800, microphone 1900, and speaker 2000. The various parts of the computer 1000 are connected by a bus 1050.
[0089] The processing circuit 1100 operates based on a program stored in the ROM 1300 or secondary storage device 1400, and controls each part. For example, the processing circuit 1100 loads the program stored in the ROM 1300 or secondary storage device 1400 into the RAM 1200 and executes processing corresponding to various programs.
[0090] ROM 1300 stores boot programs such as the BIOS (Basic Input Output System) that are executed by the processing circuit 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.
[0091] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the processing circuit 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records programs for each process of the information processing system 1 according to the embodiment of this disclosure, which is an example of program data 1450.
[0092] The communication interface 1500 is an interface for the computer 1000 to connect to the external network 1550. For example, the processing circuit 1100 can receive data from other devices or transmit data it has generated to other devices via the communication interface 1500.
[0093] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the processing circuit 1100 receives data from input devices such as a microphone 1900 or a touch panel via the input / output interface 1600. The processing circuit 1100 also transmits data to output devices such as a display unit 1700 or a speaker 2000 via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, or semiconductor memory.
[0094] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electroluminescent display (Organic Electro Luminescence Display). Alternatively, the display unit 1700 may be a touch panel display device or an image projection device.
[0095] The camera unit 1800 is an interface for the computer 1000 to capture images. The microphone 1900 is an interface for the computer 1000 to capture sound. The speaker 2000 is an interface for the computer 1000 to output processed sound. The various parts of the computer 1000 are connected by the bus 1050. Each interface does not necessarily have to be located inside the computer 1000, but may be located outside the computer 1000 via a network or the like. Furthermore, each part of the computer 1000 may be controlled by a circuit different from the processing circuit 1100. For example, the display unit 1700 may be controlled not by the processing circuit 1100, but by a circuit dedicated to display processing provided within the display unit 1700.
[0096] For example, when computer 1000 functions as an information processing system 1 according to an embodiment of this disclosure, the processing circuit 1100 of computer 1000 functions as a control unit (not shown) by executing a program loaded onto RAM 1200. The secondary storage device 1400 stores the information processing program according to this disclosure and various data stored in the map storage unit 120 and the prior knowledge storage unit 140. The processing circuit 1100 reads and executes the program data 1450 from the secondary storage device 1400, but as another example, these programs may be obtained from other devices via an external network 1550. In other words, the secondary storage device 1400 is not limited to being inside computer 1000, but may be located outside computer 1000. The processing circuit 1100 is an example of an integrated circuit, and CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered integrated circuits.
[0097] The above describes an example of the hardware configuration of a computer 1000 that can be applied to the information processing system 1 according to the embodiment of this disclosure.
[0098] <3. Modifications> Although preferred embodiments of the present disclosure have been described in detail above with reference to the attached drawings, the technical scope of the present disclosure is not limited to these examples. It is clear that a person with ordinary skill in the art of the present disclosure may conceive of various modifications or alterations within the scope of the technical ideas described in the claims, and these will naturally also fall within the technical scope of the present disclosure.
[0099] For example, the functions of the information processing system 1 according to the embodiment of this disclosure may be performed by a computer, but the type of computer into which the functions of the information processing system 1 are incorporated is not limited. For example, the functions of the information processing system 1 may be incorporated into various computers that perform self-localization while creating a map (e.g., a camera, an autonomous robot, or AR glasses).
[0100] Furthermore, the above describes the case in which measured distance data is obtained from multiple image data using the principle of triangulation. However, there are cases in which data indicating the distance between multiple positions corresponding to the recognition object can be measured using LiDAR or monocular depth estimation. In such cases, instead of obtaining measured distance data from multiple image data using the principle of triangulation, the distance measurement determination unit 130 may obtain data indicating the distance between multiple positions corresponding to the recognition object from a single image data as measured distance data.
[0101] Furthermore, among the processes described in the embodiments of this disclosure described above, all or part of the processes described as being performed automatically may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0102] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0103] Furthermore, the embodiments of this disclosure described above can be combined as appropriate in areas where the processing content is not contradictory. Also, the order of each step shown in the sequence diagram or flowchart of this embodiment can be changed as appropriate. For example, each step may be processed chronologically, repeatedly, or partially in parallel.
[0104] Furthermore, the effects described herein are merely descriptive or illustrative and not limiting. In other words, the technology relating to this disclosure may produce other effects that are obvious to those skilled in the art from the description herein, in addition to or in lieu of the effects described herein.
[0105] The following configurations also fall within the technical scope of this disclosure: (1) An information processing method performed by a processor, comprising: acquiring a map created based on image data; recognizing a recognition target based on image data obtained by a camera; and outputting a map whose scale has been corrected based on first distance data relating to the recognition target and the map. (2) The information processing method according to (1), wherein the processor obtains data indicating the distance between a plurality of positions corresponding to the recognition target as the first distance data. (3) The information processing method according to (2), wherein the recognition target includes a marker. (4) The information processing method according to (2), wherein the recognition target includes a normalized area. (5) The information processing method according to (2), wherein the recognition target includes an object. (6) The information processing method according to any one of (1) to (5), wherein the processor attempts to acquire second distance data relating to the recognition target, and if successful in acquiring the second distance data, corrects the scale of the map based on the first distance data, the second distance data, and the map. (7) The information processing method according to (6), wherein the processor attempts to obtain the second distance data from a pre-prepared machine learning model or pre-stored stored data. (8) The information processing method according to (6) or (7), wherein the processor calculates a ratio based on the second distance data and the first distance data as a scale value, and modifies the scale of the map based on the scale value and the map. (9) The information processing method according to (8), wherein the processor determines whether the scale value falls within a predetermined range, and modifies the scale of the map based on the determination that the scale value falls within the predetermined range. (10) The information processing method according to any one of (1) to (9), wherein the processor obtains the first distance data based on a plurality of image data obtained by the camera and the position of the camera on the map when the camera obtained the plurality of image data.(11) The information processing method according to (10), wherein the processor estimates the position of the camera by self-position estimation based on sensing data obtained by a predetermined sensor. (12) The information processing method according to (11), wherein the predetermined sensor includes an inertial measuring device. (13) An information processing system comprising a processor that acquires a map created based on image data, recognizes a recognition object based on image data obtained by a camera, and outputs a map whose scale has been corrected based on first distance data relating to the recognition object and the map. (14) The information processing system according to (13), wherein the processor obtains data indicating the distance between a plurality of positions corresponding to the recognition object as the first distance data. (15) The information processing system according to (13) or (14), wherein the processor attempts to acquire second distance data relating to the recognition object, and if successful in acquiring the second distance data, corrects the scale of the map based on the first distance data, the second distance data and the map. (16) The information processing system according to (15), wherein the processor attempts to obtain the second distance data from a pre-prepared machine learning model or pre-stored stored data. (17) The information processing system according to (15) or (16), wherein the processor calculates a ratio based on the second distance data and the first distance data as a scale value, and modifies the scale of the map based on the scale value and the map. (18) The information processing system according to (17), wherein the processor determines whether the scale value falls within a predetermined range, and modifies the scale of the map based on the determination that the scale value falls within the predetermined range. (19) The information processing system according to any one of (13) to (18), wherein the processor obtains the first distance data based on a plurality of image data obtained by the camera and the position of the camera on the map when the camera obtained the plurality of image data.(20) The information processing system according to (19), wherein the processor estimates the position of the camera by self-position estimation based on sensing data obtained by a predetermined sensor.
[0106] 1. Information Processing System 110 Self-Position Estimation Unit 120 Map Storage Unit 130 Distance Measurement and Determination Unit 140 Prior Knowledge Storage Unit 150 Distance Measurement Unit 160 Scale Estimation Unit 170 Scale Correction Unit
Claims
1. An information processing method performed by a processor, comprising: acquiring a map created based on image data; recognizing a recognition target based on image data obtained by a camera; and outputting a scaled map based on first distance data relating to the recognition target and the map.
2. The information processing method according to claim 1, wherein the processor obtains data indicating the distance between a plurality of locations corresponding to the recognized object as the first distance data.
3. The information processing method according to claim 2, wherein the object to be recognized includes a marker.
4. The information processing method according to claim 2, wherein the recognition target includes a standardized area.
5. The information processing method according to claim 2, wherein the object to be recognized includes an object.
6. The information processing method according to claim 1, wherein the processor attempts to acquire second distance data relating to the recognized object, and if successful in acquiring the second distance data, modifies the scale of the map based on the first distance data, the second distance data, and the map.
7. The information processing method according to claim 6, wherein the processor attempts to obtain the second distance data from a pre-prepared machine learning model or pre-stored stored data.
8. The information processing method according to claim 6, wherein the processor calculates a ratio based on the second distance data and the first distance data as a scale value, and modifies the scale of the map based on the scale value and the map.
9. The information processing method according to claim 8, wherein the processor determines whether the scale value falls within a predetermined range, and modifies the scale of the map based on the determination that the scale value falls within the predetermined range.
10. The information processing method according to claim 1, wherein the processor obtains the first distance data based on a plurality of image data obtained by the camera and the position of the camera on the map at the time the camera obtained the plurality of image data.
11. The information processing method according to claim 10, wherein the processor estimates the position of the camera by self-position estimation based on sensing data obtained by a predetermined sensor.
12. The information processing method according to claim 11, wherein the predetermined sensor includes an inertial measuring device.
13. An information processing system comprising a processor that acquires a map created based on image data, recognizes a recognition target based on image data obtained by a camera, and outputs a scaled map based on first distance data relating to the recognition target and the map.
14. The information processing system according to claim 13, wherein the processor obtains data indicating the distance between a plurality of locations corresponding to the recognized object as the first distance data.
15. The information processing system according to claim 13, wherein the processor attempts to acquire second distance data relating to the recognized object, and if successful in acquiring the second distance data, modifies the scale of the map based on the first distance data, the second distance data, and the map.
16. The information processing system according to claim 15, wherein the processor attempts to obtain the second distance data from a pre-prepared machine learning model or pre-stored stored data.
17. The information processing system according to claim 15, wherein the processor calculates a ratio based on the second distance data and the first distance data as a scale value, and modifies the scale of the map based on the scale value and the map.
18. The information processing system according to claim 17, wherein the processor determines whether the scale value falls within a predetermined range, and modifies the scale of the map based on the determination that the scale value falls within the predetermined range.
19. The information processing system according to claim 13, wherein the processor obtains the first distance data based on a plurality of image data obtained by the camera and the position of the camera on the map at the time the camera obtained the plurality of image data.
20. The information processing system according to claim 19, wherein the processor estimates the position of the camera by self-position estimation based on sensing data obtained by a predetermined sensor.