Information processing device, image processing method, and computer program

The image processing device evaluates and stores images with high symbolic and visible content to enhance map understanding, addressing the unintuitive nature of map information and memory constraints in self-location estimation.

JP2025183550APending Publication Date: 2025-12-17CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024091221
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-12-17

AI Technical Summary

Technical Problem

Map information used in self-location estimation is not intuitively understandable to humans, and storing all landscape images or videos along with the map information is impractical due to memory consumption.

Method used

An image processing device that evaluates images based on symbolicity and visibility, storing only those images with high symbolic and visible content in association with the map, using OCR and AI-OCR for character recognition, and deep learning for image evaluation.

Benefits of technology

Enables efficient understanding of maps by associating high-symbolic and visible images with map information, reducing memory usage and enhancing user comprehension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025183550000001_ABST
    Figure 2025183550000001_ABST
Patent Text Reader

Abstract

To solve the problem in which: map information used in self-location estimation is not intuitively understandable to humans, resulting that it takes time for humans to understand the map; furthermore, saving all landscape images along with the map information or saving videos therewith is not practical because it would take up too much storage space.SOLUTION: A device includes: image acquisition means for acquiring images of an environment captured by an imaging device mounted on a mobile object; location information acquisition means for acquiring location information obtained from estimating the location of the mobile object; map acquisition means for acquiring a map generated based on the location information of the mobile object; image evaluation means for evaluating the images; and image storage means for storing, based on the evaluation, the images with the evaluation higher than a predetermined threshold, in association with the map.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technology for an image processing device that performs self-position estimation. [Background technology]

[0002] For example, when a mobile object such as a self-propelled robot is caused to travel within a factory environment, in order to stably control the movement of the mobile object, a method is known in which the amount of movement is controlled based on the self-position created from image information, as in Patent Document 1. Also, Patent Document 2 discloses a technology that clearly indicates locations where self-position estimation is likely to be erroneous. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7341652 [Patent Document 2] Patent Publication No. 2023-81005 Summary of the Invention [Problem to be solved by the invention]

[0004] However, map information used in self-location estimation is not intuitively understandable to humans, and so it takes time for people to understand the map. Furthermore, storing all landscape images or videos along with the map information is not practical because it would consume memory space. The present invention has been made in consideration of the above-mentioned problems, and aims to provide an image processing device that stores images for efficiently understanding a map created by self-location estimation. [Means for solving the problem]

[0005] The image processing device of the present invention comprises an image acquisition means for acquiring an image of an environment photographed by an imaging device mounted on a mobile body, a location information acquisition means for acquiring location information that estimates the location of the mobile body, a map acquisition means for acquiring a map generated based on the location information of the mobile body, an image evaluation means for evaluating the image, and an image storage means for storing the image with the evaluation higher than a predetermined threshold value and the map in association with each other based on the evaluation. [Effects of the Invention]

[0006] According to the present invention, in a map created by self-localization estimation, a person can efficiently understand the map by checking the map together with an image having symbolic properties and visibility. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram showing the basic configuration of a system according to a first embodiment. [Figure 2] FIG. 1 shows a basic configuration of a first embodiment. [Figure 3] FIG. 1 shows an example of an image according to the first embodiment. [Figure 4] Flowchart showing the processing flow of the first embodiment [Figure 5] FIG. 10 is a diagram showing an example of displaying a map and a saved image according to the first embodiment. [Figure 6] FIG. 10 is a diagram showing an example of displaying a new image in the first embodiment. [Figure 7] FIG. 10 is a diagram showing an example of displaying a plurality of related maps according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] [Embodiment 1] Hereinafter, embodiments will be described with reference to the drawings. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the configurations shown in the drawings.

[0009] In this embodiment, a mobile body called an AGV (Automated Guided Vehicle) will be used for explanation. The following description will be given taking an AGV as an example of a mobile body.

[0010] 1 shows a system configuration diagram of this embodiment. An information processing system 100 in this embodiment is made up of multiple mobile objects 101 (101-1, 101-2, ...), a process control system 103, and a mobile object control system 102. The information processing system 100 is a logistics system, a production system, or the like.

[0011] A plurality of mobile objects 101 (101-1, 101-2, ...) are guided vehicles (AGVs) that transport objects according to a process schedule determined by a process control system. A plurality of mobile objects are moving (traveling) within the environment.

[0012] The process control system 103 manages processes executed by an information processing system. For example, it is an MES (Manufacturing Execution System) that manages processes in a factory or a logistics warehouse. It communicates with the mobile object control system 102. The mobile object control system 102 is a system that manages mobile objects. It communicates with the process control system 12. It also communicates with the mobile objects (for example, Wi-Fi communication) to send and receive operation information in both directions.

[0013] 2 shows the basic configuration of the invention in this embodiment. The image processing device 200 is composed of a self-position estimation unit 201, an image evaluation unit 202, and an image storage unit 203. In addition, there is an image acquisition unit 204 that acquires images from an imaging device such as a camera, a sensor information acquisition unit 205 that acquires information from a distance sensor, an acceleration sensor, an infrared sensor, etc., and an input device 206 that accepts user operations such as a mouse, keyboard, or touch operation. A driving control unit 208 acquires position information 207 from the self-position estimation unit 201 and controls the moving object. In other words, the self-position estimation unit 201 also serves as a position information acquisition unit.

[0014] The self-location estimation unit 201 is configured as a SLAM (Simultaneous Localization and Mapping) unit that creates map information based on the image information acquired by the image acquisition unit 204 and estimates the self-location. The SLAM technology generates three-dimensional map information to acquire information such as the positions of objects in the environment, and performs a self-location estimation process using the map information. The three-dimensional map information is composed of map elements (hereinafter referred to as keyframes). In other words, the self-location estimation unit 201 also plays the role of a map information acquisition unit.

[0015] Furthermore, the self-location estimation unit 201 generates a map based on the location estimation, that is, the self-location estimation unit 201 also functions as a map acquisition unit.

[0016] The device that processes SLAM matches image features obtained from images of the environment captured by an imaging device mounted on a moving object with image features obtained from key frames contained in pre-created 3D map information. The device then estimates the self-position of the imaging device mounted on the moving object. The positional relationship between the imaging device and the moving object is known, and the self-position of the moving object can be estimated by converting the position of the imaging device into the position of the moving object.

[0017] Many SLAM methods have been proposed and can be used. For example, a method can be used in which feature values ​​(Haar-Like features, HOG features, etc.) obtained from images taken from multiple different viewpoints using algorithms such as SIFT or ORB are saved as keyframes and then matched to estimate self-location. Another method can be used in which brightness values ​​of images taken from multiple different viewpoints are compared to estimate self-location so as to minimize error. Other well-known methods that can be used include RGB-D SLAM, which generates a map while tracking the depth of feature points detected from an image as the depth value of a depth sensor. Furthermore, a position estimation method using a deep learning-based technique can also be used.

[0018] The image evaluation unit 202 evaluates the symbolicity and visibility of the image acquired by the image acquisition unit 204 .

[0019] As a method for evaluating symbolicity, the number and size of characters in an image can be evaluated using OCR (Optical Character Recognition) analysis technology, which performs character recognition. Alternatively, the number and size of characters in an image can be evaluated using AI-OCR technology, which incorporates AI (Artificial Intelligence) technology into OCR and performs character recognition using machine learning or deep learning. Alternatively, the number and size of characters in an image can be evaluated using technology such as scene character string detection. Images containing many characters and large characters are rated higher. For example, images containing five characters are rated higher than images containing three characters and saved. Furthermore, images containing large characters are rated higher than images containing small characters. For example, a character size that is difficult for humans to recognize can be set as a predetermined threshold, and characters below the predetermined threshold, i.e., small characters, can be excluded. In this way, images with symbolicity that can be recognized by humans can be saved.

[0020] Visibility evaluation methods combine basic image features such as brightness, color, and edges in an image with information on temporal and spatial contrast. For example, because the purpose of visual recognition of signs and billboards is to make them visible to humans, they can be detected by extracting bright colors such as red and yellow from the image and performing shape recognition. Alternatively, template images of the signs and billboards to be detected may be prepared in advance for color regions and edges, and signs and billboards with high similarity may be detected. Furthermore, sign and billboard regions may be determined and evaluated using segmentation such as deep learning. For example, brightness, color, and edges that can be recognized by humans may be set as predetermined thresholds, and images below the predetermined thresholds, i.e., images with low brightness values ​​or dark colors, may be excluded. In this way, images with visibility that can be recognized by humans can be saved.

[0021] As mentioned above, signs and billboards contain the information people need to intuitively understand. Images containing such signs and billboards are highly symbolic and highly visible, making them suitable for helping people understand maps.

[0022] In this way, both symbolicity and visibility are evaluated to evaluate whether the map created by the self-location estimation unit 201 is an image that can be understood by people. Furthermore, by linking highly rated images with maps, it becomes possible to store images that make it easier for people to understand the map.

[0023] Even when the evaluation values ​​for symbolicity and visibility are low, or even when the evaluation values ​​for symbolicity and visibility are high, in addition to the evaluation of symbolicity and visibility, images of the AGV's operation and the starting points on the map, such as the start point, end point, intersections, and branches, are saved. Because the AGV's movement route and operation based on the movement route are determined by the user, the start point, end point, intersections, and branches are places that the user takes into consideration. For this reason, they are useful as images of places that help people understand the map. In this way, by saving images that take location into consideration, it is possible to save images that allow people to understand the map.

[0024] In addition, a list of landmark buildings and spaces and a detector for them may be prepared in advance, and if a captured image contains a landmark, it may be saved as an image that allows humans to understand the map. Landmarks are lighthouses, steel towers, and other distinctive buildings and spaces that serve as landmarks in a location. Images that contain such landmarks are suitable for humans to understand the map from the distinctive buildings and spaces.

[0025] The image storage unit 203 stores the evaluated image in association with the map. As a storage method, the image may be included in the map file, or the image may be stored separately from the map. Alternatively, the image may include location information or coordinates and be stored, and linked to the map.

[0026] The image acquisition unit 204 may acquire images obtained from an imaging device other than the imaging device used for self-location estimation. The imaging device used for self-location estimation may have a fixed angle for self-location estimation, which may prevent it from obtaining images with high symbolicity and visibility. For this reason, images may be acquired from a different imaging device. The imaging device may also be an omnidirectional imaging device, and an image may be cut out from an omnidirectional image. In this case, the difference in the direction of the imaging device used for self-location estimation may be saved and displayed. This allows the difference in direction from the imaging device used for self-location estimation to be displayed, thereby supplementing information about the direction of the image used for map understanding.

[0027] If the AGV is equipped with sensors such as inertial sensors such as gyros and IMUs, and encoders for acquiring the amount of tire rotation, the sensor information acquisition unit 205 inputs the sensor values. The self-position estimation unit 201 can also estimate the self-position of the image acquisition unit 204 by using the sensor values ​​in combination. In this way, by using the image information of the image acquisition unit 204 in combination with the sensor information, the self-position can be estimated with high accuracy and robustness.

[0028] The accuracy of these SLAMs and the methods for calculating the accuracy vary depending on the method. For example, there are no limitations as long as the value can be calculated using features, scores, etc. that can be obtained by an algorithm, etc.

[0029] Next, an example of an image in this embodiment will be described with reference to FIG.

[0030] Image 300 is an image in which edges, feature points, and corner points exist on walls, pipes, etc. However, it is difficult for a person to identify a location in an image that only shows walls or pipes. For this reason, image 300 can be evaluated as an image with low symbolism and visibility. Such an image is not suitable for a person to understand where the map is.

[0031] Next, image 301 has the letters AAA (place name), making it highly symbolic and easy for people to understand. Also, there is an arrow mark, and if the mark's color is red, which stands out compared to the surrounding environment, it is easy for people to see. Therefore, image 301 can be evaluated as an image with high symbolic value and visibility. Such an image is suitable for people to understand the map.

[0032] Next, the flow of this embodiment will be described with reference to FIG.

[0033] In step S401, an image is acquired from the image acquisition unit 204. Next, in step S402, the symbolism and visibility of the image are evaluated. Next, in step S403, an image is selected based on the evaluation. Next, in step S404, the selected image is saved as an image of the map. Next, in step S405, it is determined whether to end the process, and if it is determined that the process is ended, the process ends. Through this flow, an image suitable for a person to understand the map is saved.

[0034] Furthermore, when evaluating the image in step S402, moving objects may be detected and excluded from the evaluation. Unlike signs or landmarks, moving objects move from their location and are therefore not suitable as information for people to recognize locations. For example, moving objects may include people or other AGVs. However, moving objects that move within a predetermined range, such as machines that perform repetitive motion, may also be included in the evaluation.

[0035] Furthermore, when selecting images in step S403, images with high symbolic value and visibility may be selected at specified intervals such as distance, time, or key frame. For example, images may be selected at intervals of 5 minutes, 500 meters, or 50 key frames. Key frames are created based on the amount of change in the image, so by saving images based on the key frame interval, it is possible to save an appropriate number of images in accordance with changes in the scenery.

[0036] You can also select images based on the AGV's operation and the start and end points, intersections, and branches on the map. For example, you can select images of the AGV just before or after a turn, or the start and end points. By saving images like this, you can save images in locations that are convenient for the user.

[0037] By specifying the interval, the number of images saved can be reduced, thereby reducing the resources used by the image processing device. When checking multiple maps, they are displayed and checked on an information device such as a PC (Personal Computer), but the display and storage areas are limited. Therefore, by reducing resources, it becomes possible to save and display images that people need to understand the map using limited resources. In addition, since equipment such as AGVs operates in an environment with limitations such as battery life, preventing unnecessary calculations can enable them to operate for longer periods of time.

[0038] Next, an embodiment for displaying a map and a saved image will be described with reference to FIG.

[0039] In this embodiment, a window for presenting two-dimensional map information will be described, but three-dimensional map information may also be presented. A display method for the image acquired by the image acquisition unit 204, the map created by the self-position estimation unit 201, and the image evaluated by the image evaluation unit 202 and stored in the image storage unit 203 will be described.

[0040] Display example 500 is an example of displaying a map and an image. Environmental map 504 is a simple map of the real world. Map 501 shows the shape of the map created by self-location estimation. As shown in map 501, the image location and direction where the image was taken may be indicated. Also, an arrow or the like indicating the direction when the map was created may be indicated. Image 502 is an image selected by evaluating the image on the map 501 for symbolism and visibility, and is an image of point 1-A on environmental map 504.

[0041] By displaying the map in a list like this, people can see the highly symbolic and highly visible images associated with the map, which allows them to understand the map efficiently due to the image superiority effect.

[0042] Image 503 is an example in which the direction in which the image was taken is displayed on the image, and is an image taken from the opposite direction to image 502. By displaying the direction in which the image was taken by the imaging device, the directions of the map and the image can be understood, and the direction on the map can be quickly understood.

[0043] By displaying a map created by self-location estimation on the environmental map 504, it becomes possible to understand the location of the created map. Map 505 displays multiple maps created by driving through part of the environmental map 504 using dotted lines. When multiple maps are displayed on top of each other in this way, the overlapping of many lines can make it difficult to understand.

[0044] Therefore, by displaying images with high symbolic value and visibility together with the map as in the example display 500, it becomes easier for people to understand multiple maps.

[0045] Next, a method for displaying a new image will be described with reference to FIG.

[0046] When the AGV travels while referring to the map created by the self-position estimation unit 201, the image acquisition unit 204 acquires a new image, which is evaluated by the image evaluation unit 202 and stored in the image storage unit 203. This section explains how the new image is displayed.

[0047] The display example 600 displays an image 502 created at the time of map creation and an image 601 newly acquired at the location where the highly rated image 502 was previously acquired while the AGV was traveling. This is displayed after determining that the location information included in image 502 matches or is similar to the location information included in image 601. By displaying the image in this manner, the user can check the past image and the most recent image. This enables linking and understanding past and present memories. Furthermore, image 602 is displayed when an image highly rated for symbolism and visibility is saved during a trip at a location other than where images were saved in the past. In this way, displaying a newly highly rated image makes it possible to link and display the image with highly rated images in the most recent environment, thereby enabling efficient map understanding.

[0048] Next, a method for displaying a related map by specifying an image or location information will be described with reference to FIG. 7. Display example 700 is an example in which, when image 702 indicating location B on map 701 is specified, map 703 including location B is displayed. The specified image 702 includes location information indicating location B, and map 703 including the location information indicating location B is displayed. Alternatively, only location information may be specified, and a map including the specified location information may be displayed. If there is a map other than map 703 that includes location information indicating location B, that may also be displayed. Furthermore, images or symbolism and visibility associated with image 702 and map 703 may be compared, and the corresponding map 703 may be displayed if the degree of match between the images exceeds a threshold.

[0049] In this way, by displaying related maps, the map can be compared with other maps. By comparing multiple maps, the extent of the map can be understood from the differences between the map, and people can understand the map.

[0050] As described above, by displaying an image with high symbolic value and visibility together with map information, people can check the map together with the image, which makes it easier to understand the map.

[0051] In this embodiment, the moving object is not limited to an AGV (Automated Guided Vehicle). For example, the moving object may be an autonomous mobile robot, an automated guided forklift, a self-driving car, or a drone. Alternatively, the image processing device described in this embodiment may be applied.

[0052] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. The program is also called a computer program. The present invention can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions. [Explanation of symbols]

[0053] 100 Information Processing Systems 101 Mobile Systems 102 Mobile Management System 103 Process Control System

Claims

1. image acquisition means for acquiring an image of the environment captured by an imaging device mounted on the moving body; a position information acquisition means for acquiring position information that estimates the position of the moving object; a map acquisition means for acquiring a map generated based on the position information of the mobile object; an image evaluation means for evaluating the image; and an image storage means for storing the image having the evaluation higher than a predetermined threshold value and the map in association with each other based on the evaluation.

1. An image processing device comprising:

2. the image storage means stores images based on a predetermined key frame interval; 2. The image processing device according to claim 1, wherein:

3. the image evaluation means evaluates the image based on at least one of visibility determined based on the brightness, color, and edges of the image, and symbolism determined based on the size and number of characters included in the image.

2. The image processing device according to claim 1, wherein:

4. the image evaluation means evaluates the image based on the movement of the moving object when the visibility and the symbolicity are lower than predetermined thresholds.

2. The image processing device according to claim 1, wherein:

5. the image evaluation means evaluates the image based on a movement path of the moving object when the visibility and the symbolicity are lower than predetermined thresholds.

2. The image processing device according to claim 1, wherein:

6. the image acquisition means acquires the image and position information of the moving object when the moving object moves around based on the map; the image storage means stores the image based on the position information of the image stored in the image storage means and the position information of the image; 2. The image processing device according to claim 1, wherein:

7. a designation means for designating at least one of the image and the position information; a display means for displaying the image and the map; the display means displays a map including either the image designated by the designation means or the location information.

2. The image processing device according to claim 1, wherein:

8. a second image acquisition means for acquiring a second image; the image evaluation means evaluates the second image based on the symbolism and the visibility, the image storage means stores the second image having the evaluation higher than a predetermined threshold value and the location information in association with the map based on the evaluation; 2. The image processing device according to claim 1, wherein:

9. an image acquisition step of acquiring an image of the environment by an imaging device mounted on the moving body; a position information acquisition step of acquiring position information that estimates the position of the moving object; a map acquisition step of acquiring a map created by the self-location estimation of the moving object and consisting of a plurality of map elements; an image evaluation step of evaluating the image; and an image storage step of storing the image having the evaluation higher than a predetermined threshold value in association with the map based on the evaluation.

10. An image processing method executed by an information processing device.

10. A computer program for causing a computer to execute each step of the image processing method according to claim 9.

11. image acquisition means for acquiring an image of the environment captured by an imaging device mounted on the moving body; a map acquisition means for acquiring a map created by self-location estimation of the moving object; an image evaluation means for evaluating the image based on at least one of visibility determined based on the brightness, color, and edges of the image, and symbolism determined based on the size and number of characters included in the image; and an image storage means for storing the image having the evaluation higher than a predetermined threshold value and the map in association with each other based on the evaluation.

1. An image processing device comprising:

Citation Information

Patent Citations

  • Information processing apparatus, information processing method and program

    JP2023081005A

  • Information processing device, information processing method, program, and system

    JP7341652B2