Information processing device, method, and program
The information processing device enhances no-entry area setting for autonomous mobile objects by using semantic segmentation to categorize areas accurately, improving path adherence in complex environments.
Patent Information
- Application Number
- JP2023166832
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-09-28
AI Technical Summary
Setting no-entry areas for autonomous mobile objects is cumbersome and prone to errors, especially when the areas are large or complex, leading to undesired travel paths.
An information processing device that uses a semantic segmentation model to categorize areas into no-entry, permitted, and intermediate zones based on image analysis, integrating this information into a 3D data structure for accurate route planning.
Improves the setting of no-entry areas by reducing errors and ensuring autonomous mobile objects follow desired paths, even in complex environments.
Smart Images

Figure 0007763223000001 
Figure 0007763223000002 
Figure 0007763223000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, method, and program. [Background technology]
[0002] Conventionally, autonomous mobile objects such as autonomously traveling service robots have been known. Such autonomous mobile objects have a defined area in which they can move. Generally, in the case of cleaning robots that perform cleaning work in commercial facilities and security robots that perform patrol security, the area managed by a facility management company is limited to the common areas of the facility and may not include management of areas occupied by tenants. In such cases, when a facility management company uses a cleaning robot to perform cleaning work, the cleaning robot only travels in the common areas of the facility, and the security robot also only patrols in the common areas.
[0003] As such, cleaning robots and security robots, which are autonomous mobile bodies, need to travel by distinguishing between permitted and prohibited areas.
[0004] In order to set a no-entry area for an autonomous moving body, it is known to manually move the autonomous moving body and set the no-entry area based on the movement trajectory of the autonomous moving body (see Patent Document 1). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2022-86464 [Non-patent literature]
[0006] [Non-Patent Document 1] F.Dellaert,D.Fox,W.Burgard,S.Thrun, Monte Carlo localization for mobile robotS, Proc.of IEEE International Conference on RoboticS and Automation,1999 [Non-patent document 2] Dieter Fox, Adapting the Sample Size in particle filterS through KLD-Sampling, International Journal of RoboticS ReSearch,2003 [Non-patent document 3] Giorgio GriSetti,Cyrill StachniSS,Wolfram Burgard, Improved techniqueS for grid mapping with rao-blackwellized particle filterS, IEEE tranSactionS on RoboticS,2007 [Non-patent document 4] Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam, Encoder-Decoder with AtrouS Separable Convolution for Semantic Image Segmentation, European Conference on Computer ViSion,2018 Summary of the Invention [Problem to be solved by the invention]
[0007] However, when setting a no-entry area based on the movement trajectory of an autonomous mobile body, there is a problem that the setting work is cumbersome when the no-entry area is large or has a complex shape, and there is a high possibility that a setting error occurs. In this case, it is not possible to make the autonomous mobile body travel as desired. As such, there has been room for improvement in setting a no-entry area for an autonomous mobile body.
[0008] In view of the above-mentioned problems, an object of the present invention is to improve the setting of no-entry areas for autonomous moving bodies. [Means for solving the problem]
[0009] An information processing apparatus according to an embodiment of the present invention is In commercial facilities means for acquiring an image of an environment; and Determine the area to work on input to a semantic segmentation model, and as an output of the semantic segmentation model, Tenant areas of said commercial facility that are not subject to work a no-entry area in which the moving object is prohibited from entering; the common areas that are the subject of work in said commercial facility; an access permission area in which the mobile object is permitted to enter; and It is difficult to determine whether the tenant area is a common area or a common area. An intermediate area located between the no-entry area and the permitted area 、 At least three categories of Get and means. [Effects of the Invention]
[0010] According to the present invention, it is possible to improve the setting of no-entry areas for autonomous moving bodies. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram showing a configuration of a mobile object management system 101 according to a first embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of a mobile system. [Figure 3] FIG. 1 is a diagram illustrating an example of a functional configuration of a mobile system. [Figure 4] FIG. 1 is a diagram illustrating a work flow of a mobile system. [Figure 5]FIG. 2 illustrates an example of a hardware configuration of a server. [Figure 6] FIG. 2 illustrates an example of a functional configuration of a server. [Figure 7] FIG. 4 is a diagram illustrating a processing flow for creating a no-entry area according to the first embodiment. [Figure 8] FIG. 4 is a diagram showing an example of creating a no-entry area according to the first embodiment. [Figure 9] FIG. 10 is a diagram illustrating a processing flow for creating a no-entry area according to the second embodiment. [Figure 10] FIG. 11 is a diagram showing an example of creating a no-entry area according to the second embodiment. [Figure 11] FIG. 10 is a diagram showing an example of an overhead view of a no-entry area and an entry-permitted area. [Figure 12] FIG. 11 is a diagram showing an example of an overhead view of a no-entry area according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0013] In the following, as an example, a case where the present invention is applied to a mobile object management system in which a mobile object performs cleaning work will be described. However, the present invention is not limited to this, and can be applied to any system that analyzes sensor information and outputs predetermined information.
[0014] [First embodiment] <System configuration> 1 is a diagram showing the configuration of a mobile object management system 101 according to a first embodiment of the present invention. The mobile object management system 101 includes mobile object systems 110a, 11b, and 110c, a network 120, and a server 130. Note that, hereinafter, each of the mobile object systems 110a, 11b, and 110c may be referred to as a mobile object system 110.
[0015] The mobile system 110 is an autonomously moving mobile body equipped with the functions necessary for cleaning work. In the first embodiment, the mobile system 110 has a built-in computing device capable of processing sensor information such as images, but the present invention is not limited to this. For example, an external computer such as a PC may be connected to the mobile system 110, and a combination of these may be treated as the mobile system 110. PC is an abbreviation for personal computer.
[0016] The server 130 is a computer such as a PC, and has a function of analyzing and processing sensor information such as images. Furthermore, the server 130 is a device that receives input from a user and outputs information to the user (for example, displays information).
[0017] The mobile system 110 and the server 130 are communicatively connected via a network 120. The network 120 includes a plurality of routers, switches, cables, and the like that comply with a communication standard such as Ethernet (registered trademark). In this embodiment, the network 120 may be any network that enables communication between the mobile system 110 and the server 130, and may be constructed with any scale, configuration, and conforming communication standard. For example, the network 120 may be the Internet, a wired LAN, a wireless LAN, a WAN, or the like. LAN is an abbreviation for Local Area Network. WAN is an abbreviation for Wide Area Network. The network 120 may be configured to enable communication using a communication protocol that complies with the ONVIF standard, for example. ONVIF is an abbreviation for Open Network Video Interface Forum. However, the present invention is not limited to this, and the network 120 may be configured to enable communication using another communication protocol, for example, a proprietary communication protocol.
[0018] <Device configuration> (Hardware configuration of the mobile system 110) Next, the configuration of the mobile body system 110 will be described. Fig. 2 is a diagram showing an example of the hardware configuration of the mobile body system 110. The mobile body system 110 has a camera 201, a LiDAR 202, an information processing device 203, a communication unit 204, a control device 205, and a wheel unit 206. LiDAR is an abbreviation for Light Detection and Ranging. The camera 201 and the LiDAR 202 constitute a sensor unit 200.
[0019] The camera 201 is a distance imaging camera, such as a TOF camera that measures distances using a TOF method. TOF stands for Time of Flight. The camera 201 can acquire distance information in addition to RGB image information. The camera 201 includes a lens unit for focusing light, an image sensor that converts the focused light into an analog signal, and a light source. The lens unit of the camera 201 has a zoom function for adjusting the angle of view and an aperture function for adjusting the amount of light. The image sensor of the camera 201 has a gain function for adjusting the sensitivity when converting light into an analog signal. The light source of the camera 201 emits pulsed light and measures distance by measuring the time from when the light is emitted to when it hits an object and is reflected back in synchronization with the image sensor of the camera 201. These functions are adjusted based on setting values notified by the information processing device 203. The analog signal acquired by the camera 201 is converted into a digital signal by an analog-to-digital conversion circuit in the camera 201 and transferred to the information processing device 203 as an image signal.
[0020] The LiDAR 202, also known as a laser scanner or laser range finder, measures distances using the TOF method, just like the camera 201. However, by scanning the surroundings using a laser beam as a light source, it is possible to acquire distance information over a wide range. The LiDAR 202 includes a laser element and a light-receiving element that converts the light reflected from an object into an analog signal. There are various laser scanning methods, broadly categorized into rotary (mechanical) and non-rotary (solid-state) types. In this embodiment, a rotary 2D LiDAR is used, in which a laser beam scans a two-dimensional plane to obtain distance information distributed on the plane. The LiDAR 202's laser irradiation and scanning functions, as well as a gain function that adjusts the sensitivity when converting the light from the light-receiving element into an analog signal, are adjusted based on setting values notified by the information processing device 203. The analog signal acquired by the LiDAR 202 is converted into a digital signal by an analog-to-digital conversion circuit in the LiDAR 202 and transferred to the information processing device 203 as a distance signal.
[0021] The information processing device 203 has the functions of a general embedded PC device, and includes a CPU 207, a ROM 208, a RAM 209, a storage unit 210 such as an HDD or SSD, a general-purpose I / F 211 such as USB, and a system bus 212. CPU is an abbreviation for Central Processing Unit. ROM is an abbreviation for Read Only Memory. RAM is an abbreviation for Random Access Memory. HDD is an abbreviation for Hard Disk Drive. SSD is an abbreviation for Solid State Drive. USB is an abbreviation for Universal Serial Bus. I / F is an abbreviation for interface.
[0022] <Functional Configuration of Mobile System 110> The functional configuration of the mobile system according to the first embodiment of the present invention will be described below. The processing of each unit shown below is realized by software, in which a program is loaded from the ROM 208 or the like onto the RAM 209 and then executed by the CPU 207. Some or all of the functional configuration may be configured as hardware, such as an ASIC or FPGA. ASIC is an abbreviation for Application Specific Integrated Circuit. FPGA is an abbreviation for Field Programmable Gate Array.
[0023] 3 is a block diagram showing the functional configuration of the mobile system 110. The mobile system 110 includes an imaging control unit 301, a recognition unit 302, a storage unit 303, a self-position estimation unit 304, The system includes a plan determination unit 305 , a network communication unit 306 , and a control unit 307 .
[0024] The imaging control unit 301 acquires sensor information from the sensor unit 200 through the general-purpose I / F 211. In this embodiment, the imaging control unit 301 acquires image information and distance information from the camera 201, and acquires distance information from the LiDAR 202. In this embodiment, the sensor information is the image information and distance information from the camera 201, and the distance information from the LiDAR 202.
[0025] The recognition unit 302 performs predetermined recognition processing based on the sensor information acquired by the imaging control unit 301 from the sensor unit 200 and an environmental map (described later) read from the storage unit 303. Examples of recognition processing include detecting obstacles, people, and landmarks while driving, and performing recognition processing necessary to drive safely without collisions even if the environment changes dynamically.
[0026] The self-position estimation unit 304 calculates the self-position of the mobile system 110 at the time of acquiring the sensor information, based on the sensor information acquired by the imaging control unit 301 from the sensor unit 200 and an environmental map (described later) read from the storage unit 303. The self-position of the mobile system 110 is the position and attitude of the mobile system 110.
[0027] In this embodiment, MCL is used as an example of a self-localization technology using a known particle filter. MCL is an abbreviation for Monte Carlo Localization. In this embodiment, the technology described in Non-Patent Document 1 and the technology described in Non-Patent Document 2, in which the number of particles is variable, are applied and used.
[0028] When an environmental map does not yet exist, SLAM technology is used to create an environmental map using sensor information. SLAM stands for Simultaneous Localization and Mapping. Since the creation of an environmental map and self-localization are interdependent, SLAM simultaneously performs self-localization and map creation. SLAM takes sensor data as input and outputs the movement trajectory and map information of the mobile system.
[0029] In this embodiment, as an example, a SLAM technique using distance information obtained by 2D LiDAR is used, for example, by applying the technique described in Non-Patent Document 3. This is one of the techniques called sequential SLAM, which builds an environmental map each time sensor information is obtained.
[0030] The output of the self-position estimation unit 304 is sent to the plan determination unit 305 and used for driving control of the mobile system 110, or the self-position information is sent to a display device (not shown) and used for displaying the current position of the mobile system 110 on the UI. The UI includes GUI. UI is an abbreviation for User Interface. GUI is an abbreviation for Graphical User Interface.
[0031] It is also possible to transmit the sensor information acquired by the imaging control unit 301 to an external device via the network communication unit, and to acquire the position and orientation calculated by the external device by the self-position estimation unit 304 via the network communication unit 306.
[0032] The plan determination unit 305 plans and determines the next operation based on the self-position of the mobile system 110 acquired from the self-position estimation unit 304 and the state of the surrounding environment acquired from the recognition unit 302. For example, if an obstacle is detected, the plan determination unit 305 determines to plan and execute a local detour route. The plan determination unit 305 also makes decisions such as changing the cleaning mode (suction strength, movement speed, movement route, etc.) for each location. The output of the plan determination unit 305 is sent to the control unit 307. The control unit 307 receives the determination by the plan determination unit 305 and performs automatic driving control of the mobile system 110 based on the input.
[0033] <Environmental Map> The environmental map used in this embodiment will be described below. When performing self-location estimation using the LiDAR 202, for example, a point cloud map created based on distance information obtained by the LiDAR 202 can be used as the environmental map. Also, based on the point cloud map, space is divided into cells in a grid pattern, and an occupancy grid map can be used as the environmental map. The occupancy grid map is formed in a format that indicates the probability of whether each cell has been observed and whether a point cloud is included in each cell. In this embodiment, the occupancy grid map is used as the environmental map.
[0034] (How to obtain an environmental map) An environmental map is generated by moving a mobile system 110 equipped with a sensor within the environment. When an environmental map does not exist, a travel route for acquiring sensor data can be provided by, for example, operating the mobile system 110 with a remote control or by hand. The environmental map based on sensor information acquired by sensors equipped in the mobile system 110 can be generated simultaneously with the operation or offline after data collection. Alternatively, the mobile system 110 may autonomously travel within the environment, planning a travel route exploratory using maps generated up to that point and local maps around the current location.
[0035] <Workflow for mobile systems> The workflow of the mobile system 110 when performing the work will be described using the flowchart in Fig. 4. Fig. 4 is a diagram illustrating the workflow of the mobile system 110.
[0036] First, in step S401, the self-position estimation unit 304 reads a previously created environmental map from the storage unit 303. In step S402, the self-position estimation unit 304 calculates the self-position of the mobile system 110 from the environmental map read in step S401 and local information about the current location from the imaging control unit 301.
[0037] Next, in step S403, the plan determination unit 305 performs a route plan from the self-position calculated in step S402 to the work start point. In step S404, the mobile system 110 moves to the work start point while estimating its own position on the environmental map.
[0038] Next, in step S405, the plan determination unit 305 creates a route plan from the work start point to the work end point. If a task that involves working evenly within an area, such as cleaning, is assigned, the plan determination unit 305 creates a route plan for that task. Alternatively, if a route plan has been created in advance and stored in the storage unit 303, the plan determination unit 305 reads the route plan from the storage unit 303.
[0039] Next, in step S406, the mobile system 110 travels according to the route plan. At this time, it performs a predetermined task as needed. The predetermined task is a suction or wiping operation for a cleaning robot, a grass cutting operation for a lawnmower, or a tilling operation for a cultivator.
[0040] When the work end point is reached, in step S407, the plan determination unit 305 performs a route plan for returning from the current value to a predetermined position such as the travel start point. In step S408, the mobile system 110 travels to the predetermined position according to the route plan, and notifies the server 130 via the network communication unit 306 that the work has been completed.
[0041] If the mobile system 110 detects an obstacle not on the environmental map while traveling, it replans a local route to avoid the obstacle and continues traveling along that route. If the mobile system 110 cannot avoid the obstacle, it stops, indicates that an abnormality exists using the indicator light of the mobile system 110, notifies the user of a message via the network communication unit 306, and waits for an operation from the administrator.
[0042] The cleaning work time, travel route, or areas where work could not be performed due to obstacle avoidance is stored as a post-work report in the storage unit 303. Alternatively, the post-work report is sent to the server 130 via the network communication unit 306.
[0043] (Server hardware configuration) 5 is a diagram showing an example of the hardware configuration of the server 130. The server 130 is configured as a computer such as a general PC. The server 130 includes a CPU 501 which is a processor, RAM 502 and ROM 503 which are memories, a storage unit 504 such as an HDD, and a communication unit 505. The server 130 realizes various functions by the CPU 501 executing programs stored in the ROM 503, the storage unit 504, or an external storage device.
[0044] (Server functional configuration) 6 is a diagram showing an example of the functional configuration of the server 130. The server 130 includes a network communication unit 601, a control unit 602, a display unit 603, an operation unit 604, an analysis unit 605, a storage unit 606, and a setting processing unit 607.
[0045] The network communication unit 601 connects to, for example, the network 120 and communicates with the mobile system 110 and the like via the network 120. Note that this is merely an example, and the network communication unit 601 may be configured to establish a direct connection with the mobile system 110 and communicate with the mobile system 110 without going through the network 120 or other devices.
[0046] The control unit 602 controls the network communication unit 601, the display unit 603, the operation unit 604, the analysis unit 605, the storage unit 606, and the setting processing unit 607 so that they perform their respective processes.
[0047] The display unit 603 presents information to the user via, for example, a display. In this embodiment, the information is presented to the user by displaying the results of rendering by the browser on the display. Note that the information may also be presented by a method other than a screen display, such as sound or vibration.
[0048] The operation unit 604 accepts operations from the user. In this embodiment, the operation unit 604 is a mouse or a keyboard. The user operates the operation unit 604 to input user operations into the browser. However, the present invention is not limited to this, and the operation unit 604 may be any device capable of detecting the intention of other users, such as a touch panel or a microphone.
[0049] The analysis unit 605 performs the analysis work required for the mobile system 110 to perform a predetermined work. A method for setting partial areas that are no-entry areas in the environmental map based on sensor information will be described later.
[0050] The storage unit 606 stores information necessary for the mobile system 110 to perform predetermined operations. The setting processing unit 607 executes the setting process described below.
[0051] <Flow of setting no-entry areas for mobile systems> The flow of setting no-entry areas for a mobile system is described below. Each of the following processes is executed by the information processing device 203 of the mobile system 110. However, this is merely an example, and some or all of the processes described below may be implemented not only by the mobile system 110, but also by the server 130, a laptop PC or tablet terminal capable of communicating with the mobile system 110, or dedicated hardware.
[0052] A method for the mobile body system 110 to recognize no-entry areas while traveling and reflect the recognition results in a route plan will be described using the flowchart in Fig. 7. Fig. 7 is a diagram illustrating the no-entry area creation processing flow according to the first embodiment.
[0053] First, in step S701, the mobile system 110 starts autonomous traveling to execute a given task. The task here is, for example, transporting luggage or cleaning. If the task is transporting luggage, the mobile system 110 plans a route to travel from its current location to the destination without coming into contact with any obstacles and return to a predetermined location such as a waiting area, and travels according to the route plan. If the task is cleaning, the mobile system 110 plans a route to travel thoroughly through the cleaning area and return to a predetermined location such as a waiting area, and travels according to the route plan. At this time, the mobile system 110 acquires image information and distance information of the surrounding environment while traveling.
[0054] Next, in step S702, the mobile system 110 determines whether it has arrived at the destination (a predetermined location such as a waiting area). If it is determined that the mobile system 110 has arrived at the destination, it stops self-propelled and ends. If it is determined that the mobile system 110 has not arrived at the destination, the process proceeds to step S703.
[0055] In step S703, the mobile system 110 performs a process of segmenting the acquired image into no-entry areas using a semantic area segmentation model trained by a technique called machine learning, which finds hidden patterns in large amounts of data through repeated calculations.
[0056] Semantic segmentation is a task of dividing an image into subregions by inferring a category of object for each pixel. Semantic segmentation is used, for example, to divide acquired images into various categories so that a mobile system 110 operating in an indoor environment can understand the environment and act accordingly. The categories into which the acquired images are divided include, for example, people, walls, fences, floors, ceilings, pillars, chairs, desks, shelves, trash cans, signs, escalators, and elevators.
[0057] In particular, since models using deep neural networks have become capable of performing inference with high accuracy in recent years, this embodiment also uses a deep neural network model, and applies the technology described in Non-Patent Document 4, for example.
[0058] In this embodiment, the purpose is to set no-entry areas in a commercial facility. For this purpose, this area division model is used to divide the floor surface of the commercial facility from an image into no-entry areas that are mainly tenant areas and the like, and allowed areas that are mainly common areas such as corridors.
[0059] For convenience, the following description will be given assuming that tenant areas are prohibited areas and common areas are permitted areas. In reality, areas managed by the building may be categorized as prohibited areas in the same way as tenant areas. Conversely, tenant areas may be categorized as permitted areas in the same way as common areas managed by the building. For example, an area near an escalator entrance may not be a tenant area, but it can be categorized as a prohibited area. This is determined by the annotation rules of the correct answer data used for learning.
[0060] In many cases, the tenant areas and the common areas are clearly indicated by the color and texture of the floor. Therefore, at first glance, it may seem possible to separate the two areas if the area can be divided based on color and texture. However, in reality, the designs of commercial facilities are diverse, so it is not possible to completely distinguish between the two areas based on color and texture alone.
[0061] FIG. 8 is a diagram illustrating an example of creating a no-entry area according to the first embodiment. FIG. 8 illustrates an example of an image captured by the camera 201 of the mobile system 110. An intermediate area 803 with a different color and texture is included between a tenant area (no-entry area) 801 and a common area (permitted area) 802. This intermediate area 803 is highly context-dependent, and whether it is included in the tenant area or the common area cannot be determined uniformly from the image alone, depending on the design of the facility. A uniform determination would result in a high probability of an incorrect determination. Furthermore, in the example of FIG. 8, there is an area (intermediate area 803) between the tenant area (no-entry area) 801 and the common area (permitted area) 802, where it is difficult to determine whether it is included in the tenant area or the common area. However, areas where it is difficult to determine exist outside the area between the tenant area (no-entry area) and the common area (permitted area). The semantic area segmentation model outputs an area where it is difficult to determine whether it is a tenant area (no-entry area) or a shared area (permitted area) as an area where the determination is uncertain. Hereinafter, for convenience, an area where it is difficult to determine as a semantic area segmentation model is referred to as an "intermediate area." However, as mentioned above, an area where it is difficult to determine may exist in areas other than the area sandwiched between the tenant area (no-entry area) and the shared area (permitted area).
[0062] Therefore, in this embodiment, the semantic area division model is trained to output division into at least three categories: a no-entry area (tenant area), an entry-permitted area (shared area), and an intermediate area (ambiguous area). The intermediate area 803 is classified as an intermediate area (ambiguous area).
[0063] It is known that the ambiguity of the correct labels in the training data degrades the inference accuracy of the model. If we try to forcefully assign areas to either no-entry or permitted areas, the correct category will fluctuate from image to image, resulting in a loss of consistency in the correct data. By assigning new labels to intermediate areas that could be either type depending on the facility, we can reduce the ambiguity and improve the inference accuracy of the semantic region segmentation model.
[0064] Next, in step S704, the mobile system 110 integrates the area division information into a three-dimensional data structure and stores it in the storage unit 606. For environmental understanding, the mobile system 110 divides the three-dimensional space into voxels and stores in its memory a data structure that allows each voxel to be selectively placed in one of three states: occupied, unoccupied, or unmeasured. Furthermore, for those voxels that are in an occupied state, the area category can be stored as an attribute of the voxel in addition to color, normal, etc.
[0065] The mobile system 110 handles the surrounding environment as input using information from the distance imaging camera 201, but updates information based on the recognition results for voxels that correspond to the viewing frustum determined by the distance measurement results and distance measurement range of the distance imaging camera 201. However, values of voxels that are further back than the distance measurement results, even if they are within the viewing frustum, are not updated. The updating method may involve replacing the voxels with new ones each time a measurement is made, or taking a weighted average. In this way, the results of region segmentation are added and updated to the data structure that handles information on three-dimensional space.
[0066] In this embodiment, the view frustum obtained from the image information and distance information is applied to a position in a map coordinate system using the current position and orientation of the mobile system 110, and the data is stored in the data structure of the above-mentioned three-dimensional space. The map coordinate system is a world coordinate system.
[0067] Next, in step S705, the mobile body system 110 sets the no-entry area in a format that can be used for route planning. In this embodiment, the mobile body system 110 performs route planning based on an occupancy grid map (cost map) that divides space into a grid based on sensor information and indicates the probability of obstacle presence for each grid. To this end, the mobile body system 110 projects the 3D spatial data integrated in step S704 onto a 2D plane and sets the grids of the occupancy grid map (cost map) at positions that correspond to the no-entry area to an occupied state.
[0068] Furthermore, whether or not a grid in the occupancy grid map at a position corresponding to an intermediate area is to be occupied is determined by the user's settings. It is often determined whether an intermediate area is included in a tenant area or a common area for each facility or floor. Therefore, whether an intermediate area is included in a no-entry area or an allowed area can be determined according to the settings of a user who knows the facility's design rules. In this way, the no-entry areas of the traveling system 110 can be automatically set for each facility, floor, or partial area, etc.
[0069] The mobile system 110 can create a route plan that does not pass through the no-entry areas by passing the occupancy grid map thus obtained to a module that creates a route plan. Note that the no-entry areas that are automatically set in this way can also be modified by the user as needed using a GUI or the like. Once the no-entry areas have been set, the mobile system 110 continues driving towards the destination.
[0070] [Second embodiment] The second embodiment of the present invention has the same hardware configuration, functional block diagram, and the like as the first embodiment except for the flow of setting no-entry areas in a mobile system, so only the flow of setting no-entry areas in a mobile system will be described.
[0071] <Flow of setting no-entry areas for mobile systems> A method for recognizing a no-entry area while the mobile body system 110 is traveling and reflecting the recognition result in a route plan will be described below with reference to the flowchart in Fig. 9. Fig. 9 is a diagram illustrating the no-entry area creation processing flow according to the second embodiment.
[0072] The processes in steps S701 and S702 are the same as those in the first embodiment, and therefore the description thereof will be omitted.
[0073] In step S903, the mobile system 110 performs processing to divide the acquired image into no-entry areas. Fig. 10 is a diagram showing an example of creating no-entry areas according to the second embodiment. Fig. 10 is an example of an image captured by the camera 201 of the mobile system 110.
[0074] Tenant area 1001 includes a hallway 1004. The color and texture of hallway 1004 may be the same as or different from the rest of tenant area 1001.
[0075] As described above, segmentation is particularly difficult when the tenant area 1001 of a commercial facility includes additional structures such as corridors, or when the color or texture of the floor of the common area is continuous and directly connected to the tenant area. In such cases, the method of the first embodiment alone may not be sufficient for segmentation. Therefore, in the second embodiment, information about hanging walls is used to segment the area.
[0076] That is, the semantic area segmentation model used in the second embodiment is trained to segment a no-entry area 1001 for tenants, a permitted area 1002 for common areas, an intermediate area 1003 between these areas, and a hanging wall area 1005. A hanging wall is also called a hanging wall. A hanging wall is a wall that hangs down from the ceiling by several tens of centimeters to several meters and has the function of dividing a space.
[0077] Next, in step S904, the mobile system 110 integrates the area division information into a three-dimensional data structure and stores it in the storage unit 606, similar to the process in step S704.
[0078] Next, in step S905, the mobile system 110 performs area division using information obtained by projecting the three-dimensional spatial information onto a two-dimensional plane.
[0079] Fig. 11 is a diagram showing an example of an overhead view of a no-entry area. Fig. 11 is a diagram showing an example of a no-entry area 1101, a permitted area 1102, an intermediate area 1103, and an unmeasured area 1104 projected onto a two-dimensional plane from three-dimensional spatial information. In the example of Fig. 11, because there is a passageway within the tenant, the area is mistakenly divided into areas where the permitted area exists within the tenant.
[0080] Therefore, in this embodiment, region division is performed using a hanging wall region 1005 as shown in Fig. 10. Fig. 12 is a diagram showing an example of an overhead view of a no-entry region according to the second embodiment. In this embodiment, a dividing line 1201 is drawn at a position projected from the three-dimensional position of the hanging wall region 1005 onto the floor surface in the state shown in Fig. 11. In this embodiment, the entry-permitted region 1102 is divided by this dividing line 1201, and the region behind the dividing line 1201 as seen from the camera position is labeled as a no-entry region 1202. For example, the user can set, while checking on the display unit 603, whether the region divided by the dividing line 1201, i.e., the hanging wall region 1005, is an entry-permitted region or a no-entry region.
[0081] Next, in step S906, the mobile system 110 sets the no-entry area in a format that can be used for route planning, similar to the process in step S705.
[0082] According to the second embodiment, even in cases where both the passage 1004 and the common area 1002 in the tenant area are mistakenly divided as entry-permitted areas, it is possible to correctly set the entry-prohibited areas.
[0083] Whether or not the above-described area division by the hanging wall is enabled can be determined by the user's settings.
[0084] [Another embodiment] Although the present invention has been described with an example in which images captured by the mobile system 110 are analyzed and processed by the mobile system 110, the images captured by the mobile system 110 may also be analyzed and processed by the server 130.
[0085] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0086] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments and various modifications and changes are possible within the scope of the gist of the present invention.
[0087] The disclosure of this embodiment includes the following configurations: a method and a program. (Configuration 1) A means for acquiring an image of an environment in which the moving object moves; a means for inputting the acquired image into a semantic region segmentation model and acquiring at least three categories of regions in the acquired image, including no-entry regions, permitted regions, and regions that cannot be determined as no-entry regions or permitted regions, as outputs of the semantic region segmentation model; An information processing device comprising: (Configuration 2) 2. The information processing device according to configuration 1, further comprising a setting unit for setting whether the area that cannot be determined as an entry-prohibited area or an entry-permitted area is an entry-prohibited area or an entry-permitted area. (Configuration 3) The semantic area segmentation model outputs the area segmented into at least four categories: a no-entry area, an allowed area, an area in which it is not clear whether the area is a no-entry area or an allowed area, and a hanging wall area. 3. The information processing device according to configuration 1 or 2. (Configuration 4) 4. The information processing device according to configuration 3, further comprising a setting unit for setting whether the area partitioned by the hanging wall area is an entry-prohibited area or an entry-permitted area. (Method 1) acquiring an image of an environment in which the moving object moves; inputting the acquired image into a semantic region segmentation model, and acquiring at least three categories of regions in the acquired image, including no-entry regions, permitted regions, and regions that cannot be determined as no-entry regions or permitted regions, as outputs of the semantic region segmentation model; A method comprising: (Program 1) Computer, A means for acquiring an image of the environment in which the moving object moves; and a means for inputting the acquired image into a semantic region segmentation model, and acquiring at least three categories of regions in the acquired image, including no-entry regions for the moving object, permitted regions, and regions that cannot be determined as either no-entry regions or permitted regions, as outputs of the semantic region segmentation model; A program characterized by functioning as [Explanation of symbols]
[0088] 101 Mobile Management System 110 Mobile Systems 120 Network 130 servers
Claims
1. a means for acquiring an image of an environment in a commercial facility in which the mobile object moves; a means for inputting the acquired image into a semantic area segmentation model that determines the area to be worked on, and acquiring at least three categories as an output of the semantic area segmentation model: a no-entry area that prohibits entry of the moving object and is made up of a tenant area that is not the object of work in the commercial facility; an allowed area that permits entry of the moving object and is made up of a common area that is the object of work in the commercial facility; and an intermediate area that is located between the no-entry area and the allowed area and is made up of an area that is difficult to determine as either the tenant area or the common area; An information processing device comprising:
2. 2. The information processing apparatus according to claim 1, further comprising a setting unit for setting the intermediate area as an area into which the moving object is prohibited from entering or an area into which the moving object is permitted to enter.
3. The semantic area segmentation model outputs the area segmented into at least four categories: the no-entry area, the permitted area, the intermediate area, and an area partitioned at the top by a hanging wall.
2. The information processing apparatus according to claim 1, wherein:
4. 4. The information processing device according to claim 3, further comprising means for setting whether the area partitioned by the hanging wall is an area into which the moving object is prohibited from entering or an area into which the moving object is permitted to enter.
5. 2. The information processing apparatus according to claim 1, further comprising a planning unit for planning a movement route of said mobile object so as to avoid said no-entry area and said intermediate area.
6. 4. The information processing apparatus according to claim 3, wherein the moving path of the moving object is planned so as to avoid the area partitioned at the top by a hanging wall.
7. A method executed by an information processing device, comprising: acquiring an image of an environment in a commercial facility in which the mobile object moves; inputting the acquired image into a semantic area segmentation model that determines the area to be worked on, and acquiring at least three categories as an output of the semantic area segmentation model: a no-entry area that prohibits entry of the moving object and is made up of a tenant area that is not the object of work in the commercial facility; an allowed area that permits entry of the moving object and is made up of a common area that is the object of work in the commercial facility; and an intermediate area that is located between the no-entry area and the allowed area and is made up of an area that is difficult to determine as either the tenant area or the common area; A method comprising:
8. Computer, A means for acquiring images of the environment in the commercial facility where the mobile object moves; and A program characterized by functioning as a means for inputting the acquired image into a semantic area division model that determines the area to be worked on, and obtaining at least three categories as output of the semantic area division model: a no-entry area that prohibits entry of the moving object and consists of a tenant area that is not the target of work in the commercial facility; an allowed area that allows entry of the moving object and consists of a common area that is the target of work in the commercial facility; and an intermediate area that is located halfway between the no-entry area and the allowed area and consists of an area that is difficult to determine as either a tenant area or a common area.
Citation Information
Patent Citations
Deep confined space unmanned transportation equipment travelable region detection and autonomous obstacle avoidance method
CN115100622A
Robot full-coverage operation method and device and robot
CN116149314A
Circumferential environment recognition device, autonomous mobile system using the same, and circumferential environment recognition method
JP2014194729A
Information processor and mobile apparatus
JP2017211909A
Prioritizing cleaning areas
JP2017502371A