Moving body control device, moving body control system, moving body control method, and storage medium

The control device enhances self-position estimation in autonomous vehicles by excluding regions based on object type and confidence levels, addressing processing load issues and improving navigation accuracy.

WO2026069654A1PCT designated stage Publication Date: 2026-04-02HONDA MOTOR CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Conventional autonomous driving technologies face increased processing load due to multiple moving and stationary objects in images, which can hinder accurate self-position estimation of vehicles, neglecting the impact on surrounding traffic participants.

Method used

A control device for a moving object that acquires images, recognizes objects, and estimates self-position by excluding regions based on object type and confidence levels, adjusting the exclusion criteria to enhance estimation accuracy.

Benefits of technology

Reduces processing load and improves self-position estimation by considering surrounding conditions, contributing to more accurate navigation and safer autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024034923_02042026_PF_FP_ABST
    Figure JP2024034923_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A moving body control device according to the present invention includes an acquisition unit that acquires an image obtained by imaging situations in a vicinity of a moving body, a recognition unit that recognizes an object that is present in the vicinity of the moving body on the basis of the image, an own position estimation unit that estimates the own position of the moving body on the basis of the image, and a movement control unit that performs movement control of the moving body on the basis of an estimation result of the own position. The recognition unit recognizes at least a first object that is present in the vicinity of the moving body and that is moving, as an object. The own position estimation unit sets regions of objects to be excluded from a region in the image when estimating the own position, on the basis of type of each object and reliability representing likelihood, and estimates the own position on the basis of information regarding a region in the image that is not excluded.
Need to check novelty before this filing date? Find Prior Art

Description

Control device for a moving body, mobile control system, control method for a moving body, and storage medium

[0001] The present invention relates to a control device for a moving body, a mobile control system, a control method for a moving body, and a storage medium.

[0002] In recent years, efforts have been actively made to provide access to a sustainable transportation system that takes into account people in vulnerable positions among traffic participants. In order to achieve this, research and development focused on further improving traffic safety and convenience through research and development related to autonomous driving technology.

[0003] Conventionally, the practical application of a moving body that moves with a user while maintaining a preset fixed relative positional relationship, such as in front of or behind the user, has been advanced. In such a moving body, it is necessary to accurately grasp the current position (self-position) and safely guide (navigate) or follow the user to the destination according to the situation of the movement path where the user and the moving body itself will move in the future.

[0004] In relation to this, conventionally, a technique for estimating the self-position of a moving body based on the result of comparing a plurality of captured images taken in a predetermined direction at different positions with a reference image captured in advance has been disclosed (see, for example, Patent Document 1). Furthermore, conventionally, a technique related to a self-position estimation device that estimates the position of a moving body using information acquired by a sensor unit, acquires a reliability regarding the estimation result, and acquires the self-position of the moving body according to the estimation result corresponding to the highest reliability among a plurality of reliabilities respectively acquired by a plurality of position estimation units has been disclosed (see, for example, Patent Document 2).

[0005] WO 2019 / 073795 JP 2021-018638 A

[0006] Incidentally, in autonomous driving technology, a challenge is that images taken to estimate the vehicle's own position often include multiple other moving objects, such as cars, bicycles, and pedestrians, both moving and stationary. This increases the processing load on the vehicle when estimating its own position using these images. However, conventional technologies have not considered the increased processing load on other traffic participants present along the vehicle's future travel path. For example, when there are many other traffic participants in the image, the vehicle may become unable to grasp its surroundings or accurately estimate its own position.

[0007] This invention was made based on the above-mentioned problem recognition and aims to provide a control device for a moving object, a moving object control system, a moving object control method, and a storage medium that can more suitably estimate the self-position of a moving object according to the surrounding conditions of the moving object. In other words, this invention aims to reduce the processing load for estimating the self-position of a moving object by considering other traffic participants included in images captured around the moving object, and to achieve more suitable self-position estimation. Ultimately, this will contribute to the development of sustainable transportation systems.

[0008] The control device for a mobile body, a mobile body control system, a method for controlling a mobile body, and a storage medium according to this invention employ the following configuration: (1) A control device for a mobile body according to one aspect of this invention comprises: an acquisition unit that acquires an image of the surrounding conditions of a mobile body; a recognition unit that recognizes objects present in the vicinity of the mobile body based on the image; a self-position estimation unit that estimates the self-position of the mobile body based on the image; and a movement control unit that controls the movement of the mobile body based on the self-position estimation result, wherein the recognition unit recognizes at least a first moving object present in the vicinity of the mobile body as the object, and the self-position estimation unit sets a region of the object to be excluded from the region in the image when estimating the self-position based on the type of each object and a confidence level representing the likelihood of its existence, and estimates the self-position based on the information of the region in the image that is not excluded.

[0009] (2) In the embodiment of (1) above, the recognition unit further recognizes the second object, which is the first object but is currently stationary, as the object.

[0010] (3) In the embodiment of (1) above, the self-position estimation unit first sets the region of all the first objects as the region of the objects to be excluded, and then estimates its own position.

[0011] (4) In the embodiment of (2) above, if the self-position estimation unit cannot estimate its own position, it sets the region of the object whose reliability is equal to or greater than a predetermined threshold as the region of the object to be excluded, changes the degree to which the region of the object is excluded from the region in the image, and estimates its own position.

[0012] (5) In the embodiment of (4) above, the self-position estimation unit estimates the self-position by gradually changing the threshold when the self-position cannot be estimated.

[0013] (6) In the embodiment of (5) above, the self-position estimation unit estimates the self-position by lowering or increasing the degree to which it excludes the region of the object when the self-position cannot be estimated.

[0014] (7) In any one embodiment of (4) to (6) above, the self-position estimation unit stores the threshold value when it is able to estimate the self-position.

[0015] (8) In the embodiment of (7) above, when the self-position estimation unit estimates the next self-position, it first sets the regions of the objects that are greater than or equal to the stored threshold as the regions of the objects to be excluded, instead of all the regions of the objects, and then estimates the self-position.

[0016] (9): A mobile body control system according to one aspect of the present invention comprises: an acquisition unit that acquires an image of the surrounding conditions of a mobile body; a recognition unit that recognizes objects present in the surrounding area of ​​the mobile body based on the image; and a self-position estimation unit that estimates the self-position of the mobile body based on the image, wherein the recognition unit recognizes at least a first moving object present in the surrounding area of ​​the mobile body as the object; and the self-position estimation unit sets an area of ​​the object to be excluded from the area in the image when estimating the self-position based on the type of each object and a confidence level representing the likelihood of such an object, and estimates the self-position based on the information of the area in the image that is not excluded.

[0017] (10): A method for controlling a moving body according to one aspect of the present invention is a method for controlling a moving body in which a computer acquires an image of the surrounding environment of the moving body, recognizes objects present in the surrounding environment of the moving body based on the image, estimates the self-position of the moving body based on the image, controls the movement of the moving body based on the estimation result of the self-position, recognizes at least a first moving object present in the surrounding environment of the moving body as the object, sets an area of ​​the object to be excluded from the area in the image based on the type of each object and a confidence level representing the likelihood of each object, and estimates the self-position based on the information of the area in the image that has not been excluded.

[0018] (11): A storage medium according to one aspect of the present invention is a storage medium that stores a program which causes a computer to acquire an image of the surroundings of a moving object, recognize objects present in the vicinity of the moving object based on the image, estimate the self-position of the moving object based on the image, control the movement of the moving object based on the estimation result of the self-position, recognize at least a first moving object present in the vicinity of the moving object as the object when recognizing the object, set an area of ​​the object to be excluded from the area in the image based on the type of each object and a confidence level representing the likelihood of each object, and estimate the self-position based on the information of the area in the image that has not been excluded.

[0019] According to the embodiments described in (1) to (11) above, the self-position can be more preferably estimated depending on the surrounding conditions of the moving object.

[0020] This figure shows an example of the configuration of a mobile body control system including a mobile body according to the embodiment. This is a perspective view showing an example of the external configuration of the mobile body according to the embodiment. This figure shows an example of the functional configuration of the mobile body according to the embodiment. This figure is for explaining the method of estimating the self-position in the control device of the mobile body according to the embodiment. This is a flowchart showing an example of the flow of the self-position estimation process performed in the control device of the mobile body according to the embodiment. This figure shows an example of self-position estimation in the control device of the mobile body according to the embodiment.

[0021] Hereinafter, embodiments of the control device for a mobile body, a mobile body control system, a mobile body control method, and a storage medium of the present invention will be described with reference to the drawings.

[0022] [Configuration of Mobile Control System] Figure 1 is a diagram showing an example of the configuration of a mobile control system including a mobile body according to an embodiment. The mobile control system 1 includes, for example, a terminal device 2, a management device 10, an information providing device 20, and a mobile body 100. These components communicate via a network NW or the like. The network NW is any network such as a LAN (Local Area Network), WAN (Wide Area Network), or Internet connection.

[0023] [Terminal device] Terminal device 2 is a computer device such as a smartphone or a tablet. Terminal device 2 is used by a user of the mobile control system 1 and requests the use of the mobile device 100 from the management device 10 based on the user's operation, and obtains information from the management device 10 indicating that the use of the mobile device 100 has been permitted.

[0024] [Management Device] The management device 10 manages the usage status and usage reservations of mobile devices 100 within the mobile device control system 1. In response to requests received from the terminal device 2, the management device 10 sets usage rights for mobile devices 100 that are available to the user, and provides the user with information indicating that the use of the set mobile device 100 has been permitted by sending it to the terminal device 2. For example, the management device 10 generates and manages schedule information that associates pre-registered user identification information with the date and time of the usage reservation for the mobile device 100.

[0025] [Information Provisioning Device] The information provisioning device 20 provides the mobile body 100 and the terminal device 2 with information such as the location where the mobile body 100 is located, the area in which the mobile body 100 is moving, and map information of the area surrounding the area. The information provisioning device 20 may also generate a route from the current location of the mobile body 100 to its destination in response to a request from the mobile body 100, and provide the generated route to the mobile body 100. The management device 10 and the information provisioning device 20 may be implemented by, for example, a server device, or they may be devices configured by cloud computing consisting of one or more information processing devices.

[0026] [Mobile Entity] The mobile entity 100 is, for example, a mobile entity capable of autonomous movement. Autonomous movement means moving the mobile entity 100 by performing either or both of the speed control or turning control of the mobile entity 100 without relying on driving operations by a user. Turning control includes, for example, changing the direction of the mobile entity 100 by rotation or turning, or steering control if it has a steering wheel. The mobile entity 100 is, for example, a vehicle, but may also include other mobile entities capable of autonomous movement (e.g., walking robots). Vehicles include not only four-wheeled vehicles, but also all vehicles capable of movement with three or two wheels. The mobile entity 100 has a structure that can carry and transport objects such as luggage. The above objects may include people such as users. The mobile entity 100 may be capable of traveling on roadways or in predetermined areas other than roadways (e.g., sidewalks, inside buildings, public open spaces, etc.).

[0027] The mobile unit 100 is used by users based on usage rights set by the management device 10, for example. For example, a user can load luggage or other items onto the mobile unit 100 and have it follow the user's movements, lead the user to a destination, or travel alongside the user, according to the movement mode instructed by the user. The location information, usage status, and usage reservation status of the mobile unit 100 are managed by the management device 10. The mobile unit 100 acquires information from the management device 10 and the information providing device 20 and performs movement control based on usage rights and the provided information.

[0028] Figure 1 shows the configuration of a mobile control system 1 that includes one terminal device 2, one management device 10, one information providing device 20, and one mobile body 100. However, the mobile control system 1 of this embodiment may have multiple terminal devices 2, management devices 10, information providing devices 20, and mobile bodies 100, or at least one of them. In the mobile control system 1 of this embodiment, the management device 10 and the information providing device 20 may be integrated. The mobile control system 1 may also have a configuration that does not include a management device 10. In this case, at least a part of the functions of the management device 10 is provided on the mobile body 100, and the terminal device 2 communicates with the mobile body 100 via a network NW to manage things like user rights.

[0029] [External Configuration of the Mobile Body] Figure 2 is a perspective view showing an example of the external configuration of the mobile body 100 according to the embodiment. In the following description, the forward direction of the mobile body 100 is the positive X direction, the rear direction of the mobile body 100 is the negative X direction, the width direction of the mobile body 100 is the positive Y direction to the left and the negative Y direction to the right with respect to the positive X direction, and the height direction of the mobile body 100, which is the direction perpendicular to the X and Y directions, is the positive Z direction.

[0030] The mobile body 100 comprises, for example, a base 110, a door 112 provided on the base 110, and wheels (first wheel 120, second wheel 130, and third wheel 140) mounted on the base 110. For example, a user can open the openable door 112 to put luggage into a storage compartment provided on the base 110, or take luggage out of the storage compartment. The first wheel 120 and the second wheel 130 are drive wheels and rotate by power from a motor or the like. The third wheel 140 is an auxiliary wheel (driven wheel). The mobile body 100 may also be movable using configurations other than wheels, such as tracks.

[0031] A cylindrical support 150 extending in the positive Z direction is provided on the surface of the base body 110 in the positive Z direction. A camera 180 for imaging the area around the mobile body 100 is provided at the end of the support 150 in the positive Z direction.

[0032] Camera 180 is a digital camera that utilizes a solid-state image sensor such as a CCD (Charge Coupled Device) or CMOS (Complementary Metal Oxide Semiconductor). The position in which camera 180 is installed may be any position different from the above. Camera 180 periodically and repeatedly images the area around the moving object 100 (at least in front of it) at predetermined time intervals, for example. Camera 180 may be a stereo camera or a camera capable of imaging the area around the moving object 100 in a wide angle (for example, 360 degrees). Camera 180 may be composed of multiple cameras that image the front, rear, and sides of the moving object 100, respectively, to image the area around the moving object 100 in a wide angle. Camera 180 may be implemented by combining multiple 120-degree cameras or multiple 60-degree cameras, for example.

[0033] The configuration of the mobile body 100 shown in Figure 2 is merely an example, and other configurations may be added, some configurations (for example, some configurations that are not essential for realizing the functions of the present invention, such as the door portion 112) may be omitted, or other configurations may be added. For example, the mobile body 100 may be equipped with a radar device or a detection device (sensor) other than the camera 180, such as a LIDAR (Light Detection and Ranging), in order to detect objects present in the vicinity of the mobile body 100. Furthermore, the size, shape, and arrangement of each configuration in the mobile body 100 shown in Figure 2 are not limited to the example shown in Figure 2.

[0034] [Functional Configuration of the Mobile Body] Figure 3 is a diagram showing an example of the functional configuration of the mobile body 100 according to the embodiment. In addition to the configuration shown in Figure 2, the mobile body 100 includes, for example, a first motor 122, a second motor 132, a battery 134, a brake device 136, a steering device 138, a communication unit 190, and a control device 200. The first motor 122 and the second motor 132 are powered by electricity supplied from the battery 134. The first motor 122 drives the first wheel 120. The first motor 122 may be an in-wheel motor provided on the wheel of the first wheel 120. The second motor 132 drives the second wheel 130. The second motor 132 may be an in-wheel motor provided on the wheel of the second wheel 130.

[0035] The braking device 136 outputs brake torque to each wheel based on instructions from the control device 200. The steering device 138 includes an electric motor. The electric motor, for example, applies force to the rack and pinion mechanism based on instructions from the control device 200 to change the direction of the first wheel 120 or the second wheel 130, thereby changing the path of the moving body 100.

[0036] The communication unit 190 is a communication interface for sending and receiving various types of information by communicating with the terminal device 2, the management device 10, and / or the information providing device 20 via the network NW. The communication unit 190 includes, for example, a network card or a NIC (Network Interface Controller). The communication unit 190 may also communicate with other mobile devices.

[0037] The control device 200 controls the overall operation of the mobile body 100. The control device 200 is housed within the base body 110. The control device 200 includes, for example, an acquisition unit 202, an information processing unit 204, a recognition unit 206, a self-position estimation unit 208, a trajectory generation unit 210, a movement control unit 212, and a storage unit 220.

[0038] Each of the acquisition unit 202, information processing unit 204, recognition unit 206, self-position estimation unit 208, trajectory generation unit 210, and movement control unit 212 is realized, for example, by a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be realized by hardware (including circuitry) such as LSI (Large Scale Integration), SOC (System On Chip), Application Specific Integrated Circuit (ASIC), programmable logic device (e.g., Simple Programmable Logic Device (SPLD) or Complex Programmable Logic Device (CPLD), Field Programmable Gate Array (FPGA)), and GPU (Graphics Processing Unit), or by the cooperation of software and hardware. Some or all of these components may be realized by a dedicated LSI. The program may be stored in advance in a storage device such as a semiconductor memory element like ROM (Read Only Memory), RAM (Random Access Memory), or flash memory, or in a storage device with a non-transient storage medium such as a hard disk drive (HDD), or it may be stored in a removable storage medium such as a DVD or CD-ROM (a non-transient storage medium) and installed in the storage device when the storage medium is mounted in a drive device. Some or all of the functional configuration included in the control device 200 may be included in other devices. For example, the mobile device 100 may communicate with other devices (such as a server device) and cooperate to control the mobile device 100.

[0039] Some or all of the functions of each component of the acquisition unit 202, information processing unit 204, recognition unit 206, self-position estimation unit 208, and trajectory generation unit 210 may be implemented by, for example, a server device, and perform control equivalent to that of the control device 200 by communicating with the communication unit 190 of the mobile body 100 via a network NW.

[0040] The storage unit 220 is implemented by, for example, semiconductor memory elements such as ROM, RAM, and flash memory, or a storage device such as a hard disk drive (HDD) (a storage device equipped with a non-transient storage medium). The storage unit 220 stores control information 222, which includes a control program for controlling the operation of the mobile body 100 (for example, operation in a mobile mode), and map information 224, which are referenced by the movement control unit 212.

[0041] Map information 224 is, for example, map information such as the location of the mobile body 100 provided by the information providing device 20, the area in which the mobile body 100 moves, and the area surrounding the area. Map information 224 may include information such as the location of stores, floor maps of facilities such as shopping centers, art museums, and museums, associated with location information on the map (e.g., latitude and longitude). Map information 224 may also include detailed map information such as the width (road width), gradient, and curvature of roads and passages. Map information 224 may also include information on feature points corresponding to the road shape and the edges of structures (static obstacles) installed in the surrounding area. Map information 224 can be created by taking a source image for creating map information 224, retaining only static objects such as structures that are always captured in images taken by the camera 180 of the mobile body 100 and can be used for feature point extraction, and deleting dynamic objects such as other traffic participants that are not necessarily captured in images taken by the camera 180 due to movement and cannot be used for feature point extraction. Map information 224 can be created with high accuracy by, for example, the creator of map information 224 distinguishing between static objects to remain in map information 224 and dynamic objects to be removed from map information 224, and manually (by hand, etc.) removing the dynamic objects. Map information 224 may also be created by, for example, using a trained model that has been trained to distinguish between dynamic and static objects to distinguish and remove dynamic objects from map information 224.

[0042] The memory unit 220 may store user characteristic information, characteristic information of specific objects (for example, characteristic information based on shape, size, color, etc.), and information indicating the correspondence between user actions (gestures) and the operation control of the mobile body 100. At least a portion of the information stored in the memory unit 220 may be updated from time to time by the communication unit 190 communicating with other devices such as the management device 10 or the information providing device 20.

[0043] The acquisition unit 202 acquires the detection results from the camera 180, which captures images of the surrounding environment of the moving object 100. For example, the acquisition unit 202 acquires the image captured by the camera 180 (hereinafter referred to as "camera image").

[0044] The acquisition unit 202 acquires information obtained by the communication unit 190 or an operation unit (not shown). For example, the acquisition unit 202 acquires information regarding a movement mode designated by the user through an operation unit (not shown). The operation unit receives input operations by the user. The operation unit includes at least one of, for example, a touch panel, a switch, a key, etc. The operation unit may have a voice input unit (such as a microphone) that receives the user's input operation by voice input or receives the voices of surrounding people. The movement modes include at least a following mode of moving following the user, a leading mode of leading the user to move towards a destination, a parallel running mode of moving side by side with the user, etc. The movement modes may include a standby mode of retreating (standing by) at a position designated by the user and an emergency mode of performing specific control in an emergency of the user.

[0045] The camera image is an example of the "image".

[0046] The information processing unit 204 manages, for example, the information acquired from the terminal device 2, the management device 10, and the information providing device 20. For example, based on the information received by an operation unit (not shown), the information processing unit 204 transmits information to the terminal device 2, the management device 10, and the information providing device 20, or outputs the information received from each device (such as control information 222, map information 224, feature information of the user or a specific object) to each component of the control device 200, or stores it in the storage unit 220.

[0047] The information processing unit 204 performs processes such as registering characteristic information of users who use the mobile device 100 and authenticating users. For example, when registering user characteristic information, the information processing unit 204 generates characteristic information about the face, body shape (back shape), hair color, skin color, and clothing color from the user's face image and full-body image captured by the camera 180 before the user starts using the mobile device 100. The information processing unit 204 may also generate characteristic information about the user's posture and movements (walking motion). The information processing unit 204 may acquire the user's voice data and palm image to generate characteristic information about fingerprints, veins, and voice. The information processing unit 204 may also generate information about the operation of the mobile device 100 corresponding to the user's gestures and register the generated information in the storage unit 220. When performing user authentication, the information processing unit 204 generates feature information from an image of the user and audio acquired from a microphone. Based on the generated feature information, it refers to feature information previously stored in the storage unit 220. If matching feature information (feature information with a similarity of a threshold or higher) exists, it permits the user to use the mobile device 100. If no matching feature information exists in the storage unit 220, the information processing unit 204 does not permit use and instead causes an error message or a message prompting the registration of feature information to be output to an information output unit (not shown). The information output unit (not shown) is, for example, an example of a notification unit for notifying predetermined information around the mobile device 100. The information output unit includes, for example, a display unit and an audio output unit. The display unit is, for example, a liquid crystal display (LCD) or an organic electroluminescence (EL) display. The display unit may be configured as a touch panel and integrated with an operation unit (not shown). The display unit may include a light-emitting unit composed of light-emitting elements such as LEDs (Light Emitting Diodes) that emit light of a predetermined color. The audio output unit is, for example, a speaker. The audio output unit outputs (sounds) sounds corresponding to the operation of the mobile body 100 and sounds corresponding to the display (images, etc.) from the display unit.

[0048] The recognition unit 206 recognizes the surrounding environment of the moving body 100, for example, based on camera images captured by the camera 180. The recognition unit 206 recognizes, for example, objects present around the moving body 100, their positions (distance from the moving body 100 and orientation relative to the direction of movement of the moving body 100 (plus X direction)), their speed, acceleration, and other conditions, as well as the type of object. Objects include, for example, other traffic participants such as vehicles, bicycles, and pedestrians (people), obstacles such as fallen objects present (on the ground) in the movement path of the moving body 100, such as passages within facilities or roads outdoors, and some or all of structures placed (installed) in the movement path. The type is, for example, information that distinguishes vehicles, bicycles, pedestrians, obstacles, structures, etc. The type is not limited to information that distinguishes specific objects such as the vehicles, bicycles, pedestrians, obstacles, and structures mentioned above, but may also be, for example, information that distinguishes whether an object is dynamic or static. The recognition unit 206 distinguishes and recognizes moving, dynamic objects (e.g., obstacles such as other traffic participants) and stationary, static objects (e.g., obstacles such as structures) that are present around the moving body 100, based on the camera image captured by the camera 180. Furthermore, the recognition unit 206 may also distinguish and recognize objects that are dynamic but are currently stationary (remaining in their current position) (hereinafter referred to as "semi-dynamic objects"). Semi-dynamic objects are not limited to dynamic objects that remain in their current position (vehicles, bicycles, pedestrians (people), and other traffic participants), but may also include objects that can be moved, such as desks, tables, and chairs. Object recognition in the recognition unit 206 may be performed by processing such as pattern matching using pre-set object patterns for recognizing vehicles, bicycles, pedestrians, and fixed objects (including obstacles and structures). Object recognition in the recognition unit 206 may also be performed using AI (Artificial Intelligence) technology such as machine learning. More specifically, the recognition unit 206 may, for example, perform the recognition by inputting the camera image captured by the camera 180 into a trained model that has been trained to output information such as the presence, location, and type of an object when an image is input, thereby extracting features such as the presence, location, and type of an object captured in the camera image.

[0049] When a detection device (sensor) such as a radar device or LIDAR is provided in the moving body 100, the recognition unit 206 may recognize the situation around the moving body 100 by using the detection results of the radar device, LIDAR, etc. in addition to or instead of the camera image.

[0050] The recognition unit 206 outputs information regarding each object (including dynamic objects, semi-dynamic objects, and static objects) obtained by associating information indicating the presence, type, and state of the recognized object (such as position, moving speed, and acceleration) with information indicating the certainty (i.e., reliability) of the recognized object to the self-position estimation unit 208.

[0051] The object recognition information is an example of the "recognition result of an object". A dynamic object is an example of the "first object", and a semi-dynamic object is an example of the "second object". The "second object" may be included in the "first object".

[0052] The self-position estimation unit 208 estimates the current position (self-position) of the moving object 100. The self-position estimation unit 208 estimates the self-position of the moving object 100 based on, for example, the camera image acquired by the acquisition unit 202. More specifically, the self-position estimation unit 208 extracts feature points of static objects from the camera image and estimates the self-position of the moving object 100 based on the extracted feature points. In this case, for example, if the feature points of static objects are hidden by dynamic or semi-dynamic objects captured in the camera image, the self-position estimation unit 208 may not be able to extract the feature points of static objects that it should be able to extract, and may instead extract feature points from the area of ​​the dynamic or semi-dynamic object. In other words, the self-position estimation unit 208 may extract incorrect feature points and be unable to correctly estimate the self-position of the moving object 100. Therefore, based on the object recognition information output by the recognition unit 206, the self-position estimation unit 208 excludes the regions of dynamic and semi-dynamic objects captured in the camera image from the regions in the camera image from which feature points are extracted, and extracts feature points from the regions that are not excluded. In other words, the self-position estimation unit 208 extracts feature points with the regions of dynamic and semi-dynamic objects captured in the camera image excluded, and estimates the self-position of the moving object 100. One method by which the self-position estimation unit 208 excludes the regions of dynamic and semi-dynamic objects in the camera image is, for example, masking by filling the regions of dynamic and semi-dynamic objects in the camera image with black. In the following explanation, the state in which the regions of dynamic and semi-dynamic objects captured in the camera image are masked will be referred to as the state in which the regions of dynamic and semi-dynamic objects in the camera image are excluded.

[0053] However, for example, if the camera image contains many moving or semi-moving objects, masking the regions of all moving or semi-moving objects would mask a large portion of the camera image, making it impossible for the self-position estimation unit 208 to extract feature points necessary to estimate the self-position of the moving object 100. In other words, the self-position estimation unit 208 may not be able to estimate the self-position of the moving object 100 using only the feature points it has extracted. Therefore, the self-position estimation unit 208 changes the region of objects to be masked in the camera image based on the object recognition information output by the recognition unit 206, and extracts feature points in this state. Then, the self-position estimation unit 208 estimates the self-position of the moving object 100 based on the extracted feature points. Details of the functions and operation of the self-position estimation unit 208 will be described later.

[0054] The self-position estimation unit 208 outputs information representing the estimation result of the self-position of the moving body 100 (hereinafter referred to as "self-position information") to the trajectory generation unit 210 and the movement control unit 212, respectively.

[0055] Self-localization information is an example of "self-localization estimation results."

[0056] The trajectory generation unit 210 generates a target trajectory for the mobile object 100 to travel in the future, that is, a travel path to the destination, based on the surrounding conditions of the mobile object 100 recognized by the recognition unit 206 and the self-position of the mobile object 100 (self-position information) estimated by the self-position estimation unit 208. The trajectory generation unit 210 generates a target trajectory that allows the mobile object 100 to move smoothly to the target point, for example, in accordance with the operation control corresponding to the movement mode instructed by the user (e.g., follow control, lead control, parallel control, evasive control, emergency control).

[0057] For example, if the movement mode of the mobile body 100 is a follow mode in which it moves in front of the user, the trajectory generation unit 210 generates a target trajectory that follows the user at a position within a predetermined distance range from the user, and that follows the user with the target point being behind the user (which may be diagonally behind the user so that it is visible to the user). For example, if the movement mode of the mobile body 100 is a lead mode in which it moves in front of the user toward a destination, the trajectory generation unit 210 generates a target trajectory that leads the user at a position within a predetermined distance range from the user, and that follows the user with the target point being in front of the user (which may be diagonally in front of the user). For example, if the movement mode of the mobile body 100 is a parallel mode in which it moves alongside the user, the trajectory generation unit 210 generates a target trajectory that runs parallel to the user at a position within a predetermined distance range from the user, and that follows the user to the side (which may be diagonally in front of or behind the user). For example, if the mobile unit 100's movement mode is a standby mode in which it retreats (stands) to a location specified by the user, the trajectory generation unit 210 generates a target trajectory with the specified location (retreat location) as the target point. For example, if the mobile unit 100's movement mode is an emergency mode, the trajectory generation unit 210 generates a target trajectory for autonomous movement to seek help from nearby people or facilities. Operation control corresponding to these types of movement modes is performed based on information stored in, for example, control information 222.

[0058] The movement control unit 212 controls the movement of the mobile body 100 to move at a position corresponding to the movement mode set by the user, based on the self-position (self-position information) of the mobile body 100 estimated by the self-position estimation unit 208. More specifically, the movement control unit 212 controls the motors (first motor 122, second motor 132), the brake device 136, and the steering device 138 so that the mobile body 100 travels along a target trajectory generated by the trajectory generation unit 210. At this time, when the mobile body 100 travels along the target trajectory generated by the trajectory generation unit 210, the movement control unit 212 controls the mobile body 100 based on the object recognition result by the recognition unit 206 so that the mobile body 100 does not come into contact with surrounding objects and the distance between the user and the mobile body 100 according to the movement mode is within a predetermined distance range. The predetermined distance range is the distance range between the preset shortest distance and longest distance. The shortest and longest distances may be variable distances depending on factors such as the type of travel mode and surrounding conditions (shape of the travel path, degree of crowding), or they may be fixed distances.

[0059] The movement control unit 212 (which may include the trajectory generation unit 210) is an example of a "movement control unit".

[0060] [Function and Operation of Self-Position Estimation Unit] The self-position estimation unit 208 performs known image analysis processing (e.g., edge extraction, feature extraction, pattern matching, etc.) on the camera image captured by the camera 180, and extracts feature points around the moving object 100 based on the analysis results. Feature points are, for example, feature points of objects in real space included in the camera image (e.g., traffic signals, road signs, traffic participants, buildings, road structures, etc.). Objects may include road markings drawn on the road that can be recognized from the camera image, such as road lane markings and stop lines. For example, the self-position estimation unit 208 extracts a sequence of points on the edges of objects included in the camera image as feature points (feature point group).

[0061] The self-position estimation unit 208 may, for example, extract feature points using a pre-trained model that has been trained to output the edges of objects captured in a camera image as a point cloud when a camera image is input. This pre-trained model may be stored in the memory unit 220 in advance, or it may be acquired from an external device via the communication unit 190 mounted on the mobile body 100. The self-position estimation unit 208 may, for example, extract feature points using the Visual SLAM (Simultaneous Localization and Mapping) method, which is a technique for grasping the self-position in three dimensions from a camera image. The method for extracting feature points from a camera image is not limited to the above example, and other known methods may be used. Based on the extracted feature points, the self-position estimation unit 208 estimates the road (travel path) on which the mobile body 100 is traveling and estimates the position of the mobile body 100 on the road.

[0062] The self-position estimation unit 208 may estimate its own position on the road by comparing feature points obtained from the camera image with feature points included in the map information 224. In this case, the self-position estimation unit 208 may obtain the position information of the mobile body 100 by acquiring the position information of the mobile body 100 using a GPS (Global Positioning System) device (not shown) built into the mobile body 100, or by communicating with a communication device located within a predetermined distance via the communication unit 190 using a wireless communication method such as Bluetooth® to obtain the position information of the communication device, or by extracting surrounding feature points by referring to map information based on the estimated position information.

[0063] However, as mentioned above, if the camera image captured by camera 180 contains many objects (especially moving or semi-moving objects), there is a high possibility that these objects will obscure the feature points of static objects that could be extracted, or that incorrect feature points will be extracted. In this case, when the self-position estimation unit 208 estimates the self-position of the moving object 100, the matching of the extracted feature points with the feature points included in the map information 224 will not be performed correctly, and there is a high possibility that the self-position estimation will not be performed appropriately (accurately).

[0064] Therefore, the self-position estimation unit 208 sets conditions for the type and confidence level of dynamic and semi-dynamic objects to be masked (hereinafter referred to as "masking conditions") based on the position, type, and confidence level information of the recognized object included in the object recognition information output by the recognition unit 206, and masks the areas of dynamic and semi-dynamic objects captured in the camera image according to the set masking conditions. Then, the self-position estimation unit 208 extracts feature points from the areas that are not masked in the camera image and estimates the self-position of the moving object 100 based on the extracted feature points.

[0065] More specifically, the self-position estimation unit 208 initially sets masking conditions to mask the regions of all dynamic and semi-dynamic objects, extracts feature points of static objects while masking the regions of all dynamic and semi-dynamic objects recognized by the recognition unit 206, and estimates the self-position of the moving object 100 based on the extracted feature points. If the self-position can be estimated in this initial state, the self-position estimation unit 208 completes the self-position estimation process (hereinafter referred to as the "self-position estimation process"). On the other hand, if the self-position cannot be estimated in the initial state, the self-position estimation unit 208 sets masking conditions to mask the regions of some dynamic and semi-dynamic objects based on the type of recognized object and the confidence level information included in the object recognition information output by the recognition unit 206. In other words, the self-position estimation unit 208 does not mask the regions of objects with low confidence levels, thereby reducing the proportion (degree of masking) of the regions of dynamic and semi-dynamic objects that are masked in the camera image.

[0066] In this case, the self-position estimation unit 208 may set masking conditions that mask areas of dynamic or semi-dynamic objects with similar confidence levels regardless of the object type, or it may set masking conditions that differentiate the confidence levels of dynamic or semi-dynamic objects to be masked for each object type. For example, the self-position estimation unit 208 may set masking conditions such that a predetermined confidence threshold is used to mask dynamic or semi-dynamic objects of type pedestrians with high confidence levels, and a predetermined confidence threshold is used to mask dynamic or semi-dynamic objects of type vehicles with lower confidence levels than objects of type pedestrians. In this case, in the camera image, dynamic or semi-dynamic objects of type pedestrians will be masked when their confidence level is relatively high, while dynamic or semi-dynamic objects of type vehicles will be masked even if their confidence level is relatively low. This is based on the idea that pedestrians (people) are difficult to recognize because they are expected to face in various directions, whereas vehicles, such as cars, buses, and trucks, are expected to have a somewhat limited orientation relative to the direction of movement of the moving object 100 (plus X direction), making them easier to recognize than people. For this reason, if an object that is more difficult to recognize than a pedestrian (person) (a dynamic object or a semi-dynamic object) is captured in the camera image, the predetermined threshold for confidence corresponding to that object may be even lower than that for an object of the type that is a person.

[0067] The self-position estimation unit 208 then extracts feature points of static objects while masking the regions of dynamic and semi-dynamic objects that meet the set masking conditions (satisfying the masking conditions), and estimates the self-position of the moving object 100 based on the extracted feature points. If the self-position can be estimated in this state, the self-position estimation unit 208 completes the self-position estimation process. However, if the self-position cannot be estimated, it sets yet another masking condition and repeats the self-position estimation process. In other words, the self-position estimation unit 208 further reduces the degree of masking of the regions of dynamic and semi-dynamic objects that are masked in the camera image and repeats the self-position estimation process.

[0068] Here, an example of a case where the self-position estimation unit 208 sets (changes) masking conditions will be described. Figure 4 is a diagram illustrating the method of estimating the self-position in the control device 200 (self-position estimation unit 208) of the moving body 100 according to the embodiment.

[0069] Figure 4(a) schematically shows an example of the area in which the mobile body 100 moves, and Figure 4(b) schematically shows an example of a forward camera image captured by the camera 180 as the mobile body 100 moves, that is, an image capturing the area in the direction in which the mobile body 100 will travel in the future. In Figures 4(a) and 4(b), object Ob1 and object Ob2 are objects such as chairs and tables, respectively, and pedestrian P1 and pedestrian P2 are people. Here, the map information 224, which the self-position estimation unit 208 compares feature points to estimate its own position, shows the positions of object Ob1 and object Ob2, but does not show pedestrian P1 or pedestrian P2. In other words, for example, suppose the creator of the map information 224 removes dynamic objects such as people from the source image used to generate the map information 224, and determines that objects Ob1 and Ob2, which are placed side by side, are objects that can be moved, but are moved infrequently and can be used to extract feature points, and thus creates map information 224 that retains them as static objects. Furthermore, in the example shown in Figures 4(a) and 4(b), assume that the position of object Ob1 has been moved, and that pedestrians P1 and P2 have been moved. In Figure 4(b), an example of a feature point FP that the self-position estimation unit 208 can normally extract is shown with an "x".

[0070] Here, we consider the case where the self-position estimation unit 208 is set to mask the regions of all dynamic and semi-dynamic objects recognized by the recognition unit 206. Figure 4(c) shows an example where the region of the moving object Ob1 (semi-dynamic object) recognized by the recognition unit 206 is set as masking region MA1, the region of pedestrian P1 (dynamic object) is set as masking region MA2, and the region of pedestrian P2 (dynamic object) is set as masking region MA3, thereby masking the regions of all dynamic and semi-dynamic objects recognized by the recognition unit 206. In the example shown in Figure 4(c), the edge portions for extracting feature points FP1, FP2, and FP3, which are indicated by the × marks in Figure 4(b), are also masked by masking regions MA1 and MA3. This increases the likelihood that the self-position estimation unit 208 will be unable to estimate the self-position of the moving object 100 based on feature points FP other than the feature points FP1, FP2, and FP3 that it was able to extract.

[0071] On the other hand, consider the case where the self-position estimation unit 208 is set (changed) to mask the regions of some dynamic or semi-dynamic objects recognized by the recognition unit 206. Figure 4(d) shows an example where only the region of pedestrian P2 (dynamic object) recognized by the recognition unit 206 (masking region MA3) is masked. In other words, Figure 4(d) shows an example where the region of the moved object Ob1 (semi-dynamic object) (masking region MA1) and the region of pedestrian P1 (dynamic object) (masking region MA2) are not masked. In the example shown in Figure 4(d), the edge portion for extracting feature point FP1, which is masked by masking region MA1 in Figure 4(c), is not masked, and only the edge portions for extracting feature points FP2 and FP3 are masked. Therefore, the self-position estimation unit 208 can also extract feature point FP1, and based on the feature points FP other than feature points FP2 and FP3 that were extracted, the likelihood of correctly estimating the self-position of the moving object 100 increases.

[0072] Through these functions and operations, the self-position estimation unit 208 modifies the areas of dynamic and semi-dynamic objects to be masked in the camera image based on the object recognition information output by the recognition unit 206, and estimates the self-position of the mobile body 100 based on feature points extracted from the unmasked areas in the camera image. As a result, the self-position estimation unit 208 can estimate the self-position of the mobile body 100 more favorably (accurately). In other words, the self-position estimation unit 208 can stably estimate the self-position of the mobile body 100. This allows the control device 200 to move the mobile body 100 more safely and smoothly.

[0073] [Flowchart of Self-Position Estimation Process] Figure 5 is a flowchart showing an example of the flow of the self-position estimation process executed in the control device 200 of the mobile body 100 according to the embodiment. Figure 6 is a diagram showing an example of self-position estimation in the control device 200 of the mobile body 100 according to the embodiment. In the following description, the process of this flowchart shown in Figure 5 will be explained with reference to the example of self-position estimation shown in Figure 6 as appropriate. The process of this flowchart is repeatedly executed in the control device 200 (more specifically, the acquisition unit 202, the recognition unit 206, and the self-position estimation unit 208) at predetermined time intervals when the camera 180 images the area around the mobile body 100 (at least in front of it). In the following description, for the sake of simplicity, it will be assumed that the camera 180 images the area in front of the mobile body 100. In the following description, it will be assumed that the self-position estimation unit 208 estimates the self-position of the mobile body 100 by comparing feature points obtained from the camera image with feature points included in the map information 224.

[0074] When the camera 180 captures an image in front of the moving body 100, the acquisition unit 202 acquires the camera image captured by the camera 180 (step S100).

[0075] The recognition unit 206 recognizes the situation in front of the moving body 100 (objects present in front) based on the camera image acquired by the acquisition unit 202 (step S110). At this time, the recognition unit 206 distinguishes and recognizes dynamic objects, semi-dynamic objects, and static objects present in front of the moving body 100.

[0076] Figure 6(a) schematically shows an example of the objects recognized by the recognition unit 206 in the camera image of the front of the moving object 100 captured by the camera 180, and an example of the object recognition information for those objects. In the example shown in Figure 6(a), the recognition unit 206 recognizes the vehicle and the pedestrian (person) in the camera image, and the area of ​​each recognized object is designated as the object area OA. In Figure 6(a), an example of object recognition information, such as the type of object and the confidence level of the object recognized by the recognition unit 206, is shown corresponding to each object area OA. In the example of object recognition information shown in Figure 6(a), the type of vehicle recognized by the recognition unit 206 is designated as "Vehicle," the type of pedestrian (person) is designated as "Person," and the confidence level of each object is represented as "Confidence Level LV." The confidence level (LV) represents the degree of confidence of an object as a percentage.

[0077] The recognition unit 206 outputs object recognition information (see Figure 6(a)) for each recognized object to the self-position estimation unit 208.

[0078] The self-position estimation unit 208 sets masking conditions based on the position, type, and confidence level information of dynamic and semi-dynamic objects included in the object recognition information output by the recognition unit 206 (step S120). Here, we assume that this is the initial estimation of the self-position of the moving body 100. In this case, the self-position estimation unit 208 sets masking conditions to mask the regions of all dynamic and semi-dynamic objects included in the object recognition information. Then, the self-position estimation unit 208 masks the regions of all dynamic and semi-dynamic objects captured in the camera image according to the set masking conditions (step S130). Then, the self-position estimation unit 208 extracts feature points from the unmasked regions in the camera image (step S132).

[0079] Figure 6(b) schematically shows an example where a masking condition is set with the area of ​​all dynamic and semi-dynamic objects recognized by the recognition unit 206 shown in Figure 6(a) as the masking area MA, and feature points FP are extracted from the unmasked area. In Figure 6(b), an example of feature points FP extracted by the self-position estimation unit 208 is shown with an "x" mark.

[0080] The self-position estimation unit 208 estimates the self-position of the mobile object 100 based on the extracted feature points (marked with an "x" in Figure 6(b)) (step S134). More specifically, the self-position estimation unit 208 estimates the self-position of the mobile object 100 by comparing the feature points extracted in step S132 with the feature points included in the map information 224 and matching the feature points. The self-position estimation unit 208 then determines whether the estimation of the self-position of the mobile object 100 was performed correctly (step S140). This determination can be made, for example, by determining whether the extracted feature points and the feature points included in the map information 224 matched to a predetermined percentage or more, which indicates that the self-position could be correctly estimated. If the self-position estimation of the mobile object 100 is determined to have been performed correctly in step S140, the self-position estimation unit 208 proceeds to step S144.

[0081] On the other hand, if in step S140 the self-position estimation of the moving body 100 is determined to be incorrect, the self-position estimation unit 208 resets the masking conditions based on the type and confidence level information of the dynamic and semi-dynamic objects included in the object recognition information output by the recognition unit 206. Furthermore, the self-position estimation unit 208 stores the reset masking conditions (step S142). More specifically, the self-position estimation unit 208 changes and resets the masking conditions to a condition that reduces the degree of masking of the areas of dynamic and semi-dynamic objects that are masked in the camera image by masking the areas of some dynamic and semi-dynamic objects included in the object recognition information. For example, the self-position estimation unit 208 sets masking conditions in the object recognition information corresponding to each dynamic and semi-dynamic object recognized by the recognition unit 206 to mask the areas of dynamic and semi-dynamic objects whose confidence level is above a predetermined threshold. Furthermore, the self-position estimation unit 208 stores the reset masking conditions in, for example, the storage unit 220. Then, the self-position estimation unit 208 returns to step S130 and, according to the masking conditions (second masking conditions) that have been reset, masks the areas of dynamic and semi-dynamic objects captured in the camera image, and repeats the processing of steps S132 to S140.

[0082] Figure 6(c) shows an example where a predetermined threshold for confidence is set to 80%. More specifically, Figure 6(c) schematically shows an example where a masking condition is set in which the area of ​​dynamic objects and semi-dynamic objects with a confidence level of 80% (confidence level LV = 80) or higher among all dynamic and semi-dynamic objects recognized by the recognition unit 206 shown in Figure 6(a) is designated as the masking area MA, and feature points FP (marked with an "x") are extracted from the unmasked area.

[0083] As a result, the self-position estimation unit 208 performs a second estimation of the self-position of the moving object 100 (step S134) based on the feature points (marked with an "x" in Figure 6(c)) extracted using the reset masking conditions (second masking conditions), and makes a second determination (step S140) as to whether the self-position estimation was performed correctly.

[0084] If the self-position estimation of the moving object 100 is determined to be incorrect in the second step S140, the self-position estimation unit 208 changes and resets the masking conditions to further reduce the degree of masking of the areas of dynamic and semi-dynamic objects that are masked in the camera image, stores the settings (step S142), and repeats the processing from steps S130 to S140. In other words, the self-position estimation unit 208 performs the self-position estimation process for the third time.

[0085] Figure 6(d) shows an example where a predetermined threshold for confidence is set to 90%. More specifically, Figure 6(d) schematically shows an example where the masking conditions are reset to define a masking region MA as the area of ​​dynamic and semi-dynamic objects with a confidence level of 90% or higher among all dynamic and semi-dynamic objects recognized by the recognition unit 206 shown in Figure 6(a), and feature points FP (marked with an "x") are extracted from the unmasked area.

[0086] As a result, the self-position estimation unit 208 performs a third estimation of the self-position of the moving object 100 (step S134) based on the feature points (marked with an "x" in Figure 6(d)) extracted using the reset masking conditions (third masking conditions), and makes a third determination (step S140) as to whether the self-position estimation was performed correctly.

[0087] In this way, the self-position estimation unit 208 gradually changes a predetermined threshold for the reliability of the object to be masked according to the masking conditions, that is, it gradually reduces the degree of masking of the areas of dynamic and semi-dynamic objects that are masked in the camera image, and repeats the self-position estimation process until the self-position estimation is performed correctly.

[0088] In the example described above, Figure 6(c) shows an example where the predetermined confidence threshold is set to 80%, and Figure 6(d) shows an example where the predetermined confidence threshold is set to 90%. In other words, the example described above shows a case where the predetermined confidence threshold, which is changed in stages, is set in 10% increments. However, the change in the predetermined confidence threshold is not limited to the 10% increments described above. For example, the change range in the predetermined confidence threshold may be narrowed to 5% increments, or widened to 15% increments, for example.

[0089] In step S140 (or any step S140), if it is determined that the self-position estimation of the mobile body 100 has been performed correctly, the self-position estimation unit 208 outputs self-position information representing the estimated self-position of the mobile body 100 to the trajectory generation unit 210 and the movement control unit 212, respectively (step S144). As a result, the control device 200 generates a travel path (target trajectory) for the mobile body 100 based on the correctly estimated self-position, and the movement control unit 212 performs movement control of the mobile body 100.

[0090] However, in the process of step S140, it is possible that a determination result indicating that the self-position of the moving body 100 has been correctly estimated may not be obtained even when masking conditions are set that do not mask the areas of all dynamic and semi-dynamic objects. In other words, it is possible that it is impossible to estimate the self-position of the moving body 100 based on the current camera image. In this case, the self-position estimation unit 208 outputs self-position information indicating that it is impossible to estimate the self-position of the moving body 100 to both the trajectory generation unit 210 and the movement control unit 212. In this case, the control device 200 will perform self-position estimation of the moving body 100 based on the next camera image.

[0091] Then, the self-position estimation unit 208 returns the process to step S100. As a result, when the camera 180 captures more images in front of the moving body 100, that is, when the next camera image is captured and it is time to perform the self-position estimation process, the control device 200 repeats the processes from step S100 to step S144. At this time, the self-position estimation unit 208 may set the masking conditions from the previous self-position estimation process, which were stored in the process of step S142, as the initial masking conditions and perform the self-position estimation process. In this case, the control device 200 can estimate the self-position of the moving body 100 faster than the previous self-position estimation process, and furthermore, with a reduced load on the self-position estimation process.

[0092] However, even if the masking conditions from the previous self-position estimation process are set as the initial masking conditions, it may still be impossible to estimate the self-position of the moving object 100 based on the current camera image. In this case, the control device 200 discards the stored masking conditions and sets masking conditions that mask the regions of all dynamic and semi-dynamic objects included in the object recognition information based on the next camera image, and then estimates the self-position of the moving object 100 based on the next camera image.

[0093] Through this self-position estimation process flow, the control device 200's self-position estimation unit 208 gradually changes the degree of masking of the areas of dynamic and semi-dynamic objects in the camera image based on the object recognition information of the objects (dynamic and semi-dynamic objects) recognized by the recognition unit 206, and estimates the self-position of the moving body 100 based on feature points extracted from the unmasked areas in the camera image. As a result, the control device 200's self-position estimation unit 208 can estimate the self-position of the moving body 100 more favorably (accurately). In other words, the control device 200's self-position estimation unit 208 can stably estimate the self-position of the moving body 100. This allows the control device 200 to move the moving body 100 more safely and smoothly.

[0094] As described above, according to the control device for a mobile body of the embodiment, the recognition unit 206 recognizes objects present around the mobile body 100, their position, movement speed, acceleration, and other states, and the type of object, based on the camera image captured by the camera 180. Then, in the control device for a mobile body of the embodiment, the self-position estimation unit 208 estimates the self-position of the mobile body 100 based on the camera image. More specifically, in the control device for a mobile body of the embodiment, the self-position estimation unit 208 sets conditions (masking conditions) for the areas of objects to be masked in the camera image based on the object recognition information output by the recognition unit 206, and estimates the self-position of the mobile body 100 based on the camera image by extracting feature points of static objects while masking the areas of dynamic and semi-dynamic objects captured in the camera image according to the set masking conditions. As a result, in the control device for a mobile body of the embodiment, the self-position estimation unit 208 can stably and more suitably (accurately) estimate the self-position of the mobile body 100. As a result, in the control device for the mobile body of this embodiment, the trajectory generation unit 210 automatically generates a future travel path (target trajectory) for the mobile body 100 based on a more accurately estimated self-position (without relying on the operation of the driver (occupant P)), and the movement control unit 212 can control the movement of the mobile body 100 based on the generated travel path (target trajectory). As a result, the control device for the mobile body of this embodiment can operate the mobile body 100 more safely and smoothly.

[0095] The embodiment described above can be expressed as follows: A control device for a moving body comprising: a storage medium for storing computer-readable instructions; a processor connected to the storage medium, wherein the processor executing the computer-readable instructions to: acquire an image of the surrounding environment of a moving body; recognize objects present in the vicinity of the moving body based on the image; estimate the self-position of the moving body based on the image; control the movement of the moving body based on the estimation result of the self-position; when recognizing an object, at least a first moving object present in the vicinity of the moving body is recognized as the object; when estimating the self-position, a region of the object to be excluded from the region in the image is set based on the type of each object and a confidence level representing the likelihood of each object; and the self-position is estimated based on the information of the region in the image that is not excluded.

[0096] In the embodiment, the self-position estimation unit 208 initially sets a masking condition to mask the areas of all dynamic and semi-dynamic objects, and if it determines that the estimation of the self-position of the moving body 100 is not performed correctly, it estimates the self-position of the moving body 100 by gradually changing the masking condition to reduce the degree of masking of the areas of dynamic and semi-dynamic objects to be masked. However, the setting of the masking condition is not limited to the method of gradually reducing the degree of masking of the areas of dynamic and semi-dynamic objects to be masked. For example, if the self-position estimation unit 208 determines that the estimation of the self-position of the moving body 100 is not performed correctly, it may set the average reliability of all dynamic and semi-dynamic objects as a predetermined reliability threshold, and set a masking condition for the second time that masks the areas of dynamic and semi-dynamic objects that are above this threshold. Furthermore, in the second self-position estimation process, if the self-position estimation unit 208 determines that the self-position estimation of the moving object 100 has not been performed correctly, it may use the average reliability value of the masked dynamic or semi-dynamic objects as a predetermined threshold for the next reliability. If it determines that the self-position estimation of the moving object 100 has been performed correctly, it may use the average reliability value of the unmasked dynamic or semi-dynamic objects as a predetermined threshold for the next reliability. In other words, depending on the determination result in the process of step S140, the self-position estimation unit 208 may lower or increase the degree of masking of the areas of dynamic or semi-dynamic objects to be masked in the camera image, while searching for a degree of masking at which the self-position estimation of the moving object 100 is determined to have been performed correctly. This allows the self-position estimation unit 208 to estimate the self-position of the moving object 100 stably and more favorably (accurately) while reducing the load on the next self-position estimation process. In this case, the processing of the self-position estimation unit 208 (self-position estimation processing) should be equivalent to the self-position estimation processing of the self-position estimation unit 208 described in the embodiment.

[0097] In the embodiment, the case was described in which the self-position estimation unit 208 sets masking conditions to mask the areas of dynamic and semi-dynamic objects captured in the camera image based on the object recognition results (object recognition information of objects recognized by the recognition unit 206) by the recognition unit 206. However, the recognition unit 206 recognizes not only dynamic and semi-dynamic objects captured in the camera image, but also static objects such as structures. In other words, the recognition unit 206 recognizes objects other than other traffic participants that exist on the travel path that the moving body 100 will travel in the future, such as vehicles, bicycles, and pedestrians (people). In contrast, the self-position estimation unit 208 uses object recognition information of dynamic and semi-dynamic objects when setting masking conditions in the self-position estimation process. Therefore, the self-position estimation unit 208 may be configured to include, for example, a recognition unit (not shown) that recognizes only dynamic and semi-dynamic objects captured in the camera image (i.e., a recognition unit specialized in recognizing dynamic and semi-dynamic objects), and instead of the object recognition information output by the recognition unit 206, it may be configured to set masking conditions that mask the areas of dynamic and semi-dynamic objects captured in the camera image based on the object recognition information of dynamic and semi-dynamic objects recognized by this recognition unit (not shown). The self-position estimation unit 208 may also be configured to set masking conditions that mask the areas of specific types of objects, such as pedestrians (people), captured in the camera image. The self-position estimation unit 208 may also be configured to set masking conditions that mask areas other than those of static objects such as structures captured in the camera image, without masking those areas. In these cases, the recognition unit (not shown) is an example of a "recognition unit". In these cases, the processing of the self-position estimation unit 208 (self-position estimation processing) should be equivalent to the self-position estimation processing of the self-position estimation unit 208 described in the embodiment.

[0098] In the embodiment, the case in which the self-position estimation unit 208 estimates the self-position of the moving object 100 while the moving object 100 is moving was described. However, as mentioned above, for example, when creating map information 224, dynamic objects captured in the image (original image) are excluded (masked). In the above description, it was explained that the masking of dynamic objects when creating map information 224 is performed manually (by hand, etc.) by the creator of the map information 224, but when creating map information 224, a masking process based on the same concept as the self-position estimation unit 208 can be applied. In this case, the processing (manual work) required by the creator of map information 224 can be reduced. The masking process in this case should be equivalent to the self-position estimation process in the self-position estimation unit 208 described in the embodiment (more specifically, the process of setting masking conditions).

[0099] Although embodiments for carrying out the present invention have been described above using examples, the present invention is not limited in any way to these embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention.

[0100] 1...Mobile control system 2...Terminal device 10...Management device 20...Information provision device 100...Mobile body 110...Base 112...Door section 120...First wheel 122...First motor 130...Second wheel 132...Second motor 134...Battery 136...Brake device 138...Steering device 140...Third wheel 150...Support 180...Camera 190...Communication unit 200...Control device 202...Acquisition unit 204...Information processing unit 206...Recognition unit 208...Self-position estimation unit 210...Trajectory generation unit 212...Movement control unit 220...Storage unit 222...Control information 224...Map information

Claims

An acquisition unit that acquires images capturing the surrounding conditions of a moving object, Based on the aforementioned image, a recognition unit recognizes objects present around the moving object, A self-position estimation unit that estimates the self-position of the moving object based on the aforementioned image, A movement control unit that controls the movement of the moving body based on the self-position estimation result, Equipped with, The recognition unit, At a minimum, the moving first object present around the moving body is recognized as the object, The self-position estimation unit, Based on the type of each object and the confidence level representing its likelihood, a region of the object to be excluded from the region in the image when estimating its own position is set. Based on the information of the region within the image that has not been excluded, the self-position is estimated. A control device for mobile vehicles.   The recognition unit, The second object, which is the first object but is currently stationary, is further recognized as the object. A control device for a mobile body according to claim 1.   The self-position estimation unit, First, the region of all the first objects is set as the region of the object to be excluded, and the self-position is estimated. A control device for a mobile body according to claim 1.   The self-position estimation unit, If the self-position cannot be estimated, the region of the object whose confidence level is above a predetermined threshold is set as the region of the object to be excluded, the degree to which the region of the object is excluded from the region in the image is changed, and the self-position is estimated. A control device for a mobile body according to claim 2.   The self-position estimation unit, If the self-position cannot be estimated, the threshold is changed in stages to estimate the self-position. A control device for a mobile body according to claim 4.   The self-position estimation unit, If the self-position cannot be estimated, the self-position is estimated by lowering or increasing the degree to which the region of the object is excluded. A control device for a mobile body according to claim 5.   The self-position estimation unit, When the self-position can be estimated, the threshold value is stored. A control device for a mobile body according to any one of claims 4 to 6.   The self-position estimation unit, When estimating the self-position, instead of all the object regions, the regions of the object that are equal to or greater than the stored threshold are initially set as the regions of the object to be excluded, and the self-position is estimated. A control device for a mobile body according to claim 7.   An acquisition unit that acquires images capturing the surrounding conditions of a moving object, Based on the aforementioned image, a recognition unit recognizes objects present around the moving object, A self-position estimation unit that estimates the self-position of the moving object based on the aforementioned image, Equipped with, The recognition unit, At a minimum, the moving first object present around the moving body is recognized as the object, The self-position estimation unit, Based on the type of each object and the confidence level representing its likelihood, a region of the object to be excluded from the region in the image when estimating its own position is set. Based on the information of the region within the image that has not been excluded, the self-position is estimated. Mobile control system.   Computers Images are acquired that capture the surrounding environment of the moving object. Based on the aforementioned image, objects present around the moving object are recognized. Based on the aforementioned image, the self-position of the moving object is estimated. Based on the self-position estimation results, the movement of the moving body is controlled. When recognizing the aforementioned object, At a minimum, the moving first object present around the moving body is recognized as the object, When estimating the aforementioned self-position, Based on the type of each object and the confidence level representing its likelihood, a region of the object to be excluded from the region in the image is set. Based on the information of the region within the image that has not been excluded, the self-position is estimated. A method for controlling a moving object.   On the computer, The system captures images of the surrounding environment of the moving object. Based on the aforementioned image, the object present around the moving object is recognized. Based on the aforementioned image, the self-position of the moving object is estimated. Based on the self-position estimation result, the movement of the moving body is controlled. When making the aforementioned object recognized, At a minimum, the system will recognize a moving first object that is present around the moving body as the object. When estimating the aforementioned self-position, Based on the type of each object and the confidence level representing its likelihood, the region of the object to be excluded from the region in the image is set. Based on the information of the region within the image that has not been excluded, the self-position is estimated. A storage medium that stores programs.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, program and movable body

    JP2019045892A

  • Control device, control method, program, and mobile body

    JP2019053561A

  • Vehicle position estimation device

    JP7371783B2