Information processing device, information processing system, and program
The integration of image and depth data with 3D semantic segmentation in parking support systems addresses the challenge of detecting parking spaces, providing quick and comfortable parking assistance across varying environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SONY SEMICON SOLUTIONS CORP
- Filing Date
- 2025-01-23
- Publication Date
- 2026-04-27
AI Technical Summary
Existing parking support systems fail to accurately detect parking spaces considering depth information, leading to difficulties in recognizing parking spaces in various environments and resulting in inefficient and uncomfortable parking experiences.
An information processing device and program that utilize image and depth data acquisition, combined with 3D semantic segmentation, to generate a 3D semantic segmentation image for detecting parking spaces, enabling quick and smooth parking assistance regardless of environmental conditions.
Enables rapid and natural parking assistance similar to human perception, allowing for efficient parking in diverse parking spaces without the need for prior registration or clear white lines.
Smart Images

Figure 0007852098000001 
Figure 0007852098000002 
Figure 0007852098000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, information processing system and a program, and in particular, to an information processing apparatus and information processing that can realize quick and smooth parking support close to human senses without being affected by the environment of a parking space, system and a program.
Background Art
[0002] In recent years, interest in parking support systems has been increasing. There are situations where a vehicle needs to be parked in various daily driving scenes, and a parking support system that is safer, more comfortable, and more convenient is demanded.
[0003] For example, a technique has been proposed in which a parking space is detected based on a camera image and parking support for parking in the detected parking space is provided (see Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the technique described in Patent Document 1, detection of a parking space considering depth information is not performed.
[0006] Therefore, in an environment where it is difficult to detect a parking space without being able to recognize the depth, it is impossible to discriminate the parking space with a feeling close to that of a human detecting the parking space, and there is a risk that smooth parking support cannot be realized.This disclosure was made in light of these circumstances, and in particular aims to realize quick and smooth parking assistance that is close to human perception, regardless of the environment of the parking space. [Means for solving the problem]
[0008] An information processing device and program, which are aspects of this disclosure, are used to capture images of a vehicle taken by a camera. around prescribed imaging An image acquisition unit that acquires image data within a range, and a depth data acquisition unit that acquires depth data within a range that overlaps with at least a portion of the imaging range of the image data, Class information at the pixel level and The aforementioned Depth data and Correspondence This is an information processing device and program comprising a 3D semantic segmentation processing unit that generates a 3D semantic segmentation image, and a parking space detection unit that detects a parking space based on the 3D semantic segmentation image.
[0009] Information Processing, One Aspect of This Disclosure system This is a vehicle, as captured by a camera. around prescribed imaging Image acquisition: Get image data within a specified range. Department Then, depth data acquisition is performed to obtain depth data in an area that overlaps with, at least partially, the imaging range of the aforementioned image data. Department And, class information at the pixel level and The aforementioned Depth data and Correspondence 3D semantic segmentation process that generates a 3D semantic segmented image. Department Based on the aforementioned 3D semantic segmentation image, a parking space detection system is used to detect parking spaces. Department and Prepare Information processing system That is the case.
[0010] In one aspect of this disclosure, a vehicle is captured by a camera. around prescribed imagingRange image data is acquired, depth data for a range that at least partially overlaps with the imaging range of the image data is acquired, and class information and The aforementioned depth data are Correspondence used to generate a 3D semantic segmentation image, and a parking space is detected based on the 3D semantic segmentation image.
Brief Description of the Drawings
[0011] [Figure 1] It is a diagram for explaining the limitations of the parking support function. [Figure 2] It is a diagram for explaining an example of driving support by the parking support function. [Figure 3] It is a diagram for explaining the outline of the parking support function to which the technology of the present disclosure is applied. [Figure 4] It is a block diagram showing a configuration example of a vehicle control system. [Figure 5] It is a diagram showing an example of a sensing area. [Figure 6] It is a diagram for explaining a configuration example of the parking support control unit of the present disclosure. [Figure 7] It is a diagram for explaining a configuration example of the 3D semantic segmentation processing unit in FIG. 6. [Figure 8] It is a diagram for explaining 3D semantic segmentation information. [Figure 9] It is a diagram for explaining an example of detecting a parking space. [Figure 10] It is a diagram for explaining an example of detecting a parking space. [Figure 11] It is a flowchart for explaining 3D semantic segmentation information generation processing. [Figure 12] It is a flowchart for explaining parking support processing. [Figure 13] It is a flowchart for explaining parking space search mode processing. [Figure 14] It is a flowchart for explaining parking mode processing. [Figure 15]This diagram illustrates an example configuration of a general-purpose personal computer. [Modes for carrying out the invention]
[0012] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0013] The following describes the configurations for implementing this technology. The explanation will proceed in the following order. 1. Summary of this disclosure 2. Example of a vehicle control system configuration 3. Example configuration of a parking assistance control unit that realizes the parking assistance function of this disclosure 4. Examples of execution by software
[0014] <<1. Summary of this Disclosure>> This document outlines the technology that applies the technology disclosed herein to achieve rapid and smooth parking assistance that closely resembles human perception, regardless of the environment of the parking space.
[0015] While various forms of parking assistance functions, such as recognizing parking spaces and automatically parking, or guiding vehicles to the optimal parking route, have already been commercialized, all of them are subject to various limitations.
[0016] For example, the first parking assistance function, as shown by the dotted line in Figure 1, can park in parking spaces that have not been previously used, that is, even in parking spaces that have not been pre-registered in the vehicle. However, it cannot park in parking spaces that do not have white lines or where the white lines are faded.
[0017] Furthermore, for example, the second parking assistance function, as shown by the dashed line in Figure 1, allows parking in both parking spaces with and without white lines, regardless of whether or not there are white lines. However, it only functions for parking spaces that the vehicle has previously used, i.e., parking spaces that are pre-registered in the vehicle.
[0018] In other words, in order to realize the parking assistance function described above, conditions are set according to whether or not there are white lines on the target parking space and whether or not the parking space is pre-registered. In Figure 1, the horizontal axis shows the degree to which white lines are present (whether or not white lines are present), and the vertical axis shows the degree to which the action has been taken (whether or not it is pre-registered).
[0019] Therefore, in the parking assistance function described in this disclosure, as shown by the solid line in Figure 1, regardless of whether the target parking space is pre-registered in the vehicle (whether it has been parked there before) or whether there are white lines, the system appropriately recognizes the surrounding conditions such as the parking space and the direction of the parked vehicle (parallel or perpendicular), thereby achieving quick and smooth parking assistance, similar to parking operations performed by a human.
[0020] For example, consider parking assistance when a vehicle C1, equipped with sensors such as cameras Sc1-1 and Sc1-2 on the left and right sides of the main body, parks in a parking space SP1, as shown in the left side of Figure 2.
[0021] In Figure 2, the direction of the convex part indicated by the isosceles side of the isosceles triangle in the figure is assumed to be the front of vehicle C1.
[0022] As shown in the left part of Figure 2, vehicle C1 needs to pass in front of parking space SP1 at least once in order to detect the location of parking space SP1, which is an available space for parking, using sensor Sc1-1 mounted on its left side.
[0023] Furthermore, let's consider parking assistance when a vehicle C2, equipped with sensors Sc11-1 and Sc11-2, such as ultrasonic sensors, on the left and right sides of the main body as shown in the right side of Figure 2, parks in a parking space SP2.
[0024] In the case of the right side of Figure 2, the vehicle C2 needs to pass near parking space SP2 at least once in order to detect parking space SP1, which is an available parking space, using sensor Sc11-1 mounted on the left side of the vehicle C2.
[0025] In other words, as explained with reference to Figure 2, when considering parking assistance by equipping the main unit with cameras, ultrasonic sensors, etc., it is necessary to pass in front of or near the target parking space beforehand.
[0026] Therefore, in the example of parking assistance explained with reference to Figure 2, it is not possible to visually search for a suitable empty space, select the searched empty space as the target parking space, and then start the parking operation, as is the case when a human performs the parking operation.
[0027] Therefore, when parking is performed using the aforementioned parking assistance system, the vehicle will continue to patrol the parking lot until it passes in front of or beside an available parking space.
[0028] In this situation, it's possible that the person inside the vehicle might drive around searching for a parking space even in areas where they know there are no available spaces, despite being able to visually confirm that there are empty parking spaces.
[0029] As a result, it may take unnecessary time to complete parking, and the person inside the vehicle may not perceive it as a quick and smooth parking process, but rather as an awkward parking operation.
[0030] Therefore, in the parking assistance function applying the technology disclosed herein, the surrounding environment of the vehicle is grasped by object recognition processing using 3D (three-dimensional) semantic segmentation, and after identifying a parking space, the parking operation is performed.
[0031] For example, as shown in Figure 3, a sensor Sc31 that detects images and depth data is provided in front of the vehicle C11.
[0032] First, in vehicle C11, object recognition processing is performed using 3D semantic segmentation based on forward depth data detected by sensor Sc31 and images, thereby recognizing the surrounding environment. Next, based on the object recognition results, vehicle C11 searches for available parking spaces within a predetermined distance in front of vehicle C11 (for example, 15m to 30m in front).
[0033] Then, when an empty space is found from the search results in the area a predetermined distance in front of vehicle C11, the found empty space is recognized as the parking target parking space SP11, and a parking path for parking, for example, as shown by the thick solid line in the figure, is calculated, and vehicle C11 is controlled to move along the calculated parking path.
[0034] By implementing this type of parking operation, parking assistance is provided that is similar to human parking, such as visually searching for available spaces and performing the parking operation when a searched available space is recognized as a parking space.
[0035] As a result, it becomes possible to achieve quick and smooth parking assistance that looks natural to the human passenger.
[0036] Furthermore, object recognition processing using 3D semantic segmentation allows the surrounding environment to be recognized and parking spaces to be identified, making it possible to provide comfortable parking assistance regardless of the environment of the parking space, even in parking spaces located in various places.
[0037] <<2. Example of Vehicle Control System Configuration>> Figure 4 is a block diagram showing an example configuration of a vehicle control system 11, which is an example of a mobile device control system to which this technology is applied.
[0038] The vehicle control system 11 is installed in the vehicle 1 and performs processing related to driving assistance and autonomous driving of the vehicle 1.
[0039] The vehicle control system 11 includes a vehicle control ECU (Electronic Control Unit) 21, a communication unit 22, a map information storage unit 23, a GNSS (Global Navigation Satellite System) receiver unit 24, an external recognition sensor 25, an in-vehicle sensor 26, a vehicle sensor 27, a recording unit 28, a driving support / automatic driving control unit 29, a DMS (Driver Monitoring System) 30, an HMI (Human Machine Interface) 31, and a vehicle control unit 32.
[0040] The processor 21, communication unit 22, map information storage unit 23, GNSS receiver unit 24, external recognition sensor 25, in-vehicle sensor 26, vehicle sensor 27, recording unit 28, driving assistance / autonomous driving control unit 29, driver monitoring system (DMS) 30, human-machine interface (HMI) 31, and vehicle control unit 32 are interconnected and can communicate with each other via a communication network 41. The communication network 41 consists of an in-vehicle communication network or bus that conforms to digital bidirectional communication standards such as CAN (Controller Area Network), LIN (Local Interconnect Network), LAN (Local Area Network), FlexRay (registered trademark), and Ethernet (registered trademark). The communication network 41 may be used depending on the type of data being communicated; for example, CAN may be used for data related to vehicle control, and Ethernet may be used for large-capacity data. In addition, the various components of the vehicle control system 11 may be directly connected using wireless communication technologies intended for relatively short-range communication, such as Near Field Communication (NFC) or Bluetooth®, without going through the communication network 41.
[0041] In the following, when each part of the vehicle control system 11 communicates via the communication network 41, the description of the communication network 41 will be omitted. For example, when the processor 21 and the communication unit 22 communicate via the communication network 41, it will simply be described as the processor 21 and the communication unit 22 communicating.
[0042] The processor 21 is composed of various processors, such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit). The processor 21 controls the entire vehicle control system 11.
[0043] The communication unit 22 communicates with various devices inside and outside the vehicle, other vehicles, servers, base stations, etc., and transmits and receives various types of data. At this time, the communication unit 22 can communicate using multiple communication methods.
[0044] A brief explanation will be given regarding the external communication capabilities of the communication unit 22. The communication unit 22 communicates with servers (hereinafter referred to as "external servers") located on an external network via a base station or access point using wireless communication methods such as 5G (fifth-generation mobile communication system), LTE (Long Term Evolution), and DSRC (Dedicated Short Range Communications). The external network with which the communication unit 22 communicates is, for example, the internet, a cloud network, or a network specific to a carrier. The communication method used by the communication unit 22 to communicate with the external network is not particularly limited, as long as it is a wireless communication method capable of digital two-way communication at a predetermined communication speed and over a predetermined distance.
[0045] Furthermore, for example, the communication unit 22 can communicate with terminals located near the vehicle using P2P (Peer To Peer) technology. Terminals located near the vehicle include, for example, terminals worn by mobile objects moving at relatively low speeds, such as pedestrians and cyclists, terminals installed in fixed locations such as stores, or MTC (Machine Type Communication) terminals. In addition, the communication unit 22 can also perform V2X communication. V2X communication refers to communication between the vehicle and other entities, such as vehicle-to-vehicle communication with other vehicles, vehicle-to-infrastructure communication with roadside devices, etc., vehicle-to-home communication with homes, and vehicle-to-pedestrian communication with terminals carried by pedestrians, etc.
[0046] The communication unit 22 can, for example, receive externally a program to update the software that controls the operation of the vehicle control system 11. The communication unit 22 can also receive externally map information, traffic information, information about the vehicle 1's surroundings, etc. Furthermore, the communication unit 22 can transmit externally information about the vehicle 1 and information about the vehicle 1's surroundings, etc. Information about the vehicle 1 that the communication unit 22 transmits externally includes, for example, data indicating the status of the vehicle 1 and recognition results from the recognition unit 73. Furthermore, the communication unit 22 can also perform communications corresponding to vehicle emergency notification systems such as e-Call.
[0047] A brief overview of the communication capabilities of the communication unit 22 with the vehicle interior will be provided. The communication unit 22 can communicate with various devices in the vehicle, for example, using wireless communication. The communication unit 22 can communicate wirelessly with devices in the vehicle using communication methods that enable digital bidirectional communication at a predetermined or higher communication speed via wireless communication, such as Wi-Fi, Bluetooth (registered trademark), NFC, and WUSB (Wireless USB). Not limited to these, the communication unit 22 can also communicate with various devices in the vehicle using wired communication. For example, the communication unit 22 can communicate with various devices in the vehicle via wired communication through a cable connected to a connection terminal (not shown). The communication unit 22 can communicate with various devices in the vehicle using communication methods that enable digital bidirectional communication at a predetermined or higher communication speed via wired communication, such as USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface) (registered trademark), and MHL (Mobile High-definition Link).
[0048] Here, "devices inside the vehicle" refers to, for example, devices inside the vehicle that are not connected to the communication network 41. Examples of devices inside the vehicle include mobile devices and wearable devices carried by passengers such as the driver, and information devices that are brought into the vehicle and temporarily installed.
[0049] For example, the communication unit 22 receives electromagnetic waves transmitted by road traffic information communication systems (VICS (Vehicle Information and Communication System) (registered trademark)) such as radio beacons, optical beacons, and FM multiplex broadcasting.
[0050] The map information storage unit 23 stores either or both maps acquired from external sources and maps created by the vehicle 1. For example, the map information storage unit 23 stores high-precision 3D maps, global maps with lower precision than high-precision maps but covering a wide area, and so on.
[0051] High-precision maps include, for example, dynamic maps, point cloud maps, and vector maps. A dynamic map is, for example, a map consisting of four layers: dynamic information, semi-dynamic information, semi-static information, and static information, and is provided to vehicle 1 from an external server. A point cloud map is a map composed of point clouds (point cloud data). Here, a vector map refers to a map adapted for ADAS (Advanced Driver Assistance System) that maps traffic information such as the location of lanes and traffic lights to a point cloud map.
[0052] The point cloud map and vector map may be provided from, for example, an external server, or they may be created in the vehicle 1 as maps for matching with the local map described later, based on sensing results from radar 52, LiDAR 53, etc., and stored in the map information storage unit 23. In addition, if high-precision maps are provided from an external server, in order to reduce communication capacity, map data of, for example, several hundred square meters relating to the planned route that the vehicle 1 will travel will be acquired from the external server.
[0053] The GNSS receiver 24 receives GNSS signals from GNSS satellites and acquires the position information of the vehicle 1. The received GNSS signals are supplied to the driving assistance / automatic driving control unit 29. The GNSS receiver 24 is not limited to using GNSS signals; for example, it may acquire position information using beacons.
[0054] The external recognition sensor 25 is equipped with various sensors used to recognize the external conditions of the vehicle 1, and supplies sensor data from each sensor to various parts of the vehicle control system 11. The types and number of sensors equipped with the external recognition sensor 25 are arbitrary.
[0055] For example, the external recognition sensor 25 includes a camera 51, a radar 52, a LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) 53, and an ultrasonic sensor 54. However, the external recognition sensor 25 may also be configured to include one or more of the cameras 51, radar 52, LiDAR 53, and ultrasonic sensor 54. The number of cameras 51, radar 52, LiDAR 53, and ultrasonic sensor 54 is not particularly limited as long as it is a number that can be realistically installed in the vehicle 1. Furthermore, the types of sensors included in the external recognition sensor 25 are not limited to this example, and the external recognition sensor 25 may include other types of sensors. Examples of the sensing areas of each sensor included in the external recognition sensor 25 will be described later.
[0056] The shooting method of camera 51 is not particularly limited as long as it is a shooting method capable of distance measurement. For example, camera 51 can be a camera of various shooting methods such as a ToF (Time of Flight) camera, a stereo camera, a monocular camera, or an infrared camera, as needed. In addition, camera 51 may simply be for acquiring captured images (images taken) without regard to distance measurement.
[0057] Furthermore, for example, the external recognition sensor 25 may include an environmental sensor for detecting the environment relative to the vehicle 1. The environmental sensor is a sensor for detecting the environment such as weather, climate, and brightness, and may include various sensors such as a raindrop sensor, fog sensor, sunshine sensor, snow sensor, and illuminance sensor.
[0058] Furthermore, for example, the external recognition sensor 25 includes a microphone used for detecting sounds around the vehicle 1 and the location of sound sources.
[0059] The in-vehicle sensor 26 is equipped with various sensors for detecting information inside the vehicle and supplies sensor data from each sensor to various parts of the vehicle control system 11. The types and number of sensors equipped with the in-vehicle sensor 26 are not particularly limited as long as the number can realistically be installed in the vehicle 1.
[0060] For example, the in-vehicle sensor 26 can be equipped with one or more sensors from among a camera, radar, seat sensor, steering wheel sensor, microphone, and biosensor. The camera equipped in the in-vehicle sensor 26 can be a camera of various imaging types capable of distance measurement, such as a ToF camera, stereo camera, monocular camera, or infrared camera. However, it is not limited to these, and the camera equipped in the in-vehicle sensor 26 may simply be for acquiring images, regardless of distance measurement. The biosensor equipped in the in-vehicle sensor 26 is installed, for example, on the seat or steering wheel, and detects various biometric information of the driver or other passengers.
[0061] The vehicle sensor 27 is equipped with various sensors for detecting the state of the vehicle 1 and supplies sensor data from each sensor to various parts of the vehicle control system 11. The types and number of sensors equipped with the vehicle sensor 27 are not particularly limited as long as the number can realistically be installed on the vehicle 1.
[0062] For example, the vehicle sensor 27 includes a speed sensor, an acceleration sensor, an angular velocity sensor (gyro sensor), and an inertial measurement unit (IMU) that integrates them. For example, the vehicle sensor 27 includes a steering angle sensor for detecting the steering angle of the steering wheel, a yaw rate sensor, an accelerator sensor for detecting the amount of operation of the accelerator pedal, and a brake sensor for detecting the amount of operation of the brake pedal. For example, the vehicle sensor 27 includes a rotation sensor for detecting the rotation speed of the engine or motor, an air pressure sensor for detecting the air pressure of the tires, a slip ratio sensor for detecting the slip ratio of the tires, and a wheel speed sensor for detecting the rotation speed of the wheels. For example, the vehicle sensor 27 includes a battery sensor for detecting the remaining charge and temperature of the battery, and an impact sensor for detecting external impacts.
[0063] The recording unit 28 includes at least one of a non-volatile storage medium and a volatile storage medium, and stores data and programs. The recording unit 28 can be used as, for example, an EEPROM (Electrically Erasable Programmable Read Only Memory) and a RAM (Random Access Memory), and the storage medium can be a magnetic storage device such as an HDD (Hard Disk Drive), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The recording unit 28 records various programs and data used by each part of the vehicle control system 11. For example, the recording unit 28 includes an EDR (Event Data Recorder) and a DSSAD (Data Storage System for Automated Driving) to record information about vehicle 1 before and after an event such as an accident.
[0064] The driving assistance / autonomous driving control unit 29 controls the driving assistance and autonomous driving of the vehicle 1. For example, the driving assistance / autonomous driving control unit 29 includes an analysis unit 61, an action planning unit 62, and an operation control unit 63. The driving assistance / autonomous driving control unit 29 also implements the functions of the parking assistance control unit 201 (Figure 6), which realizes the parking assistance function described later in this disclosure.
[0065] The analysis unit 61 performs analysis processing on the vehicle 1 and its surroundings. The analysis unit 61 comprises a self-position estimation unit 71, a sensor fusion unit 72, and a recognition unit 73.
[0066] The self-position estimation unit 71 estimates the vehicle's position based on sensor data from the external recognition sensor 25 and a high-precision map stored in the map information storage unit 23. For example, the self-position estimation unit 71 generates a local map based on sensor data from the external recognition sensor 25 and estimates the vehicle's position by matching the local map with the high-precision map. The position of the vehicle 1 is based on, for example, the center of the rear wheel relative to the axle.
[0067] Local maps are, for example, three-dimensional high-precision maps created using technologies such as SLAM (Simultaneous Localization and Mapping), and occupancy grid maps (OGM). Three-dimensional high-precision maps are, for example, the point cloud maps mentioned above. Occupancy grid maps divide the three-dimensional or two-dimensional space around vehicle 1 into grids of a predetermined size and show the occupancy status of objects on a grid-by-grid basis. The occupancy status of objects is indicated, for example, by the presence or absence of an object or the probability of its existence. Local maps are also used, for example, in the detection and recognition processing of the external conditions of vehicle 1 by the recognition unit 73.
[0068] The self-position estimation unit 71 may estimate the vehicle 1's own position based on the GNSS signal and sensor data from the vehicle sensor 27.
[0069] The sensor fusion unit 72 performs sensor fusion processing to obtain new information by combining multiple different types of sensor data (for example, image data supplied from the camera 51 and sensor data supplied from the radar 52). Methods for combining different types of sensor data include integration, fusion, and union.
[0070] The recognition unit 73 performs a detection process to detect the external conditions of the vehicle 1, and a recognition process to recognize the external conditions of the vehicle 1.
[0071] For example, the recognition unit 73 performs detection and recognition processing of the external conditions of the vehicle 1 based on information from the external recognition sensor 25, information from the self-position estimation unit 71, information from the sensor fusion unit 72, etc.
[0072] Specifically, for example, the recognition unit 73 performs detection and recognition processing of objects around the vehicle 1. Object detection processing includes, for example, detecting the presence, size, shape, position, and movement of objects. Object recognition processing includes, for example, recognizing attributes such as the type of object or identifying a specific object. However, detection processing and recognition processing are not necessarily clearly separated and may overlap.
[0073] For example, the recognition unit 73 detects objects around the vehicle 1 by performing clustering, which classifies the point cloud based on sensor data from LiDAR 53 or radar 52 into clusters of points. This allows the presence, size, shape, and position of objects around the vehicle 1 to be detected.
[0074] For example, the recognition unit 73 detects the movement of objects around the vehicle 1 by performing tracking that follows the movement of clusters of points classified by clustering. This allows the velocity and direction of travel (movement vector) of objects around the vehicle 1 to be detected.
[0075] For example, the recognition unit 73 recognizes the types of objects around the vehicle 1 by performing object recognition processing such as semantic segmentation on the image data supplied from the camera 51.
[0076] Objects that can be detected or recognized by the recognition unit 73 include, for example, vehicles, people, bicycles, obstacles, structures, roads, traffic lights, traffic signs, and road markings.
[0077] For example, the recognition unit 73 can perform traffic rule recognition processing around the vehicle 1 based on the map stored in the map information storage unit 23, the self-position estimation result by the self-position estimation unit 71, and the recognition result of objects around the vehicle 1 by the recognition unit 73. Through this processing, the recognition unit 73 can recognize the location and status of traffic signals, the content of traffic signs and road markings, the content of traffic regulations, and the lanes that can be driven on.
[0078] For example, the recognition unit 73 can perform recognition processing of the environment surrounding the vehicle 1. The surrounding environment that the recognition unit 73 is intended to recognize may include weather, temperature, humidity, brightness, and road surface conditions.
[0079] The action planning unit 62 creates an action plan for vehicle 1. For example, the action planning unit 62 creates an action plan by performing route planning and route following processes.
[0080] Global path planning is the process of planning the general route from the start to the goal. This path planning also includes a process called local path planning, which involves generating a track that allows vehicle 1 to move safely and smoothly in its vicinity, taking into account the motion characteristics of vehicle 1 along the planned path.
[0081] Route following is the process of planning actions to safely and accurately travel the route planned by route planning within a planned time. The action planning unit 62 can, for example, calculate the target speed and target angular velocity of vehicle 1 based on the results of this route following process.
[0082] The motion control unit 63 controls the operation of the vehicle 1 in order to realize the action plan created by the action planning unit 62.
[0083] For example, the motion control unit 63 controls the steering control unit 81, brake control unit 82, and drive control unit 83, which are included in the vehicle control unit 32 described later, to perform acceleration / deceleration control and direction control so that the vehicle 1 moves along the trajectory calculated by the trajectory plan. For example, the motion control unit 63 performs cooperative control for the purpose of realizing ADAS functions such as collision avoidance or impact mitigation, follow driving, vehicle speed maintenance, collision warning for the vehicle, and lane departure warning for the vehicle. For example, the motion control unit 63 performs cooperative control for the purpose of autonomous driving, such as driving autonomously without driver operation.
[0084] The DMS30 performs driver authentication and driver status recognition based on sensor data from the in-vehicle sensors 26 and input data input to the HMI31, which will be described later. In this case, the driver status to be recognized by the DMS30 is expected to include, for example, physical condition, level of alertness, level of concentration, level of fatigue, gaze direction, level of intoxication, driving operation, and posture.
[0085] Furthermore, the DMS30 may perform authentication processing for passengers other than the driver and recognition processing for the status of said passengers. Also, for example, the DMS30 may perform recognition processing of the conditions inside the vehicle based on sensor data from the in-vehicle sensor 26. Examples of conditions inside the vehicle to be recognized include temperature, humidity, brightness, and odor.
[0086] HMI31 handles the input of various data and instructions, and presents various data to the driver and other users.
[0087] A brief explanation of data input by HMI31 is provided. HMI31 is equipped with an input device for human data input. HMI31 generates input signals based on data and instructions input by the input device and supplies them to each part of the vehicle control system 11. HMI31 is equipped with operators such as a touch panel, buttons, switches, and levers as input devices. However, HMI31 may also be equipped with input devices that allow information to be input by methods other than manual operation, such as voice or gestures. Furthermore, HMI31 may use external connected devices such as a remote control device using infrared or radio waves, or a mobile device or wearable device that corresponds to the operation of the vehicle control system 11, as input devices.
[0088] This section provides a brief overview of how HMI31 presents data. HMI31 generates visual, auditory, and tactile information for the occupant or those outside the vehicle. HMI31 also performs output control, managing the output, content, timing, and method of each of these generated pieces of information. As visual information, HMI31 generates and outputs information indicated by images and light, such as operation screens, vehicle status displays, warning displays, and monitor images showing the surroundings of vehicle 1. As auditory information, HMI31 generates and outputs information indicated by sound, such as voice guidance, warning sounds, and warning messages. Furthermore, as tactile information, HMI31 generates and outputs information that is perceived by the occupant's sense of touch through force, vibration, movement, etc.
[0089] As output devices for visual information output by HMI31, for example, a display device that presents visual information by displaying images itself, or a projector device that presents visual information by projecting images, can be applied. In addition to display devices with ordinary displays, the display device may also be a device that displays visual information within the passenger's field of view, such as a head-up display, a transparent display, or a wearable device with AR (Augmented Reality) functionality. Furthermore, HMI31 can also use display devices such as the navigation system, instrument panel, CMS (Camera Monitoring System), electronic mirrors, and lamps installed in the vehicle 1 as output devices for visual information output.
[0090] For HMI31, output devices that output auditory information can include, for example, audio speakers, headphones, and earphones.
[0091] As an output device for HMI31 to output tactile information, for example, a haptic element using haptic technology can be applied. The haptic element is installed in parts of the vehicle 1 that are in contact with by the occupant, such as the steering wheel and the seat.
[0092] The vehicle control unit 32 controls various parts of the vehicle 1. The vehicle control unit 32 includes a steering control unit 81, a brake control unit 82, a drive control unit 83, a body system control unit 84, a light control unit 85, and a horn control unit 86.
[0093] The steering control unit 81 detects and controls the state of the steering system of the vehicle 1. The steering system includes, for example, a steering mechanism with a steering wheel, an electric power steering system, etc. The steering control unit 81 includes, for example, a control unit such as an ECU that controls the steering system, an actuator that drives the steering system, etc.
[0094] The brake control unit 82 detects and controls the state of the brake system of the vehicle 1. The brake system includes, for example, a brake mechanism including a brake pedal, an ABS (Antilock Brake System), a regenerative braking mechanism, etc. The brake control unit 82 also includes, for example, a control unit such as an ECU that controls the brake system.
[0095] The drive control unit 83 detects and controls the state of the vehicle 1's drive system. The drive system includes, for example, an accelerator pedal, a drive force generating device for generating driving force such as an internal combustion engine or drive motor, and a drive force transmission mechanism for transmitting driving force to the wheels. The drive control unit 83 also includes, for example, a control unit such as an ECU that controls the drive system.
[0096] The body system control unit 84 detects and controls the state of the body system of the vehicle 1. The body system includes, for example, a keyless entry system, a smart key system, power window devices, power seats, an air conditioning system, airbags, seat belts, a shift lever, etc. The body system control unit 84 includes, for example, a control unit such as an ECU that controls the body system.
[0097] The light control unit 85 detects and controls the status of various lights on the vehicle 1. Examples of lights to be controlled include headlights, taillights, fog lights, turn signals, brake lights, projection lights, and bumper displays. The light control unit 85 includes a control unit such as an ECU that controls the lights.
[0098] The horn control unit 86 detects and controls the status of the car horn of the vehicle 1. The horn control unit 86 includes, for example, a control unit such as an ECU that controls the car horn.
[0099] Figure 5 shows examples of sensing areas using the camera 51, radar 52, LiDAR 53, and ultrasonic sensor 54 of the external recognition sensor 25 shown in Figure 4. In Figure 5, the vehicle 1 is schematically shown as viewed from above, with the left end being the front end of the vehicle 1 and the right end being the rear end of the vehicle 1.
[0100] Sensing regions 101F and 101B show examples of sensing regions of the ultrasonic sensor 54. Sensing region 101F covers the area around the front end of the vehicle 1 by multiple ultrasonic sensors 54. Sensing region 101B covers the area around the rear end of the vehicle 1 by multiple ultrasonic sensors 54.
[0101] The sensing results in sensing area 101F and sensing area 101B are used, for example, for parking assistance of vehicle 1.
[0102] Sensing areas 102F to 102B show examples of sensing areas for short-range or medium-range radar 52. Sensing area 102F covers a position further in front of vehicle 1 than sensing area 101F. Sensing area 102B covers a position further in rear of vehicle 1 than sensing area 101B. Sensing area 102L covers the rear periphery of the left side of vehicle 1. Sensing area 102R covers the rear periphery of the right side of vehicle 1.
[0103] The sensing results in sensing region 102F are used, for example, to detect vehicles or pedestrians in front of vehicle 1. The sensing results in sensing region 102B are used, for example, to prevent collisions behind vehicle 1. The sensing results in sensing regions 102L and 102R are used, for example, to detect objects in blind spots to the sides of vehicle 1.
[0104] Sensing areas 103F to 103B show examples of sensing areas by camera 51. Sensing area 103F covers a position further in front of vehicle 1 than sensing area 102F. Sensing area 103B covers a position further in rear of vehicle 1 than sensing area 102B. Sensing area 103L covers the periphery of the left side of vehicle 1. Sensing area 103R covers the periphery of the right side of vehicle 1.
[0105] The sensing results in sensing region 103F can be used, for example, for recognition of traffic lights and traffic signs, lane departure prevention support systems, and automatic headlight control systems. The sensing results in sensing region 103B can be used, for example, for parking assistance and surround view systems. The sensing results in sensing regions 103L and 103R can be used, for example, for surround view systems.
[0106] Sensing area 104 shows an example of the sensing area of LiDAR 53. Sensing area 104 covers a position further in front of vehicle 1 than sensing area 103F. On the other hand, sensing area 104 has a narrower range in the left-right direction than sensing area 103F.
[0107] The sensing results in the sensing region 104 can be used, for example, to detect objects such as surrounding vehicles.
[0108] Sensing area 105 shows an example of the sensing area of the long-range radar 52. Sensing area 105 covers a position further in front of vehicle 1 than sensing area 104. On the other hand, sensing area 105 has a narrower range in the left-right direction than sensing area 104.
[0109] The sensing results in sensing area 105 are used, for example, for ACC (Adaptive Cruise Control), emergency braking, collision avoidance, etc.
[0110] Furthermore, the sensing areas of the camera 51, radar 52, LiDAR 53, and ultrasonic sensor 54 included in the external recognition sensor 25 may take various configurations other than those shown in Figure 5. Specifically, the ultrasonic sensor 54 may also sense the sides of the vehicle 1, or the LiDAR 53 may be configured to sense the rear of the vehicle 1. Also, the installation positions of each sensor are not limited to the examples described above. In addition, there may be one or more sensors.
[0111] <<3. Example of the configuration of the parking assistance control unit that realizes the parking assistance function of this disclosure>> Next, with reference to Figure 6, an example configuration of the parking assistance control unit 201 that realizes the parking assistance function of this disclosure will be described.
[0112] The parking assistance control unit 201 is implemented by the driving assistance / autonomous driving control unit 29 in the vehicle control system 11 described above.
[0113] The parking assistance control unit 201 implements parking assistance functions based on image data, depth data (distance measurement results), and radar detection results supplied by cameras 202-1 to 202-q, ToF cameras 203-1 to 203-r, and radars 204-1 to 204-s.
[0114] Furthermore, unless there is a need to distinguish between cameras 202-1 to 202-q, ToF cameras 203-1 to 203-r, and radars 204-1 to 204-s, they shall simply be referred to as camera 202, ToF camera 203, and radar 204, and the other components shall be referred to similarly.
[0115] Camera 202 and ToF camera 203 are configured to correspond to camera 51 in Figure 4, and radar 204 is configured to correspond to radar 52 in Figure 4.
[0116] Furthermore, the parking assistance control unit 201 may implement the parking assistance function using not only the detection results of cameras 202-1 to 202-q, ToF cameras 203-1 to 203-r, and radars 204-1 to 204-s shown in Figure 6, but also the detection results of various configurations of the external recognition sensor 25 and vehicle sensor 27 shown in Figure 4.
[0117] The parking assistance control unit 201 associates the image data with the depth data on a pixel-by-pixel basis, based on the image data supplied from the camera 202, the ToF camera 203, and the radar 204, as well as the depth data.
[0118] The parking assistance control unit 201 performs 3D semantic segmentation processing using information including image data, depth data, and radar detection results, in addition to the pixel-level associated depth data, and generates a 3D semantic segmentation image that associates the depth data with the object recognition results at the pixel level.
[0119] The parking assistance control unit 201 searches for a parking space based on a 3D semantic segmentation image and assists the vehicle 1 in parking the vehicle in the searched parking space.
[0120] In this case, the parking assistance control unit 201 operates in two operating modes, a parking space search mode and a parking mode, to realize the parking assistance function.
[0121] Specifically, the parking assistance control unit 201 first operates in parking space search mode, searches for parking spaces based on 3D semantic segmentation images, and stores the search results.
[0122] Then, when a parking space is found, the parking assistance control unit 201 switches its operating mode to parking mode, plans a route to the found parking space, and controls the movement of the vehicle 1 so that parking is completed along the planned route.
[0123] This type of operation enables parking assistance, allowing the system to search for a parking space before performing the parking maneuver, just like when a human is parking. This makes it possible to achieve quick and smooth parking assistance.
[0124] More specifically, the parking assistance control unit 201 consists of an analysis unit 261, an action planning unit 262, and an operation control unit 263.
[0125] The analysis unit 261 has a configuration corresponding to the analysis unit 61 in Figure 4, the action planning unit 262 has a configuration corresponding to the action planning unit 62 in Figure 4, and the operation control unit 263 has a configuration corresponding to the operation control unit 63 in Figure 4.
[0126] The analysis unit 261 includes a self-position estimation unit 271, a sensor fusion unit 272, and a recognition unit 273.
[0127] The self-position estimation unit 271, the sensor fusion unit 272, and the recognition unit 273 have configurations corresponding to the self-position estimation unit 71, the sensor fusion unit 72, and the recognition unit 73 in Figure 4, respectively.
[0128] The self-position estimation unit 271 includes a SLAM processing unit 301 and an OMG storage unit 302 as functions for realizing automatic parking assistance processing.
[0129] The SLAM (Simultaneous Localization and Mapping) processing unit 301 simultaneously performs self-position estimation and surrounding map creation to realize the automatic parking assistance function. Specifically, the SLAM processing unit 301 creates a self-position estimation and a three-dimensional map of the surrounding area, for example as an Occupancy Grid Map (OGM), based on information (hereinafter also simply referred to as integrated information) that is integrated (fused or combined) from information supplied by multiple sensors from the sensor fusion unit 272, and stores it in the OGM storage unit 302.
[0130] The OMG storage unit 302 stores the OMG created by the SLAM processing unit 301 and supplies it to the action planning unit 262 as needed.
[0131] The recognition unit 273 includes an object detection unit 321, an object tracking unit 322, a 3D semantic segmentation processing unit 323, and a context awareness unit 324.
[0132] The object detection unit 321 detects an object based on integrated information supplied from the sensor fusion unit 272, for example, by detecting the presence, size, shape, position, and movement of the object. The object tracking unit 322 tracks the object detected by the object detection unit 321.
[0133] The 3D semantic segmentation processing unit 323 performs three-dimensional semantic segmentation (3D semantic segmentation) based on image data captured by the camera 202, depth data (distance measurement results) detected by the ToF camera 203, and detection results from the radar 204, and generates 3D semantic segmentation results. The detailed configuration of the 3D semantic segmentation processing unit 323 will be described later with reference to Figure 7.
[0134] The context awareness unit 324 consists of a recognition unit that has undergone machine learning, such as using a DNN (Deep Neural Network), and recognizes a situation (e.g., a parking space) from the relationships between objects based on the 3D semantic segmentation results.
[0135] More specifically, the context awareness unit 324 includes a parking space detection unit 324a, which detects parking spaces based on the relationships between objects, using the 3D semantic segmentation results.
[0136] The parking space detection unit 324a recognizes and detects parking spaces based on the interrelationships of multiple object recognition results, such as a space enclosed by a frame such as white lines large enough for a vehicle to park, a space large enough for a vehicle to park even without white lines but with wheel stops, and a space large enough for a vehicle to park between a vehicle and a support column, based on the 3D semantic segmentation results.
[0137] The action planning unit 262 includes a route planning unit 351, and when a parking space is detected by the recognition unit 273, it plans a route from the vehicle's current position to parking in the detected parking space. The operation control unit 263 controls the operation of the vehicle 1 in order to realize the action plan created by the action planning unit 262 to park in the parking space recognized by the recognition unit 273.
[0138] <Example configuration of the 3D semantic segmentation processing unit> Next, with reference to Figure 7, an example configuration of the 3D semantic segmentation processing unit 323 will be described.
[0139] The 3D semantic segmentation processing unit 323 includes a preprocessing unit 371, an image feature extraction unit 372, a monocular depth estimation unit 373, a 3D anchor grid generation unit 374, a dense fusion processing unit 375, a preprocessing unit 376, a point cloud feature extraction unit 377, a radar detection result feature extraction unit 378, and a type determination unit 379.
[0140] The preprocessing unit 371 applies predetermined preprocessing (such as contrast correction and edge enhancement) to the image data supplied in time series from the camera 202 and outputs it to the image feature extraction unit 372 and the type determination unit 379.
[0141] The image feature extraction unit 372 extracts image features from the pre-processed image data and outputs them to the monocular depth estimation unit 373, the point cloud feature extraction unit 377, and the type determination unit 379.
[0142] The monocular depth estimation unit 373 estimates the monocular depth (distance measurement image) based on image features and outputs it to the dense fusion processing unit 375 as dense depth data (image-based depth data). For example, the monocular depth estimation unit 373 estimates the monocular depth (distance measurement image) using distance information from a vanishing point in a single 2D image data to a feature point from which image features are extracted, and outputs it to the dense fusion processing unit 375 as dense depth data.
[0143] The 3D anchor grid generation unit 374 generates a 3D anchor grid, which is a grid of three-dimensional anchor positions based on the distance measurement results detected by the ToF camera 203, and outputs it to the dense fusion processing unit 375.
[0144] The dense fusion processing unit 375 fuses the 3D anchor grid supplied by the 3D anchor grid generation unit 374 with the dense depth supplied by the monocular depth estimation unit 373 to generate a dense fusion, which is then output to the point cloud feature extraction unit 377.
[0145] The preprocessing unit 376 performs preprocessing, such as noise reduction, on the point cloud data consisting of distance measurement results supplied from the ToF camera 203, and outputs it to the point cloud feature extraction unit 377 and the type determination unit 379.
[0146] The point cloud feature extraction unit 377 extracts point cloud features from the pre-processed point cloud data supplied by the pre-processing unit 376, based on the image features supplied by the image feature extraction unit 372 and the dense fusion supplied by the dense fusion processing unit 375, and outputs them to the classification determination unit 379.
[0147] The radar detection result feature extraction unit 378 extracts radar detection result features from the detection results of the radar 204 supplied by the radar 204 and outputs them to the type determination unit 379.
[0148] The type determination unit 379 associates depth data at the pixel level in the image data by matching the depth data at each pixel of the image data supplied by the preprocessing unit 371 based on the point cloud data (depth data) supplied by the preprocessing unit 376. In this case, the type determination unit 379 may, if necessary, generate depth data by combining the point cloud data and the radar detection result of the radar 204 (depth data (point cloud) based on the radar detection result) and associate it at the pixel level in the image data. This makes it possible to complement the ToF camera 203. That is, for example, even in scenes where the reliability of the ToF camera 203 falls below a predetermined threshold due to fog, etc., it is possible to measure distance without lowering the reliability below a predetermined threshold by using the radar detection result of the radar 204 which uses radio waves. In this case, it is desirable that the radar 204 be a so-called imaging radar with a high resolution comparable to that of the camera 202. Furthermore, by using the radar detection results from radar 204, the relative velocity to the object can be calculated, making it possible to determine whether or not the object is moving. For example, by adding pixel-level velocity information, it is possible to improve object recognition performance. More specifically, by implementing object recognition processing using six parameters (x, y, z, vx, vy, vz) for each pixel, it is possible to improve object recognition accuracy. Note that vx, vy, and vz are velocity information in the x, y, and z directions, respectively.
[0149] Furthermore, the type determination unit 379 is composed of a recognizer that performs 3D semantic segmentation processing using machine learning such as a DNN (Deep Neural Network), and performs object recognition processing on a pixel-by-pixel basis of the image data to identify the type (class) based on the image features supplied by the image feature extraction unit 372, the point cloud features supplied by the point cloud feature extraction unit 377, and the radar detection result features supplied by the radar detection result feature extraction unit 378.
[0150] In other words, the classification unit 379 associates depth data (x, y, z) at the pixel level, as shown in the grid of image P11, for example, as shown in Figure 8. Furthermore, the classification unit 379 determines the classification (class) (seg) by performing 3D semantic segmentation processing at the pixel level based on the depth data, image features, point cloud features, and radar detection result features. Then, the classification unit 379 sets 3D semantic segmentation information (x, y, z, seg) by associating the classification determination result with the depth data at the pixel level.
[0151] The type determination unit 379 generates a 3D semantic segmentation image from an image consisting of 3D semantic segmentation information (x, y, z, seg) set on a pixel-by-pixel basis. As a result, the 3D semantic segmentation image is an image in which regions for each type (class) are formed within the image.
[0152] The types (classes) recognized as objects by object recognition processing include, for example, roadways, sidewalks, pedestrians, cyclists and motorbike riders, vehicles, trucks, buses, motorbikes, bicycles, buildings, walls, guardrails, bridges, tunnels, poles, traffic signs, traffic signals, white lines, etc. For example, a pixel classified as a vehicle will have surrounding pixels similarly classified as vehicles, and the area they form as a whole will form an image that can be seen as a vehicle. Therefore, within the image, areas are formed for each type (class) classified at the pixel level.
[0153] Furthermore, we have described an example in which the type determination unit 379 determines the type by performing 3D semantic segmentation processing on a pixel-by-pixel basis, based on depth data, image features, point cloud features, and radar detection result features.
[0154] However, instead of depth data, 2D semantic segmentation may be performed using only 2D image data and image features to determine the type at the pixel level, and then 3D semantic segmentation may be achieved by associating these pixels with depth data. Alternatively, 3D semantic segmentation may be performed without using radar detection result features.
[0155] <Detection of parking spaces using context awareness (Part 1)> Next, we will explain the detection of parking spaces by the context awareness unit 324.
[0156] As mentioned above, 3D semantic segmentation information consists of pixel-level depth data and type in the captured image data.
[0157] Therefore, the context awareness unit 324 uses a process called context awareness processing to identify the relationships between objects based on 3D semantic segmentation information, using the results of type determination at the pixel level within the image captured of the surroundings, and recognizes the surrounding situation from the identified relationships between objects.
[0158] In this example, the context awareness unit 324 includes a parking space detection unit 324a. By controlling the parking space detection unit 324a, context awareness processing based on 3D semantic segmentation information is performed to detect parking spaces within the image based on the relationships between objects.
[0159] For example, in the case of image P31 in Figure 9, the parking space detection unit 324a recognizes the area indicated by the solid line as a parking space based on the pixel-level depth data and type, as well as the positional relationship between the support column Ch11 and the parked vehicle Ch12 within image P31, and the size and shape of the space.
[0160] Furthermore, for example, in the case of image P32 in Figure 9, the parking space detection unit 324a recognizes the area indicated by the solid line as a parking space based on the arrangement interval of multiple wheel stops Ch21 within image P32, as well as the size and shape of the space, based on the pixel-level depth data and type.
[0161] Furthermore, for example, in the case of image P33 in Figure 9, the parking space detection unit 324a recognizes the area indicated by a solid line as a parking space based on the pixel-level depth data and type, the spacing between the foldable wheel stops Ch31 and white lines Ch32 of the coin parking lot within image P33, and the size of the space.
[0162] Furthermore, for example, in the case of image P34 in Figure 9, the parking space detection unit 324a recognizes the area indicated by the solid line as a parking space based on the pixel-level depth data and type, the positional relationship between the vehicle Ch41 and the support pillar Ch42 for the multi-story parking within image P34, and the size and shape of the space.
[0163] Thus, the context-awareness unit 324 (specifically the parking space detection unit 324a) functions as a machine learning-based recognizer, such as a DNN (Deep Neural Network), to detect parking spaces based on the relationships between multiple recognition results, which are derived from the depth data and type (class) information included in the 3D semantic segmentation information.
[0164] Furthermore, the recognition device that implements the context awareness unit 324 (specifically the parking space detection unit 324a) using machine learning with a DNN and the recognition device that implements the type determination unit 379 using machine learning with a DNN may be different or the same.
[0165] Then, the parking space detection unit 324a of the context awareness unit 324 detects, with reference to Figure 9, any empty parking spaces where no vehicle is parked among the areas detected as parking spaces as target parking spaces for the parking assistance function.
[0166] <Detection of parking spaces using context awareness (Part 2)> The above has described methods for detecting parking spaces in parking lots, but now we will explain an example of detecting parking spaces on the street instead of in a parking lot.
[0167] For example, consider the case where an image P31, as shown in Figure 10, is captured by camera 202.
[0168] Image P31 is an image taken from vehicle C51 in the overhead view shown on the right side of Figure 10, with the upper part of the figure facing forward. In image P51, vehicles C61 through C64 are parked in a row on the left side of the road.
[0169] Each pixel in image P31 is assigned 3D semantic segmentation information consisting of depth data and type (class). Therefore, the parking space detection unit 324a of the context awareness unit 324 can recognize the distance from vehicle C51 to the right rear of vehicles C61 and C62 in image P31, as shown in the right part of the figure, and can recognize the presence of parking space SP51 from the difference in these distances.
[0170] More specifically, based on, for example, the distance DL between the right front of vehicle C62 and the white line, the distance to the right front of vehicle C62, the width DS of the visible right rear end of vehicle C61, and the distance to the right rear end of vehicle C61, the parking space detection unit 324a of the context awareness unit 324 can determine the size of the rectangular empty space indicated by the dotted line in the figure. Specifically, the size of the empty space is determined using both a parking space recognition process using an offline-learned feature extractor based on 3D semantic segmentation information and a parking space recognition process using distance information such as the distance DL and the width DS.
[0171] Therefore, when the parking space detection unit 324a of the context awareness unit 324 recognizes from the size of the identified empty space that it is large enough to park the vehicle C51, it recognizes the rectangular empty space shown by the dotted line in the figure as the parking space SP51.
[0172] Thus, based on an image like the one shown in Figure 10, image P31, and the corresponding 3D semantic segmentation information, the parking space detection unit 324a of the context awareness unit 324 can recognize the empty space between vehicles C61 and C62 as a parking space.
[0173] As a result, while it is difficult to recognize an empty space like a parking space SP51 approximately 15 to 30 meters ahead using only a 2D image P31, setting 3D semantic segmentation information makes it possible to properly detect the parking space SP51 even when only the right rear edge of the vehicle is visible.
[0174] <3D Semantic Segmentation Image Generation Process> Next, referring to the flowchart in Figure 11, we will explain the 3D semantic segmentation image generation process by which the 3D semantic segmentation processing unit 323 generates a 3D semantic segmentation image.
[0175] In step S11, the preprocessor 371 acquires the captured image data supplied from the camera 202.
[0176] In step S12, the 3D anchor grid generation unit 374 and the preprocessing unit 376 acquire depth data (distance measurement results) sensed by the ToF camera 203.
[0177] In step S13, the radar detection result feature extraction unit 378 acquires the detection result of the radar 204 as the radar detection result.
[0178] In step S14, the preprocessing unit 371 applies predetermined preprocessing (such as contrast correction and edge enhancement) to the image data supplied from the camera 202 and outputs it to the image feature extraction unit 372 and the type determination unit 379.
[0179] In step S15, the image feature extraction unit 372 extracts image features from the preprocessed image data and outputs them to the monocular depth estimation unit 373, the point cloud feature extraction unit 377, and the type determination unit 379.
[0180] In step S16, the monocular depth estimation unit 373 estimates the monocular depth (distance measurement image) based on the image features and outputs it to the dense fusion processing unit 375 as dense depth (depth data based on the image).
[0181] In step S17, the 3D anchor grid generation unit 374 generates a 3D anchor grid based on the depth data (distance measurement results) detected by the ToF camera 203, which grids the three-dimensional anchor positions, and outputs it to the dense fusion processing unit 375.
[0182] In step S18, the dense fusion processing unit 375 fuses the 3D anchor grid supplied by the 3D anchor grid generation unit 374 with the dense depth supplied by the monocular depth estimation unit 373 to generate dense fusion data, which is then output to the point cloud feature extraction unit 377.
[0183] In step S19, the preprocessing unit 376 performs preprocessing such as noise reduction on the point cloud data consisting of depth data (distance measurement results) supplied from the ToF camera 203, and outputs it to the point cloud feature extraction unit 377 and the type determination unit 379.
[0184] In step S20, the point cloud feature extraction unit 377 extracts point cloud features from the pre-processed point cloud data supplied by the pre-processing unit 376, based on the image features supplied by the image feature extraction unit 372 and the dense fusion data supplied by the dense fusion processing unit 375, and outputs them to the classification determination unit 379.
[0185] In step S21, the radar detection result feature extraction unit 378 extracts radar detection result features from the detection results of the radar 204 supplied by the radar 204 and outputs them to the type determination unit 379.
[0186] In step S22, the type determination unit 379 associates the depth data for each pixel of the image data supplied by the preprocessing unit 371 with the point cloud data supplied by the preprocessing unit 376, thereby associating the depth data on a pixel-by-pixel basis in the image data.
[0187] In step S23, the type determination unit 379 performs object recognition processing using 3D semantic segmentation processing based on pixel-level depth data, image features supplied by the image feature extraction unit 372, point cloud features supplied by the point cloud feature extraction unit 377, and radar detection result features supplied by the radar detection result feature extraction unit 378 to determine the type (class) of the image data on a pixel-by-pixel basis and generate 3D semantic segmentation information (x, y, z, seg).
[0188] In step S24, the type determination unit 379 generates a 3D semantic segmentation image by associating 3D semantic segmentation information (x, y, z, seg) with each pixel of the captured image data, and stores it in time series.
[0189] In step S25, it is determined whether or not a stop operation has been performed. If no stop operation has been performed, the process returns to step S11. In other words, the processes from steps S11 to S25 are repeated until a stop operation is performed.
[0190] Then, if a stop operation is instructed in step S25, the process terminates.
[0191] Through the above process, 3D semantic segmentation information is set on a pixel-by-pixel basis for each image data captured in time series. Furthermore, the process of generating an image consisting of 3D semantic segmentation information as a 3D semantic segmentation image is repeated and stored sequentially.
[0192] Then, 3D semantic segmentation images, which are generated sequentially over time, are used to realize the parking assistance processing described below.
[0193] Furthermore, as a result of the above processing, 3D semantic segmentation images are sequentially stored in chronological order, independently of other processes. This makes it possible to use the 3D semantic segmentation images generated in chronological order in other processes.
[0194] For example, the object detection unit 321 may detect objects based on a 3D semantic segmentation image. The object tracking unit 322 may track objects using the 3D semantic segmentation image based on the object detection result of the object detection unit 321. Furthermore, in the parking mode processing described later, it is also possible to use this to check whether the parking space has become unusable due to the detection of an obstacle or other reason, based on a 3D semantic segmentation image of the area around the parking space, while the vehicle is parking in the parking space.
[0195] <Parking assistance processing> Next, we will explain the parking assistance process with reference to the flowchart in Figure 12.
[0196] In step S41, the parking assistance control unit 201 determines whether or not to start the parking assistance process. The parking assistance control unit 201 may determine whether or not to start the parking assistance process based, for example, on whether or not an input device such as an HMI 31 that instructs the start of the parking assistance process has been operated.
[0197] Furthermore, the parking assistance control unit 201 may determine whether or not to start the parking assistance process based on whether or not information indicating the entrance to the parking lot has been detected from the 3D semantic segmentation image.
[0198] If it is determined in step S41 that the start of the parking assistance process has been instructed, the process proceeds to step S42.
[0199] In step S42, the parking assistance control unit 201 controls the HMI 31 to display information indicating that the parking assistance process has started.
[0200] In step S43, the parking assistance control unit 201 sets the operating mode to the parking space search mode and controls the HMI 31 to indicate that the current operating mode is the parking space search mode.
[0201] In step S44, the parking assistance control unit 201 controls the context awareness unit 324 in the recognition unit 273 of the analysis unit 261 to execute the parking space search mode processing, search for parking spaces, and register the parking spaces that result in the search. The details of the parking space search mode processing will be described later with reference to the flowchart in Figure 13.
[0202] In step S45, the parking assistance control unit 201 determines whether or not a parking space is registered in the context awareness unit 324. Note that "registered" does not mean that a parking space at home, etc., is registered in advance, but rather that it is temporarily registered in the context awareness unit 324.
[0203] If it is determined in step S45 that a parking space is available, the process proceeds to step S46.
[0204] In step S46, the parking assistance control unit 201 switches and sets the operating mode to parking mode and controls the HMI 31 to indicate that the operating mode is parking mode.
[0205] In step S47, the parking assistance control unit 201 controls the action planning unit 262 to execute the parking mode processing and park the vehicle. The details of the parking mode processing will be described later with reference to the flowchart in Figure 14.
[0206] In step S48, the parking assistance control unit 201 determines whether the parking assistance process has ended. More specifically, the parking assistance control unit 201 determines whether the parking assistance process has ended based on whether parking has been completed by the parking mode process, or whether it is indicated that there is no parking space and parking is not possible.
[0207] If it is determined in step S48 that the parking assistance process has been completed, the process proceeds to step S49.
[0208] In step S49, the parking assistance control unit 201 controls the HMI 31 to indicate that the parking assistance process has finished.
[0209] In step S50, the parking assistance control unit 201 determines whether or not a stopping operation has been performed to stop the operation of vehicle 1. If it is determined that no stopping operation has been performed on vehicle 1, the process returns to step S41.
[0210] On the other hand, if it is determined in step S45 that no available parking spaces are registered, the process proceeds to step S51.
[0211] In step S51, the parking assistance control unit 201 controls the HMI 31 to indicate that there is no parking space, and the process proceeds to step S48.
[0212] Furthermore, if the start of the parking assistance process is not instructed in step S41, the process proceeds to step S50.
[0213] In other words, if the start of the parking assistance process is not instructed and the vehicle 1 does not stop, the processes in steps S41 and S50 are repeated.
[0214] Then, if vehicle 1 is stopped in step S50, the process ends.
[0215] Through the above process, the parking space search mode searches for parking spaces based on 3D semantic segmentation information. If a parking space is found and registered, the vehicle 1 is parked in the registered parking space as a result of the search.
[0216] Furthermore, if no parking spaces are found and no available parking spaces are registered, a message will be displayed indicating that there are no parking spaces available and parking is not possible.
[0217] Through the above process, the system performs the parking space search mode processing followed by the parking mode processing. This allows the system to find a parking space in the surrounding area before performing the parking maneuver, just as a human would when parking, thus enabling quick and smooth parking assistance.
[0218] <Parking space search mode processing> Next, the parking space search mode processing will be explained with reference to the flowchart in Figure 13.
[0219] In step S71, the context awareness unit 324 acquires the 3D semantic segmentation image registered in the 3D semantic segmentation processing unit 323.
[0220] In step S72, the context awareness unit 324 controls the parking space detection unit 324a to search for parking spaces based on the 3D semantic segmentation image. Here, the parking space search process by the parking space detection unit 324a of the context awareness unit 324 is the context awareness process described, for example, with reference to Figures 9 and 10.
[0221] In step S73, the context awareness unit 324 registers (updates) the information of the searched parking space and its location, as a result of the parking space search conducted by the parking space detection unit 324a based on the 3D semantic segmentation image.
[0222] Through the above process, parking spaces are searched for using context-aware processing based on 3D semantic segmentation images and sequentially registered in the parking space detection unit 324a.
[0223] Furthermore, as long as the parking space search mode is active, the same process will be repeated. Therefore, if a previously searched parking space becomes unavailable due to, for example, an obstacle appearing or another vehicle parking there, the registered parking space information will be deleted and updated.
[0224] Similarly, even if a space is not currently being searched as a parking space, if a vehicle drives away from it during the parking space search mode, and the space is then searched for as a new parking space, it will be newly registered.
[0225] Furthermore, if multiple parking spaces are searched for, the location information of all parking spaces will be registered.
[0226] <Parking Mode Processing> Next, we will explain the parking mode process with reference to the flowchart in Figure 14.
[0227] In step S91, the action planning unit 262 reads the location information of the parking space registered in the parking space detection unit 324a. If location information for multiple parking spaces is registered, the location information for all of the parking spaces is read.
[0228] In step S92, the action planning unit 262 selects the parking space that is the shortest distance from the vehicle's position from the location information of the read parking spaces and sets it as the target parking space. At this time, there may be a step to confirm with the user whether it is OK to set the shortest distance parking space as the target parking space. Alternatively, multiple parking spaces may be displayed on a display located inside the vehicle, and the user may select the target parking space from among these multiple parking spaces.
[0229] In step S93, the action planning unit 262 controls the route planning unit 351 to plan the route to the target parking space as a parking route.
[0230] In step S94, the action planning unit 262 controls the HMI 31 to present the parking route to the target parking space, which has been planned by the route planning unit 351.
[0231] In step S95, the action planning unit 262 controls the motion control unit 263 to move the vehicle 1 along the parking path.
[0232] In step S96, the action planning unit 262 determines whether parking is complete or not. If it is determined in step S96 that parking is not complete, the process proceeds to step S97.
[0233] In step S97, the action planning unit 262 determines whether the target parking space is unavailable for parking, for example, based on the latest 3D semantic segmentation image. That is, it determines whether the target parking space has become unavailable for parking due to obstacles being found in the target parking space while moving along the parking path, or due to another vehicle entering the target parking space while moving along the parking path.
[0234] In step S96, if the target parking space is not unavailable, the process returns to step S94. That is, as long as the operation is performed along the parking route and the target parking space is not unavailable, the process from steps S94 to S97 is repeated until parking is completed.
[0235] Furthermore, if it is determined in step S97 that the target parking space is no longer available for parking, the process proceeds to step S98.
[0236] In step S98, the action planning unit 262 determines whether or not location information for other parking spaces exists in the location information for the parking space that has been read.
[0237] In step S98, if it is determined that the location information for another parking space exists in the location information for the parking space read, the process returns to step S92.
[0238] In other words, a new parking space is created, the target parking space and its parking route are reset, and the process in steps S94 to S97 is repeated.
[0239] Then, if it is determined in step S96 that parking is complete, the process proceeds to step S100.
[0240] In step S100, the action planning unit 262 controls the HMI 31 to display an image indicating that parking is complete.
[0241] Furthermore, if there are no other parking spaces available in step S98, the process proceeds to step S99.
[0242] In step S99, the action planning unit 262 controls the HMI 31 to display an image notifying the user that no parking space can be found and therefore parking is not possible.
[0243] Through the above process, the parking space closest to the vehicle among the registered parking space locations is set as the target parking space, a parking route is planned, and the vehicle operates along the parking route, thereby achieving automatic parking.
[0244] In this process, if, while the vehicle is attempting to park in the designated parking space, an obstacle is detected in the designated parking space, or if another vehicle parks there and parking becomes impossible, then, if location information for other parking spaces is registered, the parking space closest to the vehicle among the other available spaces will be reset as the designated parking space and parking will be performed.
[0245] Through the above series of processes, a 3D semantic segmentation image is generated by setting depth data (3D position) and object type (class) for each pixel in the image, based on the image captured by camera 202, depth data (distance measurement results) detected by ToF camera 203, and detection results by radar 204.
[0246] Furthermore, parking spaces are detected and their location information is registered through context-awareness processing based on 3D semantic segmentation images.
[0247] Then, based on the registered parking space location information, a parking route is set and the actions required for parking are controlled.
[0248] This allows for the combination of image features with 3D information such as point cloud features and radar detection result features, enabling the determination of the type (class) at the pixel level, thus achieving more accurate type determination.
[0249] Furthermore, context-awareness processing using 3D semantic segmentation images with high-precision type determination makes it possible to identify parking spaces, enabling appropriate searching of various parking spaces regardless of the environment of the parking space. In addition, since the parking space can be set using images captured by the vehicle's camera before setting the parking route, it is possible to set the parking route without passing near the parking space.
[0250] As a result, it becomes possible to identify a parking space based on information in the images captured by the camera, set a parking route, and then park the vehicle, enabling quick and smooth parking assistance similar to that performed by a human.
[0251] <<4. Example of execution by software>> Incidentally, the series of processes described above can be executed by hardware, but they can also be executed by software. When the series of processes are executed by software, the programs that make up the software are installed from a storage medium onto a computer that has dedicated hardware built in, or onto a general-purpose computer that can perform various functions by installing various programs.
[0252] Figure 15 shows an example of a general-purpose computer configuration. This personal computer has a built-in CPU (Central Processing Unit) 1001. An input / output interface 1005 is connected to the CPU 1001 via a bus 1004. A ROM (Read Only Memory) 1002 and a RAM (Random Access Memory) 1003 are connected to the bus 1004.
[0253] The input / output interface 1005 is connected to an input unit 1006 consisting of input devices such as a keyboard and mouse for the user to input operation commands, an output unit 1007 that outputs images of the processing operation screen and processing results to a display device, a storage unit 1008 consisting of a hard disk drive for storing programs and various data, and a communication unit 1009 consisting of a LAN (Local Area Network) adapter for performing communication processing via a network such as the Internet. In addition, a drive 1010 is connected to removable storage media 1011 such as magnetic disks (including flexible disks), optical disks (including CD-ROMs (Compact Disc-Read Only Memory) and DVDs (Digital Versatile Discs)), magneto-optical disks (including MDs (Mini Discs)), or semiconductor memory.
[0254] The CPU 1001 reads programs stored in the ROM 1002, or from removable storage media 1011 such as magnetic disks, optical disks, magneto-optical disks, or semiconductor memory, and installs them into the storage unit 1008. The CPU 1001 then executes various processes according to the programs loaded from the storage unit 1008 into the RAM 1003. The RAM 1003 also stores data necessary for the CPU 1001 to execute various processes as appropriate.
[0255] In a computer configured as described above, the CPU 1001 loads, for example, a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004, and executes it, thereby performing the series of processes described above.
[0256] The program executed by the computer (CPU 1001) can be provided by recording it on a removable storage medium 1011, such as a packaged media. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.
[0257] In a computer, a program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting a removable storage medium 1011 into the drive 1010. Alternatively, a program can be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Furthermore, programs can be pre-installed in the ROM 1002 or the storage unit 1008.
[0258] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.
[0259] Furthermore, the CPU 1001 in Figure 15 implements the functions of the parking assistance control unit 201 in Figure 6.
[0260] Furthermore, in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device in which multiple modules are housed in one enclosure, are both considered systems.
[0261] The embodiments described herein are not limited to those described above, and various modifications are possible without departing from the gist of this disclosure.
[0262] For example, this disclosure can take the form of cloud computing, in which a single function is shared and processed collaboratively by multiple devices over a network.
[0263] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.
[0264] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.
[0265] Furthermore, this disclosure can also be structured as follows:
[0266] <1> An image acquisition unit that acquires image data within a predetermined distance from the vehicle, captured by a camera, A depth data acquisition unit acquires depth data in a range that overlaps with at least a portion of the imaging range of the aforementioned image data. A 3D semantic segmentation processing unit generates a 3D semantic segmentation image in which class information and depth data are set for each pixel of the aforementioned image data. A parking space detection unit detects parking spaces based on the aforementioned 3D semantic segmentation image. An information processing device equipped with the following features. <2> The aforementioned class information is information classified by semantic segmentation processing based on the aforementioned image data. <1> The information processing device described above. <3> The 3D semantic segmentation processing unit generates depth data for each pixel based on the image data and the depth data, associates the depth data with each pixel of the 2D semantic segmentation image generated based on the image data, and generates a 3D semantic segmentation image. <1> The information processing apparatus according to <1>. <4> The parking space detection unit further determines whether the detected parking space is available for parking. <1> The information processing apparatus according to <1>. <5> The parking space detection unit further determines whether the detected parking space is available for parking based on the size of the vehicle. <4> The information processing apparatus according to <4>. <6> The parking space detection unit further determines whether the detected parking space is available for parking based on the distances between a plurality of objects existing within the imaging range of the image data. <4> The information processing apparatus according to <4>. <7> The 3D semantic segmentation processing unit performs 3D semantic segmentation processing based on time-series image data within a range of a predetermined distance from the vehicle and time-series depth data within a range that at least partially overlaps with the imaging range of the image data. The parking space detection unit estimates the movement of an object existing within the imaging range of the image data based on the generated time-series 3D semantic segmentation image, and updates information regarding a parking available space where the vehicle can park based on the estimated movement of the object existing within the imaging range of the image data. <1> The information processing apparatus according to <1>. <8> The 3D semantic segmentation processing unit generates depth data for each pixel based on the image data and the depth data, performs 3D semantic segmentation processing on the image data in association with the depth data for each pixel, and generates the 3D semantic segmentation image composed of a plurality of regions by setting class information for each pixel. The information processing apparatus according to <7>. <9> The depth data is generated using an ultrasonic sensor, LiDAR, optical distance measurement sensor, stereo camera, monocular camera, or infrared camera. The information processing apparatus according to <1>. <10> The vehicle further includes a radar that acquires speed information of an object moving within a range of a predetermined distance from the vehicle The information processing apparatus according to <1>. <11> The vehicle further includes a parking space search unit that registers the parking space where parking is possible The information processing apparatus according to <4>. <12> When there are a plurality of parking spaces that can be parked registered in the parking space search unit, among the plurality of parking spaces where parking is possible, the nearest parking space where parking is possible is set as the target parking space, a route to the target parking space is planned, and a parking control unit that controls the operation of the vehicle related to parking along the planned route is further provided. The information processing apparatus according to <11>. <13> Among one or more parking spaces that can be parked registered in the parking space search unit, the parking space that can be parked selected by the user is set as the target parking space, a route to the target parking space is planned, and a parking control unit that controls the operation of the vehicle related to parking along the planned route is further provided. The information processing apparatus according to <11>. <14> Performing image acquisition processing for acquiring image data within a range of a predetermined distance from the vehicle captured by a camera, Performing depth data acquisition processing for acquiring depth data of a range that at least partially overlaps with the imaging range of the image data, Performing 3D semantic segmentation processing for generating a 3D semantic segmentation image in which class information and depth data are set for each pixel of the image data, Performing parking space detection processing for detecting a parking space based on the 3D semantic segmentation image An information processing method including: <15> An image acquisition unit that acquires image data within a predetermined distance from the vehicle, captured by a camera, A depth data acquisition unit acquires depth data in a range that overlaps with at least a portion of the imaging range of the aforementioned image data. A 3D semantic segmentation processing unit generates a 3D semantic segmentation image in which class information and depth data are set for each pixel of the aforementioned image data. A parking space detection unit detects parking spaces based on the aforementioned 3D semantic segmentation image. A program that makes a computer function. [Explanation of symbols]
[0267] 201 Parking Assist Control Unit, 202, 202-1 to 202-q, 203, 203-1, 203-r ToF Camera, 204, 204-1 to 204-s, 261 Analysis Unit, 262 Action Planning Unit, 271 Self-Position Estimation Unit, 272 Sensor Fusion Unit, 273 Recognition Unit, 301 SLAM Processing Unit, 302 OGM Storage Unit, 321 Object Detection Unit, 322 Object Tracking Unit, 323 3D Semantic Segmentation Processing Unit, 324 Context Awareness Unit, 324a Parking Space Detection Unit, 351 Path Planning Unit, 371 Preprocessing Unit, 372 Image Feature Extraction Unit, 373 Monocular Depth Estimation Unit, 374 3D Anchor Grid Generation Unit, 375 Dense Fusion Processing Unit, 376 Preprocessing unit, 377 Point cloud feature extraction unit, 378 Radar detection result feature extraction unit, 379 Type determination unit
Claims
1. An image acquisition unit that acquires image data of a predetermined imaging range around the vehicle, captured by a camera, A depth data acquisition unit acquires depth data in a range that overlaps with at least a portion of the imaging range of the aforementioned image data. A 3D semantic segmentation processing unit generates a 3D semantic segmentation image in which class information and depth data are associated at the pixel level, A parking space detection unit detects multiple parking spaces based on the 3D semantic segmentation image and determines whether the detected parking spaces are available for parking. The system includes a parking control unit that sets a first parking space, which is one of a plurality of available parking spaces detected by the parking space detection unit, as a first target parking space, and plans a route to the first target parking space to control the operation of the vehicle, The parking control unit continuously confirms, based on the 3D semantic segmentation image, that the first target parking space is available for parking until the vehicle is parked. When it becomes unavailable for parking, it sets a different parking space from the first target parking space among the plurality of available parking spaces as the second target parking space, plans a route to the second target parking space, and controls the vehicle's movements related to parking along the planned route. Information processing device.
2. The aforementioned class information is information classified by semantic segmentation processing based on the aforementioned image data. The information processing apparatus according to claim 1.
3. The 3D semantic segmentation processing unit generates pixel-level depth data of the image data based on the image data and the depth data, associates the depth data with each pixel of the 2D semantic segmentation image generated based on the image data, and generates the 3D semantic segmentation image. The information processing apparatus according to claim 1.
4. The parking space detection unit further determines whether the detected parking space is suitable for parking based on the size of the vehicle. The information processing apparatus according to claim 1.
5. The parking space detection unit further determines whether the detected parking space is available for parking based on the distances between multiple objects that exist within the imaging range of the image data. The information processing apparatus according to claim 1.
6. The 3D semantic segmentation processing unit performs 3D semantic segmentation processing based on the time-series image data within a predetermined distance range from the vehicle and the time-series depth data in a range that overlaps with at least a portion of the imaging range of the image data. The parking space detection unit estimates the movement of objects within the imaging range of the image data based on the generated time-series 3D semantic segmentation image, and updates information regarding parking spaces where the vehicle can park based on the estimated movement of objects within the imaging range of the image data. The information processing apparatus according to claim 1.
7. The 3D semantic segmentation processing unit generates pixel-level depth data of the image data based on the image data and depth data, performs the 3D semantic segmentation processing on the image data in association with the pixel-level depth data of the image data, and generates a 3D semantic segmentation image consisting of multiple regions by setting the class information on a pixel-by-pixel basis of the image data. The information processing apparatus according to claim 6.
8. The aforementioned depth data is generated using an ultrasonic sensor, LiDAR, optical distance sensor, stereo camera, monocular camera, or infrared camera. The information processing apparatus according to claim 1.
9. The radar further comprises a radar that acquires velocity information of objects moving within a range that overlaps, at least partially, with the imaging range of the aforementioned image data. The information processing apparatus according to claim 1.
10. The system further comprises a parking space registration unit for registering the aforementioned available parking spaces. The information processing apparatus according to claim 1.
11. If there are multiple available parking spaces registered in the parking space registration unit, the parking control unit sets the nearest available parking space from among the multiple available parking spaces as the first target parking space or the second target parking space, plans a route to the first target parking space or the second target parking space, and controls the operation of the vehicle related to parking along the planned route. The information processing apparatus according to claim 10.
12. The parking control unit continuously confirms, based on the 3D semantic segmentation image, that the first target parking space is available for parking until the vehicle is parked. When it becomes unavailable for parking, the parking control unit sets the next nearest available parking space from among the multiple available parking spaces registered in the parking space registration unit as the new second target parking space, plans a route to the new second target parking space, and controls the vehicle's operation for parking along the planned route. The information processing apparatus according to claim 11.
13. The parking control unit sets one or more available parking spaces registered in the parking space registration unit, selected by the user, as the first target parking space, plans a route to the first target parking space, and controls the operation of the vehicle related to parking along the planned route. The information processing apparatus according to claim 10.
14. The imaging range of the aforementioned image data is the area in front of the vehicle. The information processing apparatus according to claim 1.
15. The system outputs information indicating the state of the vehicle or the surrounding conditions as visual, auditory, or tactile information. The information processing apparatus according to claim 1.
16. An image acquisition unit that acquires image data of a predetermined imaging range around the vehicle, captured by a camera, A depth data acquisition unit acquires depth data in a range that overlaps with at least a portion of the imaging range of the aforementioned image data. A 3D semantic segmentation processing unit generates a 3D semantic segmentation image in which class information and depth data are associated at the pixel level, A parking space detection unit detects multiple parking spaces based on the 3D semantic segmentation image and determines whether the detected parking spaces are available for parking. The system includes a parking control unit that sets a first parking space, which is one of a plurality of available parking spaces detected by the parking space detection unit, as a first target parking space, and plans a route to the first target parking space to control the operation of the vehicle, The parking control unit continuously confirms, based on the 3D semantic segmentation image, that the first target parking space is available for parking until the vehicle is parked. When it becomes unavailable for parking, it sets a different parking space from the first target parking space among the plurality of available parking spaces as the second target parking space, plans a route to the second target parking space, and controls the vehicle's movements related to parking along the planned route. Information processing system.
17. If there are multiple available parking spaces, the parking control unit sets the nearest available parking space from among the multiple available parking spaces as the first target parking space, plans a route to the first target parking space, and controls the operation of the vehicle related to parking along the planned route. The information processing system according to claim 16.
18. The vehicle further comprises a first external recognition sensor and a second external recognition sensor mounted on the aforementioned vehicle. The image acquisition unit acquires the image data based on the sensor data of the first external recognition sensor. The depth data acquisition unit acquires the depth data based on the sensor data from the second external recognition sensor. The information processing system according to claim 17.
19. An image acquisition unit that acquires image data of a predetermined imaging range around the vehicle, captured by a camera, A depth data acquisition unit acquires depth data in a range that overlaps with at least a portion of the imaging range of the aforementioned image data. A 3D semantic segmentation processing unit generates a 3D semantic segmentation image in which class information and depth data are associated at the pixel level, A parking space detection unit detects multiple parking spaces based on the 3D semantic segmentation image and determines whether the detected parking spaces are available for parking. The first parking space, which is one of the plurality of available parking spaces detected by the parking space detection unit, is set as the first target parking space, and the computer is made to function as a parking control unit that plans the route to the first target parking space and controls the operation of the vehicle. The parking control unit continuously confirms, based on the 3D semantic segmentation image, that the first target parking space is available for parking until the vehicle is parked. When it becomes unavailable for parking, it sets a different parking space from the first target parking space among the plurality of available parking spaces as the second target parking space, plans a route to the second target parking space, and controls the vehicle's movements related to parking along the planned route. program.
Citation Information
Patent Citations
On-vehicle parallel parking assistant device and program for on-vehicle parallel parking assistant device
JP2009220592A
Parking space detection apparatus, method, and image processing apparatus
JP2017111803A
Method and device for detecting parking area using semantic segmentation in automatic parking system
JP2020126636A
Multi-network-based path generation for vehicle parking
US20190291720A1