Information processing device, information processing method, and program

The information processing device addresses the increased load on object recognition in sensor fusion systems by detecting object regions and associating them with camera data, thereby enhancing processing efficiency and accuracy.

JP7676407B2Active Publication Date: 2025-05-14SONY SEMICON SOLUTIONS CORP
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2022537913
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-21
Filing Date
2021-07-07
Publication Date
2025-05-14
Estimated Expiration
2041-07-07

AI Technical Summary

Technical Problem

The increased load on object recognition processing due to the need to process data from multiple sensors in sensor fusion systems, particularly in associating point cloud data from LiDAR with camera images.

Method used

An information processing device that detects an object region indicating the azimuth and elevation direction of objects within the sensing range of a distance sensor, and associates this information with data from a camera whose photographing range overlaps with the sensing range, thereby reducing the processing load.

Benefits of technology

The proposed solution reduces the processing load for object recognition by efficiently associating object regions with camera image information, thereby improving the speed and accuracy of object recognition in sensor fusion systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676407000001
    Figure 0007676407000001
  • Figure 0007676407000002
    Figure 0007676407000002
  • Figure 0007676407000003
    Figure 0007676407000003
Patent Text Reader

Abstract

The present invention relates to an information processing device, an information processing method, and a program that make it possible to reduce the load of object recognition that uses sensor fusion. The information processing device is provided with an object region detection unit that detects, on the basis of three-dimensional data indicating directions and distances of measurement points measured by a ranging sensor, an object region indicating a range of azimuth directions and elevation directions within which an object is present in the sensing range of the ranging sensor, and associates the object region with information in an image, imaged by a camera, in which at least a portion of the imaging range overlaps with said sensing range. The present invention can be applied, for example, to a system that carries out object recognition.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present technology relates to an information processing device, an information processing method, and a program, and in particular to an information processing device, an information processing method, and a program suitable for use in the case of performing object recognition using sensor fusion. [Background technology]

[0002] In recent years, there has been active development of technology for recognizing objects around a vehicle using sensor fusion technology that combines multiple types of sensors, such as cameras and LiDAR (Light Detection and Ranging, laser radar), to obtain new information (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2005-284471 A Summary of the Invention [Problem to be solved by the invention]

[0004] However, when using sensor fusion, the load on object recognition increases because data from multiple sensors must be processed. For example, the load on the process of associating each measurement point in the point cloud data acquired by the LiDAR with a position in the image captured by the camera increases.

[0005] The present technology has been made in consideration of such circumstances, and is intended to reduce the load of object recognition using sensor fusion. [Means for solving the problem]

[0006] An information processing device according to one aspect of the present technology includes an object area detection unit that detects an object area indicating the range in azimuth and elevation angles in which an object is present within the sensing range of a ranging sensor based on three-dimensional data indicating the direction and distance of each measurement point measured by the ranging sensor, and associates the object area with information in an image captured by a camera whose capturing range overlaps at least a portion of the sensing range.

[0007] An information processing method according to one aspect of the present technology detects an object area indicating the azimuth and elevation ranges in which an object exists within the sensing range of a ranging sensor based on three-dimensional data indicating the direction and distance of each measurement point measured by the ranging sensor, and corresponds the object area to information in an image captured by a camera whose capturing range overlaps at least a portion of the sensing range.

[0008] A program according to one aspect of the present technology causes a computer to perform a process of detecting an object area indicating the range in the azimuth and elevation directions in which an object is present within the sensing range of a ranging sensor, based on three-dimensional data indicating the direction and distance of each measurement point measured by the ranging sensor, and associating the object area with information in an image captured by a camera whose capturing range overlaps at least a portion of the sensing range.

[0009] In one aspect of the present technology, an object area indicating the azimuth and elevation ranges in which an object exists within the sensing range of the ranging sensor is detected based on three-dimensional data indicating the direction and distance of each measurement point measured by the ranging sensor, and the object area is associated with information in an image captured by a camera whose imaging range overlaps at least a portion of the sensing range. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a vehicle control system. [Diagram 2] FIG. 2 is a diagram illustrating an example of a sensing region. [Diagram 3]1 is a block diagram showing an embodiment of an information processing system to which the present technology is applied. [Figure 4] 11A and 11B are diagrams for comparing methods of associating point cloud data with a captured image. [Diagram 5] 11 is a flowchart illustrating an object recognition process. [Figure 6] 1A to 1C are diagrams illustrating examples of sensing ranges in the mounting angle and elevation angle directions of a LiDAR. [Figure 7] FIG. 1 is a diagram showing an example of imaging of point cloud data. [Figure 8] FIG. 1 is a diagram showing an example of point cloud data when the LiDAR is scanned at equal intervals in the elevation angle direction. [Figure 9] 1 is a graph for explaining a first example of a LiDAR scanning method according to the present technology. [Figure 10] 1A to 1C are diagrams illustrating an example of point cloud data generated by a first example of the LiDAR scanning method of the present technology. [Figure 11] 11 is a diagram showing an example of point cloud data generated by a second example of the LiDAR scanning method of the present technology. [Figure 12] 3A to 3C are schematic diagrams showing examples of a virtual plane, a unit region, and an object region. [Figure 13] FIG. 13 is a diagram for explaining a method of detecting an object region. [Figure 14] FIG. 13 is a diagram for explaining a method of detecting an object region. [Figure 15] 10 is a schematic diagram showing an example in which a captured image is associated with an object region; FIG. [Figure 16] 10 is a schematic diagram showing an example in which a captured image is associated with an object region; FIG. [Figure 17] 13 is a diagram showing an example of the detection result of an object region when the upper limit of the number of detected object regions in a unit region is set to four. FIG. [Figure 18] FIG. 2 is a schematic diagram showing an example of a captured image. [Figure 19] 10 is a schematic diagram showing an example in which a captured image is associated with an object region; FIG. [Figure 20] 10A and 10B are schematic diagrams showing examples of detection results of an object region. [Figure 21] FIG. 13 is a schematic diagram showing an example of a recognition range. [Figure 22] FIG. 13 is a schematic diagram showing an example of an object recognition result. [Diagram 23] FIG. 13 is a schematic diagram showing a first example of output information. [Figure 24] FIG. 11 is a diagram showing a second example of output information. [Diagram 25] FIG. 13 is a schematic diagram showing a third example of output information. [Figure 26] 1A and 1B are schematic diagrams showing examples of a captured image and a recognition range. [Figure 27] 11 is a graph showing the relationship between the number of lines in a captured image included in a recognition range and the processing time required for object recognition. [Figure 28] FIG. 13 is a schematic diagram showing an example of setting a plurality of recognition ranges. [Figure 29] FIG. 1 is a block diagram illustrating an example of the configuration of a computer. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, an embodiment of the present technology will be described in the following order. 1. Example of vehicle control system configuration 2. Embodiment 3. Variations 4.Other

[0012] <<1. Example of vehicle control system configuration>> FIG. 1 is a block diagram showing an example of the configuration of a vehicle control system 11, which is an example of a mobility device control system to which the present technology is applied.

[0013] The vehicle control system 11 is provided in the vehicle 1 and performs processing related to driving assistance and automatic driving of the vehicle 1.

[0014] The vehicle control system 11 includes a processor 21, a communication unit 22, a map information storage unit 23, a GNSS (Global Navigation Satellite System) receiving unit 24, an external recognition sensor 25, an in-vehicle sensor 26, a vehicle sensor 27, a recording unit 28, a driving assistance / automatic driving control unit 29, a DMS (Driver Monitoring System) 30, an HMI (Human Machine Interface) 31, and a vehicle control unit 32.

[0015] The processor 21, the communication unit 22, the map information storage unit 23, the GNSS receiving unit 24, the external recognition sensor 25, the in-vehicle sensor 26, the vehicle sensor 27, the recording unit 28, the driving support / automatic driving control unit 29, the driver monitoring system (DMS) 30, the human machine interface (HMI) 31, and the vehicle control unit 32 are connected to each other via a communication network 41. The communication network 41 is configured by an in-vehicle communication network or bus conforming to any standard such as CAN (Controller Area Network), LIN (Local Interconnect Network), LAN (Local Area Network), FlexRay (registered trademark), Ethernet (registered trademark), etc. Note that each unit of the vehicle control system 11 may be directly connected to each other by, for example, near field communication (NFC (Near Field Communication)) or Bluetooth (registered trademark) without going through the communication network 41.

[0016] In the following description, when each unit of the vehicle control system 11 communicates with each other via the communication network 41, the description of the communication network 41 will be omitted. For example, when the processor 21 and the communication unit 22 communicate with each other via the communication network 41, it will be simply described that the processor 21 and the communication unit 22 communicate with each other.

[0017] The processor 21 is configured by various processors such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ECU (Electronic Control Unit), etc. The processor 21 controls the vehicle control system 11 as a whole.

[0018] The communication unit 22 communicates with various devices inside and outside the vehicle, other vehicles, servers, base stations, etc., and transmits and receives various data. As communication with the outside of the vehicle, for example, the communication unit 22 receives from the outside a program for updating software that controls the operation of the vehicle control system 11, map information, traffic information, information about the surroundings of the vehicle 1, etc. For example, the communication unit 22 transmits information about the vehicle 1 (for example, data indicating the state of the vehicle 1, recognition results by the recognition unit 73, etc.), information about the surroundings of the vehicle 1, etc., to the outside. For example, the communication unit 22 performs communication corresponding to a vehicle emergency notification system such as e-Call.

[0019] There is no particular limitation on the communication method of the communication unit 22. A plurality of communication methods may be used.

[0020] For example, the communication unit 22 performs wireless communication with devices in the vehicle using a communication method such as wireless LAN, Bluetooth, NFC, WUSB (Wireless USB), etc. For example, the communication unit 22 performs wired communication with devices in the vehicle using a communication method such as USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface, registered trademark), or MHL (Mobile High-definition Link) via a connection terminal (and a cable, if necessary) not shown.

[0021] Here, the in-vehicle device is, for example, a device that is not connected to the communication network 41 in the vehicle. For example, a mobile device or a wearable device carried by a passenger such as a driver, an information device brought into the vehicle and temporarily installed, etc. are assumed.

[0022] For example, the communication unit 22 communicates with servers and the like existing on an external network (e.g., the Internet, a cloud network, or an operator-specific network) via a base station or an access point using a wireless communication method such as 4G (fourth generation mobile communication system), 5G (fifth generation mobile communication system), LTE (Long Term Evolution), or DSRC (Dedicated Short Range Communications).

[0023] For example, the communication unit 22 communicates with a terminal (for example, a pedestrian or a store terminal, or an MTC (Machine Type Communication) terminal) present in the vicinity of the vehicle by using a P2P (Peer To Peer) technique. For example, the communication unit 22 performs V2X communication. The V2X communication includes, for example, vehicle-to-vehicle communication with other vehicles, vehicle-to-infrastructure communication with roadside devices, vehicle-to-home communication, and vehicle-to-pedestrian communication with terminals held by pedestrians.

[0024] For example, the communication unit 22 receives electromagnetic waves transmitted by a road traffic information and communication system (VICS (Vehicle Information and Communication System), registered trademark), such as a radio beacon, an optical beacon, or an FM multiplex broadcast.

[0025] The map information storage unit 23 stores maps acquired from an external source and maps created by the vehicle 1. For example, the map information storage unit 23 stores a three-dimensional high-precision map, a global map that has lower precision than a high-precision map and covers a wide area, and the like.

[0026] The high-precision map is, for example, a dynamic map, a point cloud map, a vector map (also called an ADAS (Advanced Driver Assistance System) map), etc. The dynamic map is, for example, a map consisting of four layers of dynamic information, quasi-dynamic information, quasi-static information, and static information, and is provided from an external server or the like. The point cloud map is a map configured of a point cloud (point cloud data). The vector map is a map in which information such as the positions of lanes and traffic lights is associated with the point cloud map. The point cloud map and the vector map may be provided from, for example, an external server or the like, or may be created by the vehicle 1 based on sensing results by the radar 52, the LiDAR 53, etc. as a map for matching with a local map to be described later, and may be stored in the map information storage unit 23. In addition, when a high-precision map is provided from an external server or the like, map data of, for example, several hundred meters square, related to the planned route along which the vehicle 1 will travel is acquired from the server or the like in order to reduce communication capacity.

[0027] The GNSS receiver 24 receives GNSS signals from GNSS satellites and supplies them to the driving assistance / automatic driving control unit 29.

[0028] The external recognition sensor 25 includes various sensors used to recognize the situation outside the vehicle 1, and supplies sensor data from each sensor to each unit of the vehicle control system 11. The type and number of sensors included in the external recognition sensor 25 are arbitrary.

[0029] For example, the external recognition sensor 25 includes a camera 51, a radar 52, a LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) 53, and an ultrasonic sensor 54. The number of cameras 51, radars 52, LiDARs 53, and ultrasonic sensors 54 is arbitrary, and an example of a sensing area of ​​each sensor will be described later.

[0030] As the camera 51, a camera of any imaging method, such as a ToF (Time Of Flight) camera, a stereo camera, a monocular camera, or an infrared camera, may be used as necessary.

[0031] Furthermore, for example, the external recognition sensor 25 includes an environmental sensor for detecting the weather, climate, brightness, etc. The environmental sensor includes, for example, a raindrop sensor, a fog sensor, a sunlight sensor, a snow sensor, an illuminance sensor, etc.

[0032] Furthermore, for example, the external recognition sensor 25 includes a microphone used to detect sounds around the vehicle 1 and the positions of sound sources.

[0033] The in-vehicle sensor 26 includes various sensors for detecting information inside the vehicle, and supplies sensor data from each sensor to each unit of the vehicle control system 11. The in-vehicle sensor 26 may include any type and any number of sensors.

[0034] For example, the in-vehicle sensor 26 includes a camera, a radar, a seating sensor, a steering wheel sensor, a microphone, a biosensor, etc. The camera may be a camera of any imaging method, such as a ToF camera, a stereo camera, a monocular camera, an infrared camera, etc. The biosensor is provided, for example, on a seat, a steering wheel, etc., and detects various types of bioinformation of a passenger such as a driver.

[0035] The vehicle sensor 27 includes various sensors for detecting the state of the vehicle 1, and supplies sensor data from each sensor to each unit of the vehicle control system 11. The type and number of sensors included in the vehicle sensor 27 are arbitrary.

[0036] For example, the vehicle sensor 27 includes a speed sensor, an acceleration sensor, an angular velocity sensor (gyro sensor), and an inertial measurement unit (IMU). For example, the vehicle sensor 27 includes a steering angle sensor that detects the steering angle of the steering wheel, a yaw rate sensor, an accelerator sensor that detects the amount of operation of an accelerator pedal, and a brake sensor that detects the amount of operation of a brake pedal. For example, the vehicle sensor 27 includes a rotation sensor that detects the number of rotations of the engine or the motor, an air pressure sensor that detects the air pressure of the tires, a slip ratio sensor that detects the slip ratio of the tires, and a wheel speed sensor that detects the rotation speed of the wheels. For example, the vehicle sensor 27 includes a battery sensor that detects the remaining charge and temperature of the battery, and an impact sensor that detects an impact from the outside.

[0037] The recording unit 28 includes, for example, a magnetic storage device such as a ROM (Read Only Memory), a RAM (Random Access Memory), a HDD (Hard Disc Drive), a semiconductor storage device, an optical storage device, and a magneto-optical storage device. The recording unit 28 records various programs and data used by each part of the vehicle control system 11. For example, the recording unit 28 records a rosbag file including messages transmitted and received by a ROS (Robot Operating System) on which an application program related to autonomous driving runs. For example, the recording unit 28 includes an EDR (Event Data Recorder) and a DSSAD (Data Storage System for Automated Driving), and records information on the vehicle 1 before and after an event such as an accident.

[0038] The driving support / automatic driving control unit 29 controls driving support and automatic driving of the vehicle 1. For example, the driving support / automatic driving control unit 29 includes an analysis unit 61, an action planning unit 62, and an operation control unit 63.

[0039] The analysis unit 61 performs an analysis process of the vehicle 1 and the surrounding situation. The analysis unit 61 includes a self-position estimation unit 71, a sensor fusion unit 72, and a recognition unit 73.

[0040] The self-position estimation unit 71 estimates the self-position of the vehicle 1 based on the sensor data from the external recognition sensor 25 and the high-precision map stored in the map information storage unit 23. For example, the self-position estimation unit 71 generates a local map based on the sensor data from the external recognition sensor 25, and estimates the self-position of the vehicle 1 by matching the local map with the high-precision map. The position of the vehicle 1 is based on, for example, the center of the rear wheel pair axle.

[0041] The local map is, for example, a three-dimensional high-precision map or an occupancy grid map created using a technique such as SLAM (Simultaneous Localization and Mapping). The three-dimensional high-precision map is, for example, the above-mentioned point cloud map. The occupancy grid map is a map in which the three-dimensional or two-dimensional space around the vehicle 1 is divided into grids of a predetermined size, and the occupancy state of an object is indicated on a grid-by-grid basis. The occupancy state of an object is indicated, for example, by the presence or absence of an object and the probability of its existence. The local map is also used, for example, in detection processing and recognition processing of the external situation of the vehicle 1 by the recognition unit 73.

[0042] The self-position estimation unit 71 may estimate the self-position of the vehicle 1 based on the GNSS signal and sensor data from the vehicle sensor 27.

[0043] The sensor fusion unit 72 performs a sensor fusion process to obtain new information by combining a plurality of different types of sensor data (for example, image data supplied from the camera 51 and sensor data supplied from the radar 52). Methods for combining different types of sensor data include integration, fusion, and association.

[0044] The recognition unit 73 performs detection processing and recognition processing of the situation outside the vehicle 1.

[0045] For example, the recognition unit 73 performs detection processing and recognition processing of the situation outside the vehicle 1 based on information from the external recognition sensor 25, information from the self-position estimation unit 71, information from the sensor fusion unit 72, and the like.

[0046] Specifically, for example, the recognition unit 73 performs detection processing and recognition processing of objects around the vehicle 1. The object detection processing is, for example, processing to detect the presence or absence, size, shape, position, movement, etc. of an object. The object recognition processing is, for example, processing to recognize attributes such as the type of object, or to identify a specific object. However, the detection processing and the recognition processing are not necessarily clearly separated, and may overlap.

[0047] For example, the recognition unit 73 performs clustering to classify a point cloud based on sensor data such as LiDAR or radar into clusters of points, thereby detecting objects around the vehicle 1. This allows the presence or absence, size, shape, and position of objects around the vehicle 1 to be detected.

[0048] For example, the recognition unit 73 performs tracking to follow the movement of the clusters of the point cloud classified by clustering, thereby detecting the movement of the objects around the vehicle 1. As a result, the speed and traveling direction (movement vector) of the objects around the vehicle 1 are detected.

[0049] For example, the recognition unit 73 performs object recognition processing such as semantic segmentation on the image data supplied from the camera 51 to recognize the type of object around the vehicle 1.

[0050] Note that objects to be detected or recognized may include, for example, vehicles, people, bicycles, obstacles, structures, roads, traffic lights, traffic signs, road markings, and the like.

[0051] For example, the recognition unit 73 performs a recognition process of traffic rules around the vehicle 1 based on the map stored in the map information storage unit 23, the estimation result of the vehicle's own position, and the recognition result of objects around the vehicle 1. Through this process, for example, the positions and states of traffic signals, the contents of traffic signs and road markings, the contents of traffic regulations, and lanes on which travel is possible are recognized.

[0052] For example, the recognition unit 73 performs a process of recognizing the environment around the vehicle 1. Examples of the surrounding environment to be recognized include the weather, temperature, humidity, brightness, and road surface condition.

[0053] The behavior planning unit 62 creates a behavior plan for the vehicle 1. For example, the behavior planning unit 62 creates the behavior plan by performing route planning and route following processing.

[0054] Global path planning is a process for planning a rough route from the start to the goal. This route planning also includes a process for local path planning, which is called trajectory planning and allows the vehicle 1 to proceed safely and smoothly in the vicinity of the vehicle 1, taking into account the motion characteristics of the vehicle 1 on the route planned by the route planning.

[0055] Path following is a process of planning an operation for traveling safely and accurately along a route planned by a route plan within a planned time. For example, a target speed and a target angular velocity of the vehicle 1 are calculated.

[0056] The action control unit 63 controls the action of the vehicle 1 in order to realize the action plan created by the action planning unit 62 .

[0057] For example, the operation control unit 63 controls the steering control unit 81, the brake control unit 82, and the drive control unit 83 to perform acceleration / deceleration control and direction control so that the vehicle 1 proceeds along the trajectory calculated by the trajectory plan. For example, the operation control unit 63 performs cooperative control for the purpose of realizing ADAS functions such as collision avoidance or impact mitigation, following driving, maintaining vehicle speed, collision warning for the vehicle itself, and lane departure warning for the vehicle itself. For example, the operation control unit 63 performs cooperative control for the purpose of automatic driving that autonomously drives without the driver's operation.

[0058] The DMS 30 performs authentication processing of the driver and recognition processing of the driver's state, etc., based on the sensor data from the in-vehicle sensor 26 and the input data input to the HMI 31. Examples of the driver's state to be recognized include physical condition, alertness level, concentration level, fatigue level, line of sight direction, level of intoxication, driving operation, posture, etc.

[0059] The DMS 30 may perform authentication processing of passengers other than the driver and recognition processing of the conditions of the passengers. For example, the DMS 30 may perform recognition processing of the conditions inside the vehicle based on sensor data from the in-vehicle sensor 26. Conditions inside the vehicle to be recognized include, for example, temperature, humidity, brightness, odor, and the like.

[0060] The HMI 31 is used to input various data, instructions, etc., generates input signals based on the input data, instructions, etc., and supplies them to each part of the vehicle control system 11. For example, the HMI 31 includes operation devices such as a touch panel, buttons, microphones, switches, and levers, as well as operation devices that allow input by means of voice, gestures, etc. other than manual operation. The HMI 31 may be, for example, a remote control device that uses infrared rays or other radio waves, or an externally connected device such as a mobile device or wearable device that supports the operation of the vehicle control system 11.

[0061] The HMI 31 also performs output control to generate and output visual information, auditory information, and tactile information for the occupant or the outside of the vehicle, as well as to control the output content, output timing, output method, etc. Visual information is information displayed by images or light, such as an operation screen, a status display of the vehicle 1, a warning display, a monitor image showing the situation around the vehicle 1, etc. Auditory information is information displayed by sound, such as guidance, warning sound, warning message, etc. Tactile information is information provided to the occupant's sense of touch by force, vibration, movement, etc.

[0062] Possible devices for outputting visual information include, for example, a display device, a projector, a navigation device, an instrument panel, a CMS (Camera Monitoring System), an electronic mirror, a lamp, etc. The display device may be a device having a normal display, or may be a device that displays visual information within the field of view of a passenger, such as, for example, a head-up display, a see-through display, or a wearable device with an AR (Augmented Reality) function.

[0063] Devices that output auditory information include, for example, audio speakers, headphones, earphones, and the like.

[0064] As a device for outputting tactile information, for example, a haptic element using haptic technology is assumed, and the haptic element is provided, for example, on a steering wheel, a seat, or the like.

[0065] The vehicle control unit 32 controls each part of the vehicle 1. The vehicle control unit 32 includes a steering control unit 81, a brake control unit 82, a drive control unit 83, a body system control unit 84, a light control unit 85, and a horn control unit 86.

[0066] The steering control unit 81 detects and controls the state of the steering system of the vehicle 1. The steering system includes, for example, a steering mechanism including a steering wheel, an electric power steering, etc. The steering control unit 81 includes, for example, a control unit such as an ECU that controls the steering system, and an actuator that drives the steering system.

[0067] The brake control unit 82 detects and controls the state of the brake system of the vehicle 1. The brake system includes, for example, a brake mechanism including a brake pedal, an ABS (Antilock Brake System), etc. The brake control unit 82 includes, for example, a control unit such as an ECU that controls the brake system, and an actuator that drives the brake system.

[0068] The drive control unit 83 detects and controls the state of the drive system of the vehicle 1. The drive system includes, for example, an accelerator pedal, a drive force generating device for generating drive force such as an internal combustion engine or a drive motor, a drive force transmission mechanism for transmitting the drive force to the wheels, etc. The drive control unit 83 includes, for example, a control unit such as an ECU that controls the drive system, and an actuator that drives the drive system.

[0069] The body system control unit 84 detects and controls the states of the body system systems of the vehicle 1. The body system systems include, for example, a keyless entry system, a smart key system, a power window device, a power seat, an air conditioning system, an airbag, a seat belt, a shift lever, etc. The body system control unit 84 includes, for example, a control unit such as an ECU that controls the body system systems, and an actuator that drives the body system systems.

[0070] The light control unit 85 detects and controls the states of various lights of the vehicle 1. Examples of lights to be controlled include headlights, backlights, fog lights, turn signals, brake lights, projections, and bumper displays. The light control unit 85 includes a control unit such as an ECU that controls the lights, and an actuator that drives the lights.

[0071] The horn control unit 86 detects and controls the state of the car horn of the vehicle 1. The horn control unit 86 includes, for example, a control unit such as an ECU that controls the car horn, and an actuator that drives the car horn.

[0072] FIG. 2 is a diagram showing an example of a sensing area by the camera 51, the radar 52, the LiDAR 53, and the ultrasonic sensor 54 of the external recognition sensor 25 in FIG.

[0073] The sensing area 101F and the sensing area 101B are examples of the sensing area of ​​the ultrasonic sensor 54. The sensing area 101F covers the periphery of the front end of the vehicle 1. The sensing area 101B covers the periphery of the rear end of the vehicle 1.

[0074] The sensing results in the sensing area 101F and the sensing area 101B are used for, for example, parking assistance for the vehicle 1.

[0075] Sensing area 102F to sensing area 102B show examples of sensing areas of a short-range or medium-range radar 52. Sensing area 102F covers a position farther in front of the vehicle 1 than sensing area 101F. Sensing area 102B covers a position farther in the rear of the vehicle 1 than sensing area 101B. Sensing area 102L covers the rear periphery of the left side of the vehicle 1. Sensing area 102R covers the rear periphery of the right side of the vehicle 1.

[0076] The sensing results in the sensing area 102F are used, for example, to detect vehicles, pedestrians, and the like that exist in front of the vehicle 1. The sensing results in the sensing area 102B are used, for example, for a collision prevention function behind the vehicle 1. The sensing results in the sensing area 102L and the sensing area 102R are used, for example, to detect objects in blind spots on the sides of the vehicle 1.

[0077] Sensing area 103F to sensing area 103B show examples of sensing areas by camera 51. Sensing area 103F covers a position farther in front of vehicle 1 than sensing area 102F. Sensing area 103B covers a position farther in the rear of vehicle 1 than sensing area 102B. Sensing area 103L covers the periphery of the left side of vehicle 1. Sensing area 103R covers the periphery of the right side of vehicle 1.

[0078] The sensing results in the sensing area 103F are used, for example, for recognition of traffic lights and traffic signs, lane departure prevention support systems, etc. The sensing results in the sensing area 103B are used, for example, for parking support and surround view systems, etc. The sensing results in the sensing areas 103L and 103R are used, for example, for surround view systems, etc.

[0079] Sensing area 104 illustrates an example of the sensing area of ​​the LiDAR 53. Sensing area 104 covers a position farther forward than sensing area 103F in front of the vehicle 1. On the other hand, sensing area 104 has a narrower range in the left-right direction than sensing area 103F.

[0080] The sensing results in the sensing area 104 are used for, for example, emergency braking, collision avoidance, pedestrian detection, and the like.

[0081] Sensing area 105 shows an example of the sensing area of ​​the long-distance radar 52. Sensing area 105 covers a position farther ahead of the vehicle 1 than sensing area 104. On the other hand, sensing area 105 has a narrower range in the left-right direction than sensing area 104.

[0082] The sensing results in the sensing area 105 are used for, for example, ACC (Adaptive Cruise Control) and the like.

[0083] The sensing area of ​​each sensor may have various configurations other than that shown in Fig. 2. Specifically, the ultrasonic sensor 54 may also sense the sides of the vehicle 1, and the LiDAR 53 may sense the rear of the vehicle 1.

[0084] <<2. Preferred embodiment>> Next, an embodiment of the present technology will be described with reference to FIGS.

[0085] <Configuration example of information processing system 201> FIG. 3 shows an example of the configuration of an information processing system 201 to which the present technology is applied.

[0086] The information processing system 201 is mounted on, for example, the vehicle 1 in FIG.

[0087] The information processing system 201 includes a camera 211, a LiDAR 212, and an information processing unit 213.

[0088] The camera 211 constitutes, for example, a part of the camera 51 in FIG. 1, captures an image of the area ahead of the vehicle 1, and supplies the obtained image (hereinafter, referred to as a captured image) to the information processing unit 213.

[0089] The LiDAR 212, for example, constitutes a part of the LiDAR 53 in FIG. 1, senses the area ahead of the vehicle 1, and at least a part of the sensing range overlaps with the shooting range of the camera 211. For example, the LiDAR 212 scans a laser pulse, which is a measurement light, in an azimuth direction (horizontal direction) and an elevation direction (height direction) in front of the vehicle 1, and receives the reflected light of the laser pulse. The LiDAR 212 calculates the direction and distance of a measurement point, which is a reflection point on an object that reflects the laser pulse, based on the scanning direction of the laser pulse and the time required to receive the reflected light. The LiDAR 212 generates point cloud data (point cloud), which is three-dimensional data indicating the direction and distance of each measurement point, based on the calculation result. The LiDAR 212 supplies the point cloud data to the information processing unit 213.

[0090] Here, the azimuth direction is the direction corresponding to the width direction (lateral direction, horizontal direction) of the vehicle 1. The elevation direction is the direction perpendicular to the traveling direction (distance direction) of the vehicle 1 and corresponding to the height direction (longitudinal direction, vertical direction) of the vehicle 1.

[0091] The information processing unit 213 includes an object region detection unit 221, an object recognition unit 222, an output unit 223, and a scan control unit 224. The information processing unit 213 constitutes, for example, a part of the vehicle control unit 32, the sensor fusion unit 72, and the recognition unit 73 in FIG.

[0092] The object region detection unit 221 detects a region where an object may exist (hereinafter referred to as an object region) in front of the vehicle 1 based on the point cloud data. The object region detection unit 221 associates the detected object region with information in the captured image (for example, a region in the captured image). The object region detection unit 221 supplies the captured image, the point cloud data, and information indicating the detection result of the object region to the object recognition unit 222.

[0093] Typically, as shown in FIG. 4, the point cloud data obtained by sensing the sensing range S1 in front of the vehicle 1 is converted into three-dimensional data in the world coordinate system shown at the bottom of the figure, and then each measurement point in the point cloud data is matched with a corresponding position in the captured image.

[0094] On the other hand, the object region detection unit 221 detects an object region indicating a range in the azimuth angle direction and the elevation angle direction in which an object may exist in the sensing range S1 based on the point cloud data. More specifically, as described below, the object region detection unit 221 detects an object region indicating a range in the elevation angle direction in which an object may exist for each vertically elongated rectangular unit region obtained by dividing the sensing range S1 in the azimuth angle direction based on the point cloud data. Then, the object region detection unit 221 associates each unit region with an area in the captured image. This reduces the processing required for associating point cloud data with the captured image.

[0095] The object recognition unit 222 performs object recognition in front of the vehicle 1 based on the object area detection result and the captured image. The object recognition unit 222 supplies the captured image, point cloud data, and information indicating the object area and the object recognition result to the output unit 223.

[0096] The output unit 223 generates and outputs output information indicating the result of object recognition, etc.

[0097] The scan control unit 224 controls the scanning of the laser pulse of the LiDAR 212. For example, the scan control unit 224 controls the scanning direction and scanning speed of the laser pulse of the LiDAR 212, etc.

[0098] In the following, the scanning of the laser pulse of the LiDAR 212 is also simply referred to as the scanning of the LiDAR 212. For example, the scanning direction of the laser pulse of the LiDAR 212 is also simply referred to as the scanning direction of the LiDAR 212.

[0099] <Object recognition processing> Next, the object recognition process executed by the information processing system 201 will be described with reference to the flowchart of FIG.

[0100] This process is started, for example, when an operation is performed to start the vehicle 1 and start driving, for example, when an ignition switch, a power switch, a start switch, or the like of the vehicle 1 is turned on. This process is ended, for example, when an operation is performed to end driving of the vehicle 1, for example, when an ignition switch, a power switch, a start switch, or the like of the vehicle 1 is turned off.

[0101] In step S1, the information processing system 201 acquires captured images and point cloud data.

[0102] Specifically, the camera 211 captures an image of the area ahead of the vehicle 1 , and supplies the captured image to an object region detection unit 221 of the information processing unit 213 .

[0103] The LiDAR 212 scans the front of the vehicle 1 with a laser pulse in the azimuth angle direction and the elevation angle direction under the control of the scan control unit 224, and receives the reflected light of the laser pulse. The LiDAR 212 calculates the distance to each measurement point in front of the vehicle 1 based on the time required to receive the reflected light. The LiDAR 212 generates point cloud data indicating the direction (elevation angle and azimuth angle) and distance of each measurement point, and supplies the point cloud data to the object region detection unit 221.

[0104] Here, an example of a scanning method of the LiDAR 212 by the scanning control unit 224 will be described with reference to FIGS.

[0105] FIG. 6 shows an example of the sensing range of the LiDAR 212 in the mounting angle and elevation angle directions.

[0106] 6A, the LiDAR 212 is installed on the vehicle 1 with a slight downward tilt. Therefore, a center line L1 in the elevation angle direction of the sensing range S1 is slightly tilted downward with respect to the road surface 301 from the horizontal direction.

[0107] 6B, the horizontal road surface 301 appears as an uphill slope from the LiDAR 212. That is, in the point cloud data of a relative coordinate system seen from the LiDAR 212 (hereinafter referred to as the LiDAR coordinate system), the road surface 301 appears as an uphill slope.

[0108] In contrast, typically, the coordinate system of the point cloud data is transformed from the LiDAR coordinate system to an absolute coordinate system (e.g., a world coordinate system), and then road surface estimation is performed based on the point cloud data.

[0109] Fig. 7A shows an example of imaging the point cloud data acquired by the LiDAR 212. Fig. 7B shows the point cloud data of Fig. 7A viewed from the side.

[0110] The horizontal plane indicated by the auxiliary line L2 in Fig. 7B corresponds to the center line L1 of the sensing range S1 in Fig. 6A and B, and indicates the mounting direction (mounting angle) of the LiDAR 212. The LiDAR 212 scans with a laser pulse in the elevation angle direction with the horizontal plane 212 as the center.

[0111] Here, when the laser pulses are scanned at equal intervals in the elevation angle direction, the closer the scanning direction of the laser pulses is to the road surface 301, the longer the interval at which the laser pulses are irradiated onto the road surface 301. Therefore, the farther the object 302 (FIG. 6) on the road surface 301 is from the vehicle 1, the longer the interval in the distance direction of the laser pulses reflected by the object 302 becomes. That is, the interval in the distance direction at which the object 302 can be detected becomes longer. For example, in the distant region R1 in FIG. 7, the interval in the distance direction at which the object can be detected is several meters. Also, the farther the object 302 is from the vehicle 1, the smaller the size of the object 302 as seen from the vehicle 1 becomes. Therefore, in order to improve the detection accuracy of distant objects, it is desirable to narrow the scanning interval in the elevation angle direction of the laser pulses as the scanning direction of the laser pulses is closer to the road surface 301.

[0112] On the other hand, as the angle (irradiation angle) at which the laser pulse is irradiated to the road surface 301 increases, the interval in the distance direction at which the laser pulse is irradiated to the road surface decreases, and the interval in the distance direction at which an object can be detected decreases. For example, in region R2 in Fig. 7, the interval in the distance direction at which the laser pulse is irradiated becomes shorter compared to region R1. Also, the closer an object is to the vehicle 1, the larger it appears to the vehicle 1. Therefore, as the irradiation angle of the laser pulse to the road surface 301 increases, the accuracy of object detection hardly decreases even if the scanning interval in the elevation angle direction of the laser pulse is increased to a certain extent.

[0113] Moreover, the objects to be recognized above the vehicle 1 are mainly traffic lights, road signs, guide boards, etc., and the risk of collision of the vehicle 1 is low. Furthermore, the further upward the scanning direction of the laser pulse faces, the shorter the interval in the distance direction at which the laser pulse is irradiated to the object above the vehicle 1 becomes, and the shorter the interval in the distance direction at which the object can be detected becomes. For example, in the region R3 in FIG. 7, the interval in the distance direction at which the laser pulse is irradiated becomes shorter compared to the region R1. Therefore, even if the scanning interval in the elevation angle direction of the laser pulse is increased to some extent as the scanning direction of the laser pulse faces upward, the detection accuracy of the object hardly decreases.

[0114] Figure 8 shows an example of point cloud data when laser pulses are scanned at equal intervals in the elevation angle direction. The right figure in Figure 8 shows an example of the point cloud data visualized as an image. The left figure in Figure 8 shows an example of the measurement points of the point cloud data arranged at the corresponding positions in the captured image.

[0115] As shown in this figure, if the LiDAR 212 is scanned at equal intervals in the elevation angle direction, the number of measurement points on the road surface near the vehicle 1 will be greater than necessary. This increases the processing load for the measurement points on the road surface near the vehicle 1, which may cause delays in object recognition.

[0116] In response to this, the scan control unit 224 controls the scan interval in the elevation angle direction of the LiDAR 212 based on the elevation angle.

[0117] Fig. 9 is a graph showing an example of a scanning interval in the elevation angle direction of the LiDAR 212. The horizontal axis of Fig. 9 indicates the elevation angle (unit: °), and the vertical axis indicates the scanning interval in the elevation angle direction (unit: °).

[0118] In this example, the scanning interval in the elevation angle direction of the LiDAR 212 becomes shorter as it approaches a predetermined elevation angle θ0, and is shortest at the elevation angle θ0.

[0119] The elevation angle θ0 is set according to the mounting angle of the LiDAR 212, and is set, for example, to an angle at which a laser pulse is irradiated to a position a predetermined reference distance away from the vehicle 1 on a horizontal road surface in front of the vehicle 1. The reference distance is set, for example, to the maximum distance in front of the vehicle 1 at which an object to be recognized (for example, a vehicle ahead) is desired to be recognized.

[0120] As a result, the scanning interval of the LiDAR 212 becomes shorter in an area closer to the reference distance, and the interval between measurement points in the distance direction becomes shorter.

[0121] On the other hand, the farther the area is from the reference distance, the longer the scanning interval of the LiDAR 212 becomes, and the longer the distance between the measurement points becomes. Therefore, the distance between the measurement points on the road surface in front of and near the vehicle 1 and in the area above the vehicle 1 becomes longer.

[0122] Fig. 10 shows an example of point cloud data when the scanning in the elevation angle direction of the LiDAR 212 is controlled as described above with reference to Fig. 9. The right diagram of Fig. 10 shows an example of the point cloud data being visualized, similar to the right diagram of Fig. 8. The left diagram of Fig. 10 shows an example of the measurement points of the point cloud data being arranged at the corresponding positions of the captured image, similar to the left diagram of Fig. 8.

[0123] As shown in this figure, the intervals in the distance direction of the measurement points become denser as they approach an area that is a predetermined reference distance away from the vehicle 1, and become sparser as they move away from an area that is a predetermined reference distance away from the vehicle 1. This makes it possible to thin out the measurement points of the LiDAR 212 and reduce the amount of calculation without reducing the object recognition accuracy.

[0124] FIG. 11 shows a second example of a scanning method for the LiDAR 212.

[0125] The right diagram in Fig. 11 shows an example of point cloud data visualized as in the right diagram in Fig. 8. The left diagram in Fig. 11 shows an example of each measurement point of the point cloud data arranged at the corresponding position in the captured image as in the left diagram in Fig. 8.

[0126] In this example, the scanning interval in the elevation angle direction of the laser pulse is controlled so that the scanning interval in the distance direction for the horizontal road surface ahead of the vehicle 1 is equal. This makes it possible to reduce the number of measurement points on the road surface, especially in the vicinity of the vehicle 1, and therefore the amount of calculation required when estimating the road surface based on point cloud data, for example.

[0127] Returning to FIG. 5, in step S2, object region detection section 221 detects an object region for each unit region based on the point cloud data.

[0128] FIG. 12 is a schematic diagram showing an example of a virtual plane, a unit region, and an object region.

[0129] 12 indicates a virtual plane. The virtual plane indicates the sensing range (scanning range) in the azimuth angle direction and elevation angle direction of the LiDAR 212. Specifically, the width of the virtual plane indicates the sensing range of the LiDAR 212 in the azimuth angle direction, and the height of the virtual plane indicates the sensing range of the LiDAR 212 in the elevation angle direction.

[0130] A virtual plane is divided in the azimuth direction into a number of vertically long rectangular (strip-like) regions, which represent unit areas. The width of each unit area may be equal or different. In the former case, the virtual plane is divided equally in the azimuth direction, and in the latter case, the virtual plane is divided at different angles.

[0131] The rectangular area indicated by diagonal lines in each unit area indicates an object area, which indicates the range of elevation angles in which an object may exist in each unit area.

[0132] Now, an example of a method for detecting an object region will be described with reference to FIGS.

[0133] FIG. 13 shows an example of distribution of point cloud data within one unit area (ie, within a range of a predetermined azimuth angle) in the case where a vehicle 351 is present at a position a distance d1 ahead of the vehicle 1.

[0134] A in Fig. 13 shows an example of a histogram of distances of measurement points of point cloud data in a unit area. The horizontal axis shows the distance from vehicle 1 to each measurement point. The vertical axis shows the number (frequency) of measurement points present at the distance shown on the horizontal axis.

[0135] FIG. 13B shows an example of the distribution of the elevation angle and distance of the measurement points of the point cloud data in a unit area. The horizontal axis shows the elevation angle in the scanning direction of the LiDAR 212. Note that here, the lower end of the sensing range of the LiDAR 212 in the elevation angle direction is set to 0°, and the upward direction is set to the positive direction. The vertical axis shows the distance to the measurement point existing in the direction of the elevation angle shown on the horizontal axis.

[0136] 13A, the frequency of the distance of the measurement point in the unit area is maximum immediately before the vehicle 1, and decreases as it approaches the distance d1 where the vehicle 351 is located. The frequency of the distance of the measurement point in the unit area also peaks near the distance d1, and becomes substantially zero between the distance d1 and the distance d2. After the distance d2, the frequency of the distance of the measurement point in the unit area becomes substantially constant at a value smaller than the frequency immediately before the distance d1. The distance d2 is, for example, the shortest distance to the point (measurement point) where the laser pulse reaches after passing the vehicle 351.

[0137] Note that there are no measurement points in the range from distance d1 to distance d2, so it is difficult to determine whether the area corresponding to that range is an occlusion area hidden behind an object (in this example, the vehicle 351) or an area without any object such as the sky.

[0138] 13B, the distance between measurement points in a unit area increases as the elevation angle increases in a range of elevation angles from 0° to angle θ1, and is substantially constant at distance d1 in a range of elevation angles from angle θ1 to angle θ2. Note that angle θ1 is the minimum elevation angle at which the laser pulse is reflected by vehicle 351, and angle θ2 is the maximum elevation angle at which the laser pulse is reflected by vehicle 351. The distance between measurement points in a unit area increases as the elevation angle increases in a range of elevation angles equal to or greater than angle θ2.

[0139] In the data in FIG. 13B, unlike the data in FIG. 13A, it is possible to quickly determine that the area corresponding to the elevation angle range where no measurement points exist (the elevation angle range where distance cannot be measured) is an area where no objects such as sky exist.

[0140] Then, object region detection section 221 detects an object region based on the distribution of the elevation angles and distances of the measurement points shown in B of Fig. 13. Specifically, object region detection section 221 differentiates the distribution of the distances of the measurement points in each unit region by the elevation angle for each unit region. Specifically, for example, object region detection section 221 calculates the difference in distance between adjacent measurement points in the elevation angle direction in each unit region.

[0141] Fig. 14 shows an example of the results of differentiating the distances of measurement points with respect to the elevation angle when the distances of the measurement points in a unit area are distributed as shown in B of Fig. 13. The horizontal axis shows the elevation angle, and the vertical axis shows the difference in distance between adjacent measurement points in the elevation angle direction (hereinafter referred to as the distance difference value).

[0142] For example, the distance difference value for a road surface on which no object exists is estimated to be within range R11. That is, the distance difference value is estimated to increase within a predetermined range as the elevation angle increases.

[0143] On the other hand, if an object is present on the road surface, the distance difference value is estimated to be within range R12. That is, the distance difference value is estimated to be equal to or less than a predetermined threshold value TH1, regardless of the elevation angle.

[0144] For example, the object region detection unit 221 determines that an object exists within a range of elevation angles from angle θ1 to angle θ2 in the example of Fig. 14. Then, the object region detection unit 221 detects the range of elevation angles from angle θ1 to angle θ2 in the target unit region as an object region.

[0145] In order to separate object regions corresponding to different objects in each unit region, it is desirable to set the number of detectable object regions in each unit region to 2 or more. On the other hand, in order to reduce the processing load, it is desirable to set an upper limit on the number of detected object regions in each unit region. For example, the upper limit on the number of detected object regions in each unit region is set within the range of 2 to 4.

[0146] Returning to FIG. 5, in step S3, object region detection section 221 detects a target object region based on the object region.

[0147] First, the object region detection unit 221 associates each object region with the captured image. Specifically, the mounting position and mounting angle of the camera 211 and the mounting position and mounting angle of the LiDAR 212 are known, and the positional relationship between the shooting range of the camera 211 and the sensing range of the LiDAR 212 is known. Therefore, the relative relationship between the virtual plane and each unit region and the region in the captured image is also known. Based on this known information, the object region detection unit 221 calculates a region corresponding to each object region in the captured image based on the position of each object region in the virtual plane, thereby associating each object region with the captured image.

[0148] 15 is a schematic diagram showing an example of correspondence between a captured image and an object region. A vertically long rectangular (strip-shaped) region in the captured image is the object region.

[0149] In this way, each object region is associated with the captured image based only on its position in the virtual plane, regardless of the content of the captured image, making it possible to quickly associate each object region with an area in the captured image with a small amount of calculation.

[0150] Furthermore, the object region detection unit 221 converts the coordinates of the measurement points in each object region from the LiDAR coordinate system to the camera coordinate system. That is, the coordinates of the measurement points in each object region are converted from coordinates represented by the azimuth angle, elevation angle, and distance in the LiDAR coordinate system to horizontal (x-axis direction) and vertical (y-axis direction) coordinates in the camera coordinate system. Furthermore, the depth direction (z-axis direction) coordinates of each measurement point are calculated based on the distance of the measurement point in the LiDAR coordinate system.

[0151] Next, the object region detection unit 221 performs a combining process to combine object regions estimated to correspond to the same object based on the relative positions between the object regions and the distances of the measurement points included in each object region. For example, the object region detection unit 221 combines adjacent object regions when the difference in distance is within a predetermined threshold based on the distances of the measurement points included in each of the adjacent object regions.

[0152] As a result, for example, each object region in FIG. 15 is separated into an object region including a vehicle and an object region including buildings in the background, as shown in FIG.

[0153] 15 and 16, the upper limit of the number of detected object regions in each unit region is set to 2. Therefore, for example, as shown in Fig. 16, a building and a streetlight may be included without being separated in the same object region, or a building, a streetlight, and the space between them may be included without being separated.

[0154] In contrast to this, for example, it is possible to detect object regions more accurately by setting the upper limit of the number of detected object regions in each unit region to 4. In other words, it becomes easier to separate object regions into individual objects.

[0155] FIG. 17 shows an example of the object region detection results when the upper limit of the number of object regions detected in each unit region is set to 4. The diagram on the left shows an example in which each object region is superimposed on the corresponding region of the captured image. The vertically long rectangular regions in the diagram are object regions. The diagram on the right shows an example of an image in which depth information is added to each object region. The depth direction length of each object region can be calculated based on, for example, the distance between measurement points in each object region.

[0156] By setting the upper limit of the number of detected object regions in each unit region to 4, it becomes easier to separate object regions corresponding to tall objects and object regions corresponding to short objects, for example, as shown in regions R21 and R22 in the diagram on the left. Also, it becomes easier to separate object regions corresponding to individual objects located far away, for example, as shown in region R23 in the diagram on the right.

[0157] Next, the object region detection section 221 detects an object region that may include an object that is an object to be recognized, from among the object regions after the combining process, based on the distribution of measurement points within each object region.

[0158] For example, object region detection section 221 calculates the size (area) of each object region based on the distribution of measurement points included in each object region in the x-axis direction and the y-axis direction. Also, object region detection section 221 calculates the tilt angle of each object region based on the range (dy) in the height direction (y-axis direction) and the range (dz) in the distance direction (z-axis direction) of measurement points included in each object region.

[0159] Then, the object region detection unit 221 extracts, from the object regions after the combining process, object regions whose area is equal to or larger than a predetermined threshold and whose inclination angle is equal to or larger than a predetermined threshold as target object regions. For example, when an object to be recognized is an object that needs to avoid a front collision, 2 An object region having an inclination angle of 30° or more is detected as a target object region.

[0160] For example, a rectangular object region is associated with a captured image as shown in Fig. 18, as shown in Fig. 19. Then, after a combining process of the object regions in Fig. 19 is performed, a target object region shown by a rectangular region in Fig. 20 is detected.

[0161] The object region detection unit 221 supplies the captured image, the point cloud data, and information indicating the detection results of the object region and the target object region to the object recognition unit 222.

[0162] Returning to FIG. 5, in step S4, the object recognition unit 222 sets a recognition range based on the target object region.

[0163] For example, as shown in Fig. 21, a recognition range R31 is set based on the detection result of the object region shown in Fig. 20. In this example, the width and height of the recognition range R31 are set to ranges obtained by adding a predetermined margin to the horizontal and vertical ranges in which the object region exists.

[0164] In step S5, the object recognition unit 222 performs object recognition within the recognition range.

[0165] For example, when the object to be recognized by the information processing system 201 is a vehicle ahead of the vehicle 1, as shown in FIG. 22, a vehicle 341 surrounded by a rectangular frame is recognized within a recognition range R31.

[0166] The object recognition unit 222 supplies the captured image, the point cloud data, and information indicating the detection result of the object region, the detection result of the target object region, the recognition range, and the object recognition result to the output unit 223.

[0167] In step S6, the output unit 223 outputs the result of the object recognition. Specifically, the output unit 223 generates output information indicating the result of the object recognition, etc., and outputs it to a subsequent stage.

[0168] 23 to 25 show specific examples of the output information.

[0169] 23 is a schematic diagram showing an example of output information in which an object recognition result is superimposed on a captured image. Specifically, a frame 361 surrounding a recognized vehicle 341 is superimposed on the captured image. In addition, information indicating the category of the recognized vehicle 341 (vehicle), information indicating the distance to the vehicle 341 (6.0 m), and information indicating the size of the vehicle 341 (width 2.2 m × height 2.2 m) are superimposed on the captured image.

[0170] The distance to the vehicle 341 and the size of the vehicle 341 are calculated based on, for example, the distribution of measurement points in the object region corresponding to the vehicle 341. The distance to the vehicle 341 is calculated based on, for example, the distribution of distances to measurement points in the object region corresponding to the vehicle 341. The size of the vehicle 341 is calculated based on, for example, the distribution of measurement points in the x-axis direction and y-axis direction in the object region corresponding to the vehicle 341.

[0171] Also, for example, only one of the distance to the vehicle 341 and the size of the vehicle 341 may be superimposed on the captured image.

[0172] Fig. 24 shows an example of output information in which images corresponding to each object region are arranged two-dimensionally based on the distribution of measurement points in each object region. Specifically, for example, based on the position of each object region in the virtual plane before the combining process, an image of an area in the captured image corresponding to each object region is associated with each object region. In addition, based on the direction (azimuth angle and elevation angle) and distance of the measurement point in each object region, the position of the azimuth angle direction, elevation angle direction, and distance direction of each object region are obtained. Then, the image corresponding to each object region is arranged two-dimensionally based on the position of each object region, thereby generating the output information shown in Fig. 24.

[0173] For example, an image corresponding to a recognized object may be displayed so as to be distinguishable from other images.

[0174] FIG. 25 shows an example of output information in which rectangular parallelepipeds corresponding to each object region are arranged two-dimensionally based on the distribution of measurement points in each object region. Specifically, the length of each object region in the depth direction is obtained based on the distance of the measurement points in each object region before the combining process. The length of each object region in the depth direction is calculated based on the difference in distance between the measurement point closest to the vehicle 1 and the measurement point farthest from the vehicle 1 among the measurement points in each object region, for example. In addition, the position of each object region in the azimuth angle direction, the elevation angle direction, and the distance direction is obtained based on the direction (azimuth angle and elevation angle) and distance of the measurement point in each object region. Then, the output information shown in FIG. 25 is generated by arranging rectangular parallelepipeds representing the width in the azimuth angle direction, the height in the elevation angle direction, and the length in the depth direction of each object region two-dimensionally based on the position of each object region.

[0175] For example, a rectangular parallelepiped corresponding to a recognized object may be displayed so as to be distinguishable from other rectangular parallelepipeds.

[0176] Thereafter, the process returns to step S1, and the processes from step S1 onwards are executed.

[0177] In this way, the load of object recognition using sensor fusion can be reduced.

[0178] Specifically, the scanning interval in the elevation angle direction of the LiDAR 212 is controlled based on the elevation angle, and the measurement points are thinned out, thereby reducing the processing load on the measurement points.

[0179] In addition, the object area is associated with an area in the captured image based only on the positional relationship between the sensing range of the LiDAR 212 and the shooting range of the camera 211. Therefore, the load is significantly reduced compared to the case where the measurement points of the point cloud data are associated with the corresponding positions in the captured image.

[0180] Furthermore, the target object region is detected based on the object region, and the recognition range is limited based on the target object region, thereby reducing the load on object recognition.

[0181] 26 and 27 show examples of the relationship between the recognition range and the processing time required for object recognition.

[0182] Fig. 26 shows an example of a captured image and a recognition range. Recognition range R41 shows an example of a recognition range in which the range for object recognition is restricted to an arbitrary shape based on the target object region. In this way, it is also possible to set a region other than a rectangle as the recognition range. Recognition range R42 is a recognition range in which the range for object recognition is restricted only in the height direction of the captured image based on the target object region.

[0183] When the recognition range R41 is used, it is possible to significantly reduce the processing time required for object recognition. On the other hand, when the recognition range R42 is used, the processing time cannot be reduced as much as when the recognition range R41 is used, but the processing time can be predicted in advance according to the number of lines in the recognition range R42, making system control easier.

[0184] 27 is a graph showing the relationship between the number of lines in a captured image included in recognition range R42 and the processing time required for object recognition. The horizontal axis shows the number of lines, and the vertical axis shows the processing time (in ms).

[0185] Curves L41 to L44 show the processing time when object recognition is performed using different algorithms for the recognition range in the captured image. As shown in this graph, in almost the entire range, the processing time becomes shorter as the number of lines in recognition range R42 decreases, regardless of the difference in algorithm.

[0186] <<3. Modifications>> Below, a modification of the above-described embodiment of the present technology will be described.

[0187] For example, it is also possible to set the object region to a shape other than a rectangle (for example, a rectangle with rounded corners, an ellipse, etc.).

[0188] For example, the object region may be associated with information other than the region in the captured image. For example, the object region may be associated with information (e.g., pixel information, metadata, etc.) of the region corresponding to the object region in the captured image.

[0189] For example, a plurality of recognition ranges may be set within a captured image. For example, when the positions of detected object regions are far apart, a plurality of recognition ranges may be set so that each object region is included in one of the recognition ranges.

[0190] In addition, for example, each recognition range may be classified based on the shape, size, position, distance, etc. of the object area contained in each recognition range, and object recognition may be performed using a method corresponding to the class of each recognition range.

[0191] For example, in the example of FIG. 28, recognition ranges R51 to R53 are set. Recognition range R51 includes a vehicle ahead and is classified as a class that requires precise object recognition. Recognition range R52 is classified as a class that includes tall objects such as road signs, traffic lights, street lights, utility poles, and overpasses. Recognition range R53 is classified as a class that includes distant background areas. Then, object recognition algorithms suited to the classes of each recognition range are applied to recognition ranges R51 to R53, and object recognition is performed. This improves the accuracy and speed of object recognition.

[0192] For example, the recognition range may be set based on the object region before the combining process or the object region after the combining process, without detecting the object region.

[0193] For example, without setting a recognition range, object recognition may be performed based on the object region before the combining process or the object region after the combining process.

[0194] The above-mentioned detection conditions for the object region are one example, and can be changed depending on, for example, the object to be recognized or the purpose of object recognition.

[0195] This technology can also be applied to cases where object recognition is performed using a distance measurement sensor (e.g., a millimeter wave radar, etc.) other than the LiDAR 212 for sensor fusion. This technology can also be applied to cases where object recognition is performed using sensor fusion using three or more types of sensors.

[0196] This technology can be applied not only to distance measurement sensors that scan measurement light such as laser pulses in azimuth and elevation directions, but also to distance measurement sensors that emit measurement light radially in azimuth and elevation directions and receive reflected light.

[0197] The present technology can also be applied to object recognition for applications other than the above-mentioned in-vehicle applications.

[0198] For example, this technology can be applied to cases where objects around a moving body other than a vehicle are recognized. For example, moving bodies such as motorcycles, bicycles, personal mobility, airplanes, ships, construction machinery, agricultural machinery (tractors), etc. are assumed. In addition, moving bodies to which this technology can be applied also include moving bodies that are remotely driven (operated) without a user on board, such as drones and robots.

[0199] For example, the present technology can also be applied to cases where object recognition is performed in a fixed location, such as a surveillance system.

[0200] <<4.Other>> <Computer configuration example> The above-mentioned series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed in a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, capable of executing various functions by installing various programs.

[0201] FIG. 29 is a block diagram showing an example of the hardware configuration of a computer that executes the above-mentioned series of processes by a program.

[0202] In a computer 1000, a central processing unit (CPU) 1001, a read only memory (ROM) 1002, and a random access memory (RAM) 1003 are interconnected by a bus 1004.

[0203] An input / output interface 1005 is further connected to the bus 1004. An input unit 1006, an output unit 1007, a recording unit 1008, a communication unit 1009, and a drive 1010 are connected to the input / output interface 1005.

[0204] The input unit 1006 includes an input switch, a button, a microphone, an image sensor, etc. The output unit 1007 includes a display, a speaker, etc. The recording unit 1008 includes a hard disk, a non-volatile memory, etc. The communication unit 1009 includes a network interface, etc. The drive 1010 drives removable media 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0205] In the computer 1000 configured as described above, the CPU 1001 performs the above-mentioned series of processes by, for example, loading a program recorded in the recording unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.

[0206] The program executed by the computer 1000 (CPU 1001) can be provided by being recorded on a removable medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0207] In the computer 1000, the program can be installed in the recording unit 1008 via the input / output interface 1005 by mounting the removable medium 1011 in the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the recording unit 1008. Alternatively, the program can be installed in the ROM 1002 or the recording unit 1008 in advance.

[0208] In addition, the program executed by the computer may be a program in which processing is performed chronologically in the order described in this specification, or it may be a program in which processing is performed in parallel or at the required timing, such as when called.

[0209] In this specification, a system refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all the components are in the same case. Therefore, multiple devices housed in separate cases and connected via a network, and a single device in which multiple modules are housed in a single case, are both systems.

[0210] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the gist of the present technology.

[0211] For example, the present technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.

[0212] Furthermore, each step described in the above flow chart can be executed by one device, or can be shared and executed by a plurality of devices.

[0213] Furthermore, when a single step includes multiple processes, the multiple processes included in the single step can be executed by a single device, or can be shared and executed by multiple devices.

[0214] <Examples of configuration combinations> The present technology can also be configured as follows.

[0215] (1) an object region detection unit that detects an object region indicating a range in an azimuth angle direction and an elevation angle direction in which an object exists within a sensing range of the distance measuring sensor based on three-dimensional data indicating a direction and a distance of each measurement point measured by the distance measuring sensor, and associates the object region with information in an image captured by a camera having an imaging range at least partially overlapping with the sensing range; An information processing device. (2) The object region detection unit detects an object region, which indicates a range in an elevation angle direction in which an object exists, for each unit region obtained by dividing the sensing range in an azimuth angle direction. The information processing device according to (1). (3) The object region detection unit is capable of detecting the object regions in each unit region, the number of which is equal to or less than a predetermined upper limit. The information processing device according to (2). (4) The object region detection unit detects the object region based on a distribution of elevation angles and distances of the measurement points in the unit region. The information processing device according to (2) or (3). (5) an object recognition unit that performs object recognition based on the captured image and the detection result of the object region; The information processing device according to any one of (1) to (4) further comprises: (6) The object recognition unit sets a recognition range in which object recognition is performed in the captured image based on a detection result of the object region, and performs object recognition within the recognition range. The information processing device according to (5). (7) the object region detection unit performs a process of combining the object regions based on the relative positions between the object regions and the distances between the measurement points included in each of the object regions, and detects an object region in which an object to be recognized may exist based on the object regions after the process of combining; The object recognition unit sets the recognition range based on a detection result of the object region. The information processing device according to (6). (8) The object region detection unit detects the object region based on a distribution of the measurement points in each of the object regions after the combining process. The information processing device according to (7). (9) The object region detection unit calculates a size and a tilt angle of each of the object regions based on a distribution of the measurement points in each of the object regions after the combining process, and detects the target object region based on the size and tilt angle of each of the object regions. The information processing device according to (8). (10) The object recognition unit classifies the recognition range based on the object region included in the recognition range, and recognizes the object by a method according to the class of the recognition range. The information processing device according to any one of (7) to (9). (11) the object region detection unit calculates at least one of a size and a distance of the recognized object based on a distribution of the measurement points in the object region corresponding to the recognized object; an output unit that generates output information in which at least one of the size and distance of the recognized object is superimposed on the captured image; The information processing device according to any one of (7) to (10) further comprises: (12) an output unit that generates output information in which images corresponding to each of the object regions are arranged in two dimensions based on the distribution of the measurement points within each of the object regions; The information processing device according to any one of (1) to (10) further comprises: (13) an output unit that generates output information in which rectangular parallelepipeds corresponding to each of the object regions are arranged two-dimensionally based on the distribution of the measurement points within each of the object regions; The information processing device according to any one of (1) to (10) further comprises: (14) The object region detection unit performs a process of combining the object regions based on the relative positions between the object regions and the distances between the measurement points included in each of the object regions. The information processing device according to any one of (1) to (6). (15) The object region detection unit detects an object region in which an object to be recognized may exist based on a distribution of the measurement points in each of the object regions after the combining process. The information processing device according to (14). (16) a scanning control unit that controls a scanning interval in an elevation angle direction of the distance measuring sensor based on the elevation angle of the sensing range; The information processing device according to any one of (1) to (15) further comprises: (17) The distance measuring sensor senses the area ahead of the vehicle, The scanning control unit shortens a scanning interval in the elevation angle direction of the distance measuring sensor as the scanning direction in the elevation angle direction of the distance measuring sensor approaches an angle at which the measurement light of the distance measuring sensor is irradiated to a position a predetermined distance away from the vehicle on a horizontal road surface in front of the vehicle. The information processing device according to (16). (18) The distance measuring sensor senses the area ahead of the vehicle, The scanning control unit controls a scanning interval in an elevation angle direction of the distance measuring sensor so that a scanning interval in a distance direction with respect to a horizontal road surface ahead of the vehicle becomes equal. The information processing device according to (16). (19) Based on three-dimensional data indicating the direction and distance of each measurement point measured by the distance measuring sensor, an object area indicating the range in the azimuth angle direction and elevation angle direction in which an object exists within the sensing range of the distance measuring sensor is detected, and the object area is associated with information in an image captured by a camera whose capturing range at least partially overlaps with the sensing range. Information processing methods. (20) Based on three-dimensional data indicating the direction and distance of each measurement point measured by the distance measuring sensor, an object area indicating the range in the azimuth angle direction and elevation angle direction in which an object exists within the sensing range of the distance measuring sensor is detected, and the object area is associated with information in an image captured by a camera whose capturing range at least partially overlaps with the sensing range. A program that causes a computer to execute a process.

[0216] It should be noted that the effects described in this specification are merely examples and are not limiting, and other effects may also be obtained. [Explanation of symbols]

[0217] 1 vehicle, 11 vehicle control system, 32 vehicle control unit, 51 camera, 53 LiDAR, 72 sensor fusion unit, 73 recognition unit, 201 information processing system, 211 camera, 212 LiDAR, 213 information processing unit, 221 object region detection unit, 222 object recognition unit, 223 output unit, 224 scanning control unit

Claims

1. an object region detection unit that detects an object region indicating a range in which an object exists based on a distribution of elevation angles and distances of the measurement points within unit regions obtained by dividing a sensing range of the distance measurement sensor in an azimuth angle direction, the distribution being based on three-dimensional data indicating the direction and distance of each measurement point measured by the distance measurement sensor, and that associates the object region with information within an image captured by a camera having an imaging range at least partially overlapping with the sensing range; An information processing device.

2. The object region detection unit detects, for each unit region, the object region that indicates a range in an elevation angle direction in which an object exists. The information processing device according to claim 1 .

3. The object region detection unit is capable of detecting the object regions in each unit region, the number of which is equal to or less than a predetermined upper limit. The information processing device according to claim 1 .

4. an object recognition unit that performs object recognition based on the captured image and the detection result of the object region; The information processing device according to claim 1 further comprising:

5. The object recognition unit sets a recognition range in which object recognition is performed in the captured image based on a detection result of the object region, and performs object recognition within the recognition range. The information processing device according to claim 4.

6. the object region detection unit performs a process of combining the object regions based on the relative positions between the object regions and the distances between the measurement points included in each of the object regions, and detects an object region in which an object to be recognized may exist based on the object regions after the process of combining; The object recognition unit sets the recognition range based on a detection result of the object region. The information processing device according to claim 5 .

7. The object region detection unit detects the object region based on a distribution of the measurement points in each of the object regions after the combining process. The information processing device according to claim 6.

8. The object region detection unit calculates a size of each of the object regions and an inclination angle with respect to a distance direction based on a distribution of the measurement points in each of the object regions after the combining process, and detects the target object region based on the size and the inclination angle of each of the object regions. The information processing device according to claim 7.

9. The object recognition unit classifies the recognition range based on the object region included in the recognition range, and recognizes the object by a method according to the class of the recognition range. The information processing device according to claim 6.

10. the object region detection unit calculates at least one of a size and a distance of the recognized object based on a distribution of the measurement points within the object region corresponding to the recognized object; an output unit that generates output information in which at least one of the size and distance of the recognized object is superimposed on the captured image; The information processing device according to claim 6 .

11. an output unit that generates output information in which images corresponding to the object regions are two-dimensionally arranged based on the distribution of the measurement points in each of the object regions; The information processing device according to claim 1 further comprising:

12. an output unit that generates output information in which rectangular parallelepipeds corresponding to each of the object regions are arranged two-dimensionally based on the distribution of the measurement points within each of the object regions; The information processing device according to claim 1 further comprising:

13. The object region detection unit performs a process of combining the object regions based on the relative positions between the object regions and the distances between the measurement points included in each of the object regions. The information processing device according to claim 1 .

14. The object region detection unit detects an object region in which an object to be recognized may exist based on a distribution of the measurement points in each of the object regions after the combining process. The information processing device according to claim 13.

15. a scanning control unit that controls a scanning interval in an elevation angle direction of the distance measuring sensor based on the elevation angle of the sensing range; The information processing device according to claim 1 further comprising:

16. The distance measuring sensor senses the area ahead of the vehicle, The scanning control unit shortens a scanning interval in the elevation angle direction of the distance measuring sensor as the scanning direction in the elevation angle direction of the distance measuring sensor approaches an angle at which the measurement light of the distance measuring sensor is irradiated to a position a predetermined distance away from the vehicle on a horizontal road surface in front of the vehicle. The information processing device according to claim 15.

17. The distance measuring sensor senses the area ahead of the vehicle, The scanning control unit controls a scanning interval in an elevation angle direction of the distance measuring sensor so that a scanning interval in a distance direction with respect to a horizontal road surface ahead of the vehicle becomes equal. The information processing device according to claim 15.

18. An object area indicating a range in which an object exists is detected based on a distribution of elevation angles and distances of the measurement points within unit areas obtained by dividing a sensing range of the distance measuring sensor in an azimuth direction, the distribution being based on three-dimensional data indicating the direction and distance of each measurement point measured by the distance measuring sensor, and the object area is associated with information within an image captured by a camera having an imaging range at least partially overlapping with the sensing range. Information processing methods.

19. An object area indicating a range in which an object exists is detected based on a distribution of elevation angles and distances of the measurement points within unit areas obtained by dividing a sensing range of the distance measuring sensor in an azimuth direction, the distribution being based on three-dimensional data indicating the direction and distance of each measurement point measured by the distance measuring sensor, and the object area is associated with information within an image captured by a camera having an imaging range at least partially overlapping with the sensing range. A program that causes a computer to execute a process.

Citation Information

Patent Citations

  • Vehicle-to-vehicle distance controller

    JP1996045000A

  • Device for monitoring outside of vehicle

    JP2003151094A

  • Image processing apparatus and method

    JP2005284471A

  • Obstacle detecting device and method

    JP2006140636A

  • On-vehicle image processing device

    JP2006151125A