Information processing device, information processing method, program, mobile body control device, and mobile body
By applying geometric transformation and object recognition model to image and sensor images, the problem of insufficient recognition accuracy of cameras and millimeter wave radars in the prior art is solved, and a higher precision object recognition is achieved.
Patent Information
- Application Number
- CN201980079021.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-07
- Filing Date
- 2019-11-22
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2039-11-22
AI Technical Summary
The prior art has failed to effectively improve the accuracy of identifying objects using cameras and millimeter wave radars.
By performing geometric transformation of the image and sensor images, their coordinate systems are matched, and object recognition models are used to identify objects.
It improves the recognition accuracy of object objects and enhances the recognition ability of mobile objects.
Smart Images

Figure CN113168691B_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to an information processing device, an information processing method, a program, a mobile body control device, and a mobile body, and particularly to an information processing device, an information processing method, a program, a mobile body control device, and a mobile body that aim to improve the accuracy of recognizing an object. Background Art
[0002] Conventionally, there has been a proposal to projectively transform the radar plane and the imaging plane to superimpose and display the position information of obstacles detected by a millimeter-wave radar on a captured image (for example, see Patent Document 1).
[0003] Citation List
[0004] Patent Literature
[0005] Patent Document 1: Japanese Patent Application Publication No. 2005-175603 Summary of the Invention
[0006] Technical issues
[0007] However, Patent Document 1 does not discuss improving the accuracy of recognizing an object such as a vehicle using a camera and a millimeter wave radar.
[0008] The present technology has been made in view of the above-described circumstances, and aims to improve the accuracy of identifying an object.
[0009] An information processing device according to a first aspect of the present technology includes: a geometric transformation section that transforms at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, the captured image being obtained by an image sensor, the sensor image indicating a sensing result of the sensor, and a sensing range of the sensor at least partially overlapping with a sensing range of the image sensor; and an object recognition section that performs processing for recognizing an object based on the captured image and the sensor image whose coordinate systems have been matched with each other.
[0010] According to the first aspect of the present technology, an information processing method is performed by an information processing device, and the information processing method includes: transforming at least one of a captured image and a sensor image, and matching the coordinate systems of the captured image and the sensor image, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and the sensing range of the sensor at least partially overlaps with the sensing range of the image sensor; and performing processing for identifying an object based on the captured image and the sensor image whose coordinate systems have been matched with each other.
[0011] According to the first aspect of the present technology, a program causes a computer to perform processing including the following steps: transforming at least one of a captured image and a sensor image and matching the coordinate systems of the captured image and the sensor image, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and the sensing range of the sensor at least partially overlaps with the sensing range of the image sensor; and performing processing for identifying an object based on the captured image and the sensor image whose coordinate systems have been matched with each other.
[0012] According to the second aspect of the present technology, a mobile body control device includes: a geometric transformation part that transforms at least one of a captured image and a sensor image so that the coordinate systems of the captured image and the sensor image match, the captured image is obtained by an image sensor, the sensor image indicates the sensing result of the sensor, and the sensing range of the sensor at least partially overlaps with the sensing range of the image sensor; an object recognition part that performs processing for identifying an object based on the captured image and the sensor image whose coordinate systems have been matched with each other; and a motion controller that controls the action of the mobile body based on the recognition result of the object.
[0013] According to the third aspect of the present technology, a mobile body control device includes: an image sensor; a sensor, the sensing range of the sensor at least partially overlaps with the sensing range of the image sensor; a geometric transformation part, which transforms at least one of the captured image and the sensor image so that the coordinate systems of the captured image and the sensor image match, the captured image is obtained by the image sensor, and the sensor image indicates the sensing result of the sensor; an object recognition part, which performs processing to recognize an object based on the captured image and the sensor image whose coordinate systems have been matched with each other; and an action controller, which controls an action based on the recognition result of the object.
[0014] In a first aspect of the present technology, at least one of a captured image and a sensor image is transformed so that the coordinate systems of the captured image and the sensor image match, the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and the sensing range of the sensor at least partially overlaps with the sensing range of the image sensor; and processing for identifying an object is performed based on the captured image and the sensor image whose coordinate systems have been matched with each other.
[0015] In a second aspect of the present technology, at least one of a captured image and a sensor image is transformed so that the coordinate systems of the captured image and the sensor image match, the captured image is obtained by an image sensor that takes an image of the surrounding environment of a moving body, the sensor image indicates a sensing result of the sensor, and the sensing range of the sensor at least partially overlaps with the sensing range of the image sensor; processing of identifying an object is performed based on the captured image and the sensor image whose coordinate systems have been matched with each other; and the action of the moving body is controlled based on the recognition result of the object.
[0016] In a third aspect of the present technology, at least one of a captured image and a sensor image is transformed so that the coordinate systems of the captured image and the sensor image match, the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and the sensing range of the sensor at least partially overlaps with the sensing range of the image sensor; processing for identifying an object is performed based on the captured image and the sensor image whose coordinate systems have been matched with each other; and the action of a moving body is controlled based on the recognition result of the object. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a block diagram showing an example of the configuration of a vehicle control system to which the present technology is applied.
[0018] Figure 2 is a block diagram showing a first embodiment of a data acquisition section and a first embodiment of a vehicle exterior information detector.
[0019] Figure 3 An example of the configuration of an object recognition model is shown.
[0020] Figure 4 An example of the configuration of a learning system is shown.
[0021] Figure 5 is a flowchart for describing the learning process performed on the object recognition model.
[0022] Figure 6 An example of a low-resolution image is shown.
[0023] Figure 7 A diagram showing an example of correct answer data.
[0024] Figure 8 An example of a millimeter wave image is shown.
[0025] Figure 9 Examples of a geometrically transformed signal intensity image and a geometrically transformed velocity image are shown.
[0026] Figure 10An example of the result of object recognition processing performed using only millimeter wave data is shown.
[0027] Figure 11 is a flowchart for describing a first embodiment of object recognition processing.
[0028] Figure 12 An example of the recognition result of the object is shown.
[0029] Figure 13 It is a diagram used to describe the effects provided by this technology.
[0030] Figure 14 This is a block diagram showing a second embodiment of the vehicle exterior information detection device.
[0031] Figure 15 is a flowchart for describing a second embodiment of the object recognition process.
[0032] Figure 16 An example of the relationship between a captured image and a cropped image is shown.
[0033] Figure 17 is a block diagram showing a second embodiment of a data acquisition portion and a third embodiment of a vehicle exterior information detector.
[0034] Figure 18 This is a block diagram showing a fourth embodiment of the vehicle exterior information detection device.
[0035] Figure 19 This is a block diagram showing a fifth embodiment of the vehicle exterior information detection device.
[0036] Figure 20 is a block diagram showing a third embodiment of the data acquisition portion and a sixth embodiment of the vehicle exterior information detector.
[0037] Figure 21 is a block diagram showing a fourth embodiment of the data acquisition portion and a seventh embodiment of the vehicle exterior information detector.
[0038] Figure 22 An example of the configuration of a computer is shown. DETAILED DESCRIPTION
[0039] The following describes an embodiment for implementing the present technology. The description is given in the following order.
[0040] 1. First Embodiment (First Example Using a Camera and Millimeter-Wave Radar)
[0041] 2. Second embodiment (image cropping example)
[0042] 3. Third Embodiment (First Example Using Camera, Millimeter Wave Radar, and LiDAR)
[0043] 4. Fourth Embodiment (Second Example Using a Camera and Millimeter-Wave Radar)
[0044] 5. Fifth Embodiment (Second Example Using Camera, Millimeter Wave Radar, and LiDAR)
[0045] 6. Sixth Embodiment (Third Example Using a Camera and Millimeter-Wave Radar)
[0046] 7. Seventh Embodiment (Fourth Example Using a Camera and Millimeter Wave Radar)
[0047] 8. Modifications
[0048] 9. Other examples
[0049] <<1. First embodiment>>
[0050] First, refer to Figures 1 to 13 A first embodiment of the present technology is described.
[0051] <Configuration Example of Vehicle Control System 100>
[0052] Figure 1 1 is a block diagram showing an example of a schematic functional configuration of a vehicle control system 100 as an example of a moving body control system to which the present technology can be applied.
[0053] Note that when the vehicle 10 provided with the vehicle control system 100 is to be distinguished from other vehicles, the vehicle provided with the vehicle control system 100 will be referred to as a host car or a host vehicle hereinafter.
[0054] The vehicle control system 100 includes an input unit 101, a data acquisition unit 102, a communication unit 103, an onboard device 104, an output controller 105, an output unit 106, a powertrain controller 107, a powertrain system 108, a vehicle body-related controller 109, a vehicle body-related system 110, a memory 111, and an autonomous driving controller 112. The input unit 101, the data acquisition unit 102, the communication unit 103, the output controller 105, the powertrain controller 107, the vehicle body-related controller 109, the memory 111, and the autonomous driving controller 112 are interconnected via a communication network 121. For example, the communication network 121 includes a bus or an onboard communication network conforming to any standard, such as a controller area network (CAN), a local interconnect network (LIN), a local area network (LAN), or FlexRay (registered trademark). Note that the various components of the vehicle control system 100 can be directly connected to each other without using the communication network 121.
[0055] Note that when the various components of the vehicle control system 100 communicate with each other via the communication network 121, the description of the communication network 121 will be omitted below. For example, when the input section 101 and the autonomous driving controller 112 communicate with each other via the communication network 121, it will be simply stated that the input section 101 and the autonomous driving controller 112 communicate with each other.
[0056] The input section 101 includes devices used by onboard personnel to input various data, instructions, and the like. For example, the input section 101 includes operating devices such as a touch panel, buttons, microphones, switches, and joysticks; operating devices that allow input to be performed by methods other than manual operation, such as voice or gestures; and the like. Alternatively, for example, the input section 101 may be an externally connected device such as a remote control device using infrared rays or other radio waves, or a mobile device or wearable device compatible with the operation of the vehicle control system 100. The input section 101 generates input signals based on the data, instructions, and the like input by the onboard personnel, and provides the generated input signals to the various components of the vehicle control system 100.
[0057] The data acquisition portion 102 includes various sensors and the like for acquiring data used for processing performed by the vehicle control system 100 , and supplies the acquired data to the respective constituent elements of the vehicle control system 100 .
[0058] For example, the data acquisition section 102 includes various sensors for detecting, for example, the state of the vehicle. Specifically, for example, the data acquisition section 102 includes a gyroscope; an acceleration sensor; an inertial measurement unit (IMU); and sensors for detecting the amount of accelerator pedal operation, the amount of brake pedal operation, the steering angle of the steering wheel, the number of engine revolutions, the number of motor revolutions, the number of wheel speeds, and the like.
[0059] In addition, for example, the data acquisition section 102 includes various sensors for detecting information about the exterior of the vehicle. Specifically, for example, the data acquisition section 102 includes image capture devices such as a time-of-flight (ToF) camera, a stereo camera, a monocular camera, an infrared camera, and other cameras. In addition, for example, the data acquisition section 102 includes environmental sensors for detecting weather, meteorological phenomena, and the like, as well as surrounding information detection sensors for detecting objects around the vehicle. For example, environmental sensors include raindrop sensors, fog sensors, sunlight sensors, snow sensors, and the like. Surrounding information detection sensors include ultrasonic sensors, radars, LiDAR (light detection and ranging, laser imaging detection and ranging), sonars, and the like.
[0060] Furthermore, for example, the data acquisition section 102 includes various sensors for detecting the current position of the vehicle. Specifically, for example, the data acquisition section 102 includes a global navigation satellite system (GNSS) receiver that receives GNSS signals from GNSS satellites.
[0061] Furthermore, for example, the data acquisition section 102 includes various sensors for detecting information about the vehicle interior. Specifically, for example, the data acquisition section 102 includes a camera for capturing an image of the driver, a biometric sensor for detecting the driver's biological information, and a microphone for collecting sounds from the vehicle interior. For example, the biometric sensor is attached to a seat surface, a steering wheel, or the like, and detects biological information of a person sitting in the seat or the driver holding the steering wheel.
[0062] The communication section 103 communicates with the in-vehicle device 104, various off-vehicle devices, servers, base stations, and the like, transmitting data provided by the various components of the vehicle control system 100 and providing received data to the various components of the vehicle control system 100. Note that the communication protocols supported by the communication section 103 are not particularly limited. The communication section 103 can also support multiple types of communication protocols.
[0063] For example, the communication section 103 wirelessly communicates with the in-vehicle device 104 using wireless LAN, Bluetooth (registered trademark), near field communication (NFC), wireless USB (WUSB), etc. In addition, for example, the communication section 103 communicates with the in-vehicle device 104 by wire via a connection terminal (not shown) (and, if necessary, a cable) by using a universal serial bus (USB), a high-definition multimedia interface (HDMI) (registered trademark), a mobile high-definition link (MHL), etc.
[0064] In addition, for example, the communication section 103 communicates with a device (e.g., an application server or a control server) located in an external network (e.g., the Internet, a cloud network, or a carrier-specific network) through a base station or an access point. In addition, for example, the communication section 103 communicates with a terminal (e.g., a terminal of a pedestrian or a store, or a machine type communication (MTC) terminal) located near the vehicle using a point-to-point (P2P) technology. In addition, for example, the communication section 103 performs V2X communication such as vehicle-to-vehicle communication, vehicle-to-infrastructure communication, vehicle-to-residence communication between the vehicle and a residence, vehicle-to-vehicle communication, and the like. In addition, for example, the communication section 103 includes a beacon receiver that receives radio waves or electromagnetic waves emitted from, for example, a radio station installed on the road, and obtains information about, for example, the current position, traffic congestion, traffic regulations, or necessary time.
[0065] Examples of the in-vehicle device 104 include mobile devices or wearable devices of people on board, information devices brought into or attached to the host car, and navigation devices that search for routes to any destination.
[0066] The output controller 105 controls the output of various information to occupants of the vehicle or to the exterior of the vehicle. For example, the output controller 105 generates an output signal including at least one of visual information (such as image data) or audio information (such as sound data), supplies the output signal to the output section 106, and thereby controls the output of the visual and audio information from the output section 106. Specifically, for example, the output controller 105 synthesizes the data of images captured by the different image capture devices of the data acquisition section 102 to generate a bird's-eye view image, a panoramic image, etc., and supplies the output signal including the generated image to the output section 106. Furthermore, for example, the output controller 105 generates sound data including, for example, a warning buzzer or a warning message warning of dangers such as collision, contact, or entry into a dangerous area, and supplies the output signal including the generated sound data to the output section 106.
[0067] The output portion 106 includes a device capable of outputting visual information or audio information to occupants of the vehicle or to the exterior of the vehicle. For example, the output portion 106 includes a display device, an instrument panel, audio speakers, headphones, a wearable device such as a glasses-type display for a person in the vehicle, a projector, a light, and the like. Instead of a device including a commonly used display, the display device included in the output portion 106 may be a device that displays visual information in the driver's field of view, such as a head-up display, a transparent display, or a device including an augmented reality (AR) display function.
[0068] The powertrain controller 107 generates various control signals and supplies them to the powertrain system 108, thereby controlling the powertrain system 108. In addition, the powertrain controller 107 supplies control signals to components other than the powertrain system 108 as needed, for example, to notify them of the state of controlling the powertrain system 108.
[0069] The powertrain system 108 includes various devices related to the powertrain of the vehicle. For example, the powertrain system 108 includes a driving force generating device, such as an internal combustion engine and a drive motor, for generating driving force; a driving force transmission mechanism for transmitting driving force to the wheels; a steering mechanism for adjusting the steering angle; a braking device for generating braking force; an anti-lock braking system (ABS); an electronic stability control (ESC) system; an electric power steering system; and the like.
[0070] The vehicle body controller 109 generates various control signals and supplies them to the vehicle body system 110, thereby controlling the vehicle body system 110. Furthermore, the vehicle body controller 109 supplies control signals to components other than the vehicle body system 110 as needed, for example, to notify them of the status of controlling the vehicle body system 110.
[0071] The vehicle body-related systems 110 include various vehicle body-related devices installed on the vehicle body. For example, the vehicle body-related systems 110 include a keyless entry system, a smart key system, power windows, power seats, a steering wheel, an air conditioner, and various lights (such as headlights, taillights, brake lights, blinkers, and fog lights).
[0072] For example, the storage device 111 includes a read-only memory (ROM), a random access memory (RAM), a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, a magneto-optical storage device, etc. The storage portion 111 stores various programs, data, etc. used by the various components of the vehicle control system 100. For example, the storage device 111 stores map data such as a three-dimensional high-precision map, a global map, and a local map. The high-precision map is a dynamic map, etc. The global map has a lower degree of accuracy and covers a wider area than the high-precision map. The local map includes information about the surrounding environment of the vehicle.
[0073] The autonomous driving controller 112 performs control related to autonomous driving, such as autonomous driving or driving assistance. Specifically, for example, the autonomous driving controller 112 performs cooperative control aimed at realizing the functions of an advanced driver assistance system (ADAS), including collision avoidance or impact reduction of the vehicle, driving behind a preceding vehicle based on the distance between vehicles, driving while maintaining vehicle speed, warning of a collision of the vehicle, warning of lane departure of the vehicle, etc. In addition, for example, the autonomous driving controller 112 performs cooperative control aimed at realizing, for example, autonomous driving, which is autonomous driving without any operation performed by the driver. The autonomous driving controller 112 includes a detector 131, an own position estimator 132, a state analyzer 133, a planning section 134, and an action controller 135.
[0074] The detector 131 detects various information required for controlling the automatic driving, and includes an external information detector 141 , an internal information detector 142 , and a vehicle state detector 143 .
[0075] The vehicle exterior information detector 141 performs processing to detect information related to the exterior of the vehicle based on data or signals from various components of the vehicle control system 100. For example, the vehicle exterior information detector 141 performs processing to detect, identify, and track objects around the vehicle, as well as to detect the distance to the objects. Examples of objects to be detected include vehicles, people, obstacles, structures, roads, traffic lights, traffic signs, and road markings. Furthermore, for example, the vehicle exterior information detector 141 performs processing to detect the vehicle's surrounding environment. Examples of surrounding environments to be detected include weather, temperature, humidity, brightness, and road conditions. The vehicle exterior information detector 141 provides data indicating the results of the detection processing to, for example, the vehicle position estimator 132; the map analyzer 151, traffic rule recognition unit 152, and state recognition unit 153 of the state analyzer 133; and the emergency avoidance unit 171 of the action controller 135.
[0076] The vehicle interior information detector 142 performs processing for detecting information about the interior of the vehicle based on data or signals from each component of the vehicle control system 100. For example, the vehicle interior information detector 142 performs processing for authenticating and identifying the driver, processing for detecting the driver's state, processing for detecting people on board the vehicle, and processing for detecting the interior environment of the vehicle. Examples of the driver's state as a detection target may include physical state, alertness, concentration, fatigue, and line of sight. Examples of the interior environment of the vehicle as a detection target may include temperature, humidity, brightness, and odor. The vehicle interior information detector 142 provides data indicating the results of the detection processing to, for example, the state recognition part 153 of the state analyzer 133 and the emergency avoidance part 171 of the action controller 135.
[0077] The vehicle state detector 143 performs processing to detect the vehicle's state based on data or signals from various components of the vehicle control system 100. Examples of vehicle states to be detected include speed, acceleration, steering angle, the presence and nature of abnormalities, driving operation status, power seat position and tilt, door lock status, and the status of other onboard equipment. The vehicle state detector 143 provides data representing the detection results to, for example, the state identification unit 153 of the state analyzer 133 and the emergency avoidance unit 171 of the action controller 135.
[0078] The own position estimator 132 performs estimation processing of the position or posture of the own vehicle based on data or signals from each component of the vehicle control system 100 (e.g., the external information detector 141, the state recognition part 153 of the state analyzer 133). In addition, the own position estimator 132 generates a local map for estimating the own position (hereinafter referred to as the own position estimation map) as needed. For example, the own position estimation map is a high-precision map using technologies such as simultaneous localization and mapping (SLAM). The own position estimator 132 provides data indicating the results of the estimation processing to, for example, the map analyzer 151 of the state analyzer 133, the traffic rule recognition part 152, and the state recognition part 153. In addition, the own position estimator 132 stores the own position estimation map in the memory 111.
[0079] The state analyzer 133 performs a process of analyzing the state of the vehicle and its surrounding environment. The state analyzer 133 includes a map analyzer 151, a traffic regulation recognition section 152, a state recognition section 153, and a state prediction section 154.
[0080] Map analyzer 151 analyzes various maps stored in storage 111, using data or signals from various components of vehicle control system 100 (such as vehicle position estimator 132 and vehicle exterior information detector 141) as needed, to construct a map containing information necessary for autonomous driving. Map analyzer 151 supplies the constructed map to, for example, traffic regulation recognition section 152, state recognition section 153, and state prediction section 154, as well as route planning section 161, behavior planning section 162, and action planning section 163 of planning section 134.
[0081] The traffic regulation recognition section 152 performs processing to identify traffic regulations surrounding the vehicle based on data or signals from various components of the vehicle control system 100 (such as the vehicle position estimator 132, the vehicle exterior information detector 141, and the map analyzer 151). This recognition process enables the identification of the position and status of traffic lights surrounding the vehicle, the details of traffic regulations surrounding the vehicle, and drivable lanes. The traffic regulation recognition section 152 provides data indicating the results of the recognition process to, for example, the state prediction section 154.
[0082] The state recognition section 153 performs processing to identify the state of the vehicle based on data or signals from various components of the vehicle control system 100 (such as the vehicle position estimator 132, the vehicle exterior information detector 141, the vehicle interior information detector 142, the vehicle state detector 143, and the map analyzer 151). For example, the state recognition section 153 performs processing to identify the state of the vehicle, the state of the vehicle's surroundings, the state of the vehicle's driver, and the like. Furthermore, the state recognition section 153 generates a local map (hereinafter referred to as a state recognition map) for identifying the state of the vehicle's surroundings, as needed. The state recognition map is, for example, an occupancy grid map.
[0083] Examples of the state of the vehicle to be identified include its position, posture, movement (such as speed, acceleration, and direction of movement), and the presence or absence of anomalies and their content. Examples of the state of the vehicle's surrounding environment to be identified include the type and position of stationary objects in the vehicle's surrounding environment; the type, position, and movement (such as speed, acceleration, and direction of movement) of moving objects around the vehicle; the road structure and road surface conditions around the vehicle; and the weather, temperature, humidity, and brightness around the vehicle. Examples of the driver's state to be identified include physical condition, alertness, concentration, fatigue, line of sight movement, and driving operation.
[0084] The state recognition section 153 supplies data indicating the result of the recognition processing (including a state recognition map as necessary) to, for example, the own position estimator 132 and the state prediction section 154. Furthermore, the state recognition section 153 stores the state recognition map in the memory 111.
[0085] The state prediction section 154 performs processing for predicting the state of the vehicle based on data or signals from various components of the vehicle control system 100, such as the map analyzer 151, the traffic regulation recognition section 152, and the state recognition section 153. For example, the state prediction section 154 performs processing for predicting the state of the vehicle, the state of the vehicle's surrounding environment, the state of the driver, and the like.
[0086] Examples of the vehicle's state as a prediction target include the vehicle's behavior, occurrence of an abnormality in the vehicle, and the vehicle's drivable distance. Examples of the vehicle's surrounding environment as a prediction target include the behavior of a moving object, changes in the state of a traffic light, and changes in the vehicle's surrounding environment (such as the weather). Examples of the driver's state as a prediction target include the driver's behavior and physical condition.
[0087] The state prediction section 154 supplies data indicating the prediction processing result together with data from the traffic regulation recognition section 152 and the state recognition section 153 to, for example, the route planning section 161 , the behavior planning section 162 and the action planning section 163 of the planning section 134 .
[0088] The route planning section 161 plans a route to a destination based on data or signals from various components of the vehicle control system 100 (such as the map analyzer 151 and the state prediction section 154). For example, the route planning section 161 sets a route from the current location to the designated destination based on a global map. Furthermore, the route planning section 161 may appropriately change the route based on conditions such as traffic jams, accidents, traffic control, construction, and the driver's physical condition. The route planning section 161 provides data representing the planned route to, for example, the behavior planning section 162.
[0089] The behavior planning section 162 plans the behavior of the vehicle based on data or signals from various components of the vehicle control system 100 (such as the map analyzer 151 and the state prediction section 154) so that the vehicle can safely travel along the route planned by the route planning section 161 within the time planned by the route planning section 161. For example, the behavior planning section 162 plans information regarding starting movement, stopping, travel direction (such as forward movement, backward movement, left turn, right turn, and direction change), lane for travel, travel speed, and overtaking. The behavior planning section 162 provides data representing the planned behavior of the vehicle to the action planning section 163, for example.
[0090] The action planning section 163 plans the actions of the vehicle based on data or signals from various components of the vehicle control system 100 (such as the map analyzer 151 and the state prediction section 154) to achieve the behavior planned by the behavior planning section 162. For example, the action planning section 163 plans acceleration, deceleration, and a travel path. The action planning section 163 provides data representing the planned actions of the vehicle to, for example, the acceleration / deceleration controller 172 and the direction controller 173 of the action controller 135.
[0091] The behavior controller 135 controls the behavior of the vehicle. The behavior controller 135 includes an emergency avoidance section 171, an acceleration / deceleration controller 172, and a direction controller 173.
[0092] Emergency avoidance unit 171 detects emergency events (such as collisions, contact, entry into dangerous areas, driver mishaps, or vehicle mishaps) based on the detection results of exterior vehicle information detector 141, interior vehicle information detector 142, and vehicle state detector 143. When emergency avoidance unit 171 detects an emergency, it plans a vehicle maneuver (such as an emergency stop or sharp turn) to avoid the emergency. Emergency avoidance unit 171 provides data indicating the planned vehicle maneuver to, for example, acceleration / deceleration controller 172 and steering controller 173.
[0093] The acceleration / deceleration controller 172 controls acceleration / deceleration to realize the action of the vehicle planned by the action planning section 163 or the emergency avoidance section 171. For example, the acceleration / deceleration controller 172 calculates control target values for the driving force generating device or the braking device to realize the planned acceleration, planned deceleration, or planned emergency stop, and outputs a control command representing the calculated control target value to the powertrain controller 107.
[0094] The direction controller 173 controls the direction for realizing the maneuver of the vehicle planned by the maneuver planning section 163 or the emergency avoidance section 171. For example, the direction controller 173 calculates a control target value for the steering mechanism for realizing the driving path planned by the maneuver planning section 163 or the sharp turn planned by the emergency avoidance section 171, and provides a control instruction representing the calculated control target value to the powertrain controller 107.
[0095] <Configuration Example of Data Acquisition Section 102A and Vehicle Exterior Information Detector 141A>
[0096] Figure 2 Shown as Figure 1 The data acquisition section 102A of the first embodiment of the data acquisition section 102 in the vehicle control system 100 and the data acquisition section 102A of the first embodiment of the vehicle control system 100 Figure 1 FIG. 1 is a portion of an example of a configuration of a vehicle exterior information detector 141A of the first embodiment of the vehicle exterior information detector 141 in the vehicle control system 100 .
[0097] The data acquisition section 102A includes a camera 201 and a millimeter wave radar 202. The vehicle exterior information detector 141A includes an information processor 211. The information processor 211 includes an image processor 221, a signal processor 222, a geometric transformation section 223, and an object recognition section 224.
[0098] The camera 201 includes an image sensor 201A. Any type of image sensor (e.g., a CMOS image sensor or a CCD image sensor) can be used as the image sensor 201A. The camera 201 (image sensor 201A) captures an image of the area in front of the vehicle 10 and provides the obtained image (hereinafter referred to as a captured image) to the image processor 221.
[0099] The millimeter-wave radar 202 senses an area located in front of the vehicle 10, and the sensing ranges of the millimeter-wave radar 202 and the camera 201 at least partially overlap. For example, the millimeter-wave radar 202 sends a transmission signal including millimeter waves in front of the vehicle 10, and uses a receiving antenna to receive a reception signal that is a signal reflected from an object (reflector) located in front of the vehicle 10. For example, a plurality of receiving antennas are arranged at specified intervals in the lateral direction (width direction) of the vehicle 10. In addition, a plurality of receiving antennas can also be arranged in the height direction. The millimeter-wave radar 202 provides the signal processor 222 with data (hereinafter referred to as millimeter-wave data) indicating the strength of the reception signal received using each receiving antenna in chronological order.
[0100] The image processor 221 performs specified image processing on the captured image. For example, the image processor 221 performs pixel reduction processing or filtering processing on the captured image, and reduces the number of pixels in the captured image (reduces the resolution) according to the image size that can be processed by the object recognition unit 224. The image processor 221 provides the captured image with reduced resolution (hereinafter referred to as a low-resolution image) to the object recognition unit 224.
[0101] The signal processor 222 performs specified signal processing on the millimeter wave data to generate a millimeter wave image, which is an image indicating the result of the sensing performed by the millimeter wave radar 202. Note that the signal processor 222 generates two types of millimeter wave images, for example, a signal intensity image and a speed image. The signal intensity image is a millimeter wave image that represents the position of each object located in front of the vehicle 10 and the intensity of the signal (received signal) reflected from the object. The speed image is a millimeter wave image that represents the position of each object located in front of the vehicle 10 and the relative speed of the object with respect to the vehicle 10. The signal processor 222 provides the signal intensity image and the speed image to the geometric transformation part 223.
[0102] The geometric transformation part 223 performs a geometric transformation on the millimeter wave image to transform the millimeter wave image into an image whose coordinate system is the same as the coordinate system of the captured image. In other words, the geometric transformation part 223 transforms the millimeter wave image into an image obtained when viewed from the same viewpoint as the captured image (hereinafter referred to as a geometrically transformed millimeter wave image). More specifically, the geometric transformation part 223 transforms the coordinate systems of the signal intensity image and the velocity image from the coordinate system of the millimeter wave image to the coordinate system of the captured image. Note that the signal intensity image and the velocity image on which the geometric transformation has been performed are referred to as a geometrically transformed signal intensity image and a geometrically transformed velocity image, respectively. The geometric transformation part 223 provides the geometrically transformed signal intensity image and the geometrically transformed velocity image to the object recognition part 224.
[0103] The object recognition section 224 performs processing to recognize an object located in front of the vehicle 10 based on the low-resolution image, the geometrically transformed signal intensity image, and the geometrically transformed speed image. The object recognition section 224 provides data indicating the recognition result of the object to, for example, the own position estimator 132; the map analyzer 151, the traffic regulation recognition section 152, and the state recognition section 153 of the state analyzer 133; and the emergency avoidance section 171 of the action controller 135. The data indicating the recognition result of the object includes, for example, the position and size of the object in the captured image, as well as the type of the object.
[0104] Note that the target object is an object to be identified by object recognition section 224, and any object can be set as the target object. However, it is advantageous to set the target object as an object that includes a portion having a high reflectivity for the transmission signal of millimeter wave radar 202. The following describes a case where the target object is a vehicle as an example.
[0105] <Configuration Example of Object Recognition Model 251>
[0106] Figure 3 An example of the configuration of the object recognition model 251 used for the object recognition section 224 is shown.
[0107] Object recognition model 251 is a model obtained through machine learning. Specifically, object recognition model 251 is a model obtained through deep learning, a type of machine learning that uses deep neural networks. More specifically, object recognition model 251 is composed of a single shot multibox detector (SSD), which is one type of object recognition model using deep neural networks. Object recognition model 251 includes a feature extraction unit 261 and a recognition unit 262.
[0108] The feature amount extraction section 261 includes VGG16 271 a to VGG16 271 c as convolution layers using a convolutional neural network, and an adder 272 .
[0109] The VGG16 271 a extracts feature quantities of the captured image Pa and generates a feature map (hereinafter referred to as a captured image feature map) that two-dimensionally represents the distribution of the feature quantities. The VGG16 271 a supplies the captured image feature map to the adder 272 .
[0110] The VGG16 271 b extracts the feature quantity of the geometrically transformed signal intensity image Pb and generates a feature map (hereinafter referred to as a signal intensity image feature map) that two-dimensionally represents the distribution of the feature quantity. The VGG16 271 b supplies the signal intensity image feature map to the adder 272 .
[0111] The VGG16 271 c extracts the feature quantity of the geometrically transformed velocity image Pc and generates a feature map (hereinafter referred to as a velocity image feature map) that two-dimensionally represents the distribution of the feature quantity. The VGG16 271 c supplies the velocity image feature map to the adder 272 .
[0112] The adder 272 adds the feature map of the captured image, the signal intensity image feature map, and the speed image feature map to generate a composite feature map, and provides the composite feature map to the recognition section 262 .
[0113] The recognition part 262 includes a convolutional neural network. Specifically, the recognition part 262 includes convolutional layers 273a to 273c.
[0114] The convolution layer 273a performs a convolution operation on the synthetic feature map. The convolution layer 273a performs a process of identifying an object based on the synthetic feature map on which the convolution operation has been performed. The convolution layer 273a provides the synthetic feature map on which the convolution operation has been performed to the convolution layer 273b.
[0115] The convolution layer 273b performs a convolution operation on the synthetic feature map provided by the convolution layer 273a. The convolution layer 273b performs a process of identifying an object based on the synthetic feature map on which the convolution operation has been performed. The convolution layer 273a provides the synthetic feature map on which the convolution operation has been performed to the convolution layer 273c.
[0116] The convolution layer 273c performs a convolution operation on the synthetic feature map provided by the convolution layer 273b. The convolution layer 273b performs a process of identifying an object based on the synthetic feature map on which the convolution operation has been performed.
[0117] The object recognition model 251 outputs data indicating the result of recognition of the object performed by the convolutional layers 273 a to 273 c .
[0118] Note that the size (number of pixels) of the synthetic feature map becomes smaller in order from the convolution layer 273a, and is smallest in the convolution layer 273c. In addition, if the synthetic feature map has a larger size, an object with a small size when viewed from the vehicle 10 is recognized with higher accuracy, and if the synthetic feature map has a smaller size, an object with a large size when viewed from the vehicle 10 is recognized with higher accuracy. Therefore, for example, when the object is a vehicle, a small vehicle at a distance is easily recognized in a synthetic feature map with a large size, and a large vehicle nearby is easily recognized in a synthetic feature map with a small size.
[0119] <Configuration Example of Learning System 301>
[0120] Figure 4 is a block diagram illustrating an example of the configuration of the learning system 301 .
[0121] Learning Systems 301 Figure 3 The object recognition model 251 performs learning processing. The learning system 301 includes an input part 311, an image processor 312, a correct answer data generator 313, a signal processor 314, a geometric transformation part 315, a training data generator 316 and a learning part 317.
[0122] The input section 311 includes various input devices and is used, for example, to input data required for generating training data and operations performed by the user. For example, when a captured image is input, the input section 311 provides the captured image to the image processor 312. For example, when millimeter wave data is input, the input section 311 provides the millimeter wave data to the signal processor 314. For example, the input section 311 provides data indicating user instructions input through operations performed by the user to the correct answer data generator 313 and the training data generator 316.
[0123] Image processor 312 performs the Figure 2 The image processor 312 performs a process similar to that performed by the image processor 221. In other words, the image processor 312 performs a specified image process on the captured image to generate a low-resolution image. The image processor 312 provides the low-resolution image to the correct answer data generator 313 and the training data generator 316.
[0124] Correct answer data generator 313 generates correct answer data based on the low-resolution image. For example, a user specifies the position of a vehicle in the low-resolution image through input section 311. Correct answer data generator 313 generates correct answer data indicating the vehicle's position in the low-resolution image based on the vehicle's position specified by the user. Correct answer data generator 313 provides the correct answer data to training data generator 316.
[0125] Signal processor 314 performs the Figure 2 The signal processor 314 performs a similar process to the process performed by the signal processor 222. In other words, the signal processor 314 performs a specified signal process on the millimeter wave data to generate a signal intensity image and a velocity image. The signal processor 314 provides the signal intensity image and the velocity image to the geometric transformation section 315.
[0126] The geometric transformation section 315 performs the Figure 2 The same process is performed by the geometric transformation section 223. In other words, the geometric transformation section 315 geometrically transforms the signal intensity image and the velocity image. The geometric transformation section 315 provides the geometrically transformed signal intensity image and the geometrically transformed velocity image obtained by the geometric transformation to the training data generator 316.
[0127] The training data generator 316 generates training data including input data and correct answer data, the input data including the low-resolution image, the geometrically transformed signal intensity image, and the geometrically transformed velocity image, and provides the training data to the learning part 317 .
[0128] The learning section 317 uses the training data to perform a learning process on the object recognition model 251. The learning section 317 outputs the object recognition model 251 on which learning has been performed.
[0129] <Learning Process for Object Recognition Model>
[0130] Next, refer to Figure 5 The flowchart of FIG. 3 describes the learning process of the object recognition model performed by the learning system 301.
[0131] Note that the data used to generate training data is collected before starting this process. For example, while vehicle 10 is actually traveling, camera 201 and millimeter-wave radar 202 installed on vehicle 10 sense the area in front of vehicle 10. Specifically, camera 201 captures an image of the area in front of vehicle 10 and stores the captured image in memory 111. Millimeter-wave radar 202 detects objects in front of vehicle 10 and stores the acquired millimeter-wave data in memory 111. Training data is generated based on the captured image and the millimeter-wave data accumulated in memory 111.
[0132] In step S1 , the learning system 301 generates training data.
[0133] For example, the user inputs a captured image and millimeter wave data acquired substantially simultaneously to the learning system 301 through the input section 311. In other words, the captured image and millimeter wave data obtained by performing sensing at substantially the same time are input to the learning system 301. The captured image is provided to the image processor 312, and the millimeter wave data is provided to the signal processor 314.
[0134] The image processor 312 performs image processing such as number reduction processing on the captured image and generates a low-resolution image. The image processor 312 provides the low-resolution image to the correct answer data generator 313 and the training data generator 316.
[0135] Figure 6 An example of a low-resolution image is shown.
[0136] The correct answer data generator 313 generates correct answer data indicating the position of the object in the low-resolution image based on the position of the object specified by the user through the input unit 311. The correct answer data generator 313 supplies the correct answer data to the training data generator 316.
[0137] Figure 7 Shown for Figure 6 An example of correct answer data generated from a low-resolution image. The white box area indicates the location of the vehicle, which is the object.
[0138] The signal processor 314 performs specified signal processing on the millimeter wave data to estimate the position and speed of the object that reflected the transmission signal in the area in front of the vehicle 10. The position of the object is represented by, for example, the distance from the vehicle 10 to the object and the direction (angle) of the object relative to the optical axis direction of the millimeter wave radar 202 (the direction of travel of the vehicle 10). Note that, for example, when the transmission signal is radially transmitted, the optical axis direction of the millimeter wave radar 202 is the same as the direction of the center of the range in which the radial transmission is performed, and when scanning is performed using the transmission signal, the optical axis direction of the millimeter wave radar 202 is the same as the direction of the center of the range in which the scanning is performed. The speed of the object is represented by, for example, the relative speed of the object relative to the vehicle 10.
[0139] The signal processor 314 generates a signal intensity image and a velocity image based on the result of estimating the position and velocity of the object, and supplies the signal intensity image and the velocity image to the geometric transformation part 315 .
[0140] Figure 8An example of a signal strength image is shown. The x-axis of the signal strength image represents the lateral direction (the width direction of vehicle 10), and the y-axis of the signal strength image represents the optical axis direction of millimeter-wave radar 202 (the direction of travel of vehicle 10, the depth direction). The signal strength image shows the positions of objects in front of vehicle 10 and the distribution of reflection intensity of each object, that is, the distribution of the intensity of the received signal reflected from the objects in front of vehicle 10, using a bird's-eye view.
[0141] Note that, similar to the case of the signal intensity image, the speed image is an image in which the positions of objects located in front of the vehicle 10 and the distribution of the relative speeds of the objects are given using a bird's-eye view, although illustration thereof is omitted.
[0142] The geometric transformation unit 315 performs a geometric transformation on the signal intensity image and the velocity image, converting them into images whose coordinate systems are the same as the coordinate system of the captured image, thereby generating a geometrically transformed signal intensity image and a geometrically transformed velocity image. The geometric transformation unit 315 provides the geometrically transformed signal intensity image and the geometrically transformed velocity image to the training data generator 316.
[0143] Figure 9 Examples of a geometrically transformed signal intensity image and a geometrically transformed velocity image are shown. Figure 9 A shows an example of a geometrically transformed signal intensity image, Figure 9 B shows an example of a geometrically transformed velocity image. Note that Figure 9 The geometrically transformed signal intensity image and the geometrically transformed velocity image in the image are based on the Figure 7 The low-resolution image is generated by acquiring the millimeter-wave data substantially simultaneously with the captured image.
[0144] In the geometrically transformed signal intensity image, parts with higher signal intensity are brighter, and parts with lower signal intensity are darker. In the geometrically transformed velocity image, parts with higher relative velocity are brighter, parts with lower relative velocity are darker, and parts where relative velocity cannot be detected (no object exists) are pure black.
[0145] As described above, when geometric transformation is performed on the millimeter wave image (signal intensity image and velocity image), not only the position of the object in the lateral and depth directions but also the position of the object in the height direction is given.
[0146] However, for the millimeter wave radar 202, the resolution in the height direction becomes lower as the distance increases. Therefore, the height of a distant object may be detected to be higher than its actual height.
[0147] On the other hand, when geometric transformation section 315 performs geometric transformation on millimeter wave images, it limits the height of objects located at a specified distance or greater. Specifically, when geometric transformation is performed on millimeter wave images, if an object located at a specified distance or greater has a height greater than a specified upper limit, geometric transformation section 315 limits the height of the object to the upper limit to perform the geometric transformation. This prevents misidentification caused by detecting a distant vehicle as being higher than its actual height, for example, when the object is a vehicle.
[0148] The training data generator 316 generates training data including input data and correct answer data, the input data including the captured image, the geometrically transformed signal intensity image, and the geometrically transformed velocity image, and provides the generated training data to the learning part 317.
[0149] In step S2, the learning section 317 causes the object recognition model 251 to perform learning. Specifically, the learning section 317 inputs the input data included in the training data to the object recognition model 251. The object recognition model 251 performs processing to recognize the object and outputs data indicating the recognition result. The learning section 317 compares the recognition result performed by the object recognition model 251 with the correct answer data and adjusts, for example, the parameters of the object recognition model 251 to reduce the error.
[0150] In step S3, the learning section 317 determines whether to continue learning. For example, when the learning performed by the object recognition model 251 has not ended, the learning section 317 determines to continue learning, and the process returns to step S1.
[0151] Thereafter, the processes of steps S1 to S3 are repeatedly performed until it is determined in step S3 that the learning is to be terminated.
[0152] On the other hand, in step S3 , the learning section 317 determines that when, for example, the learning has ended, the learning performed by the object recognition model 251 is terminated, and the learning process performed on the object recognition model is terminated.
[0153] As described above, the object recognition model 251 on which learning has been performed is generated.
[0154] Notice, Figure 10 An example of the result of recognition performed by the object recognition model 251 that performs learning using only millimeter wave data without using captured images is shown.
[0155] Figure 10 FIG. 8A shows an example of a geometrically transformed signal intensity image generated based on millimeter wave data.
[0156] Figure 10 B shows an example of the result of recognition performed by the object recognition model 251. Specifically, Figure 10 The viewpoint transformed brightness image of A is superimposed on the image produced by Figure 10 The millimeter wave data of the viewpoint-converted luminance image A are acquired substantially simultaneously on the captured image, and the frame area indicates the position of the vehicle where the object has been recognized.
[0157] As shown in this example, the object recognition model 251 also enables recognition of a vehicle as a target object with an accuracy not less than a specified accuracy when using only millimeter wave data (a geometrically transformed signal intensity image and a geometrically transformed speed image).
[0158] <Object Recognition Processing>
[0159] Next, refer to Figure 11 The flowchart of FIG. 1 describes the object recognition process performed by the vehicle 10 .
[0160] This process is initiated, for example, when an operation is performed to start the vehicle 10 and start driving, that is, when the ignition switch, power switch, start switch, etc. of the vehicle 10 is turned on. Furthermore, this process is terminated, for example, when an operation is performed to terminate the driving of the vehicle 10, that is, when the ignition switch, power switch, start switch, etc. of the vehicle 10 is turned off.
[0161] In step S101 , the camera 201 and the millimeter-wave radar 202 perform sensing on an area located in front of the vehicle 10 .
[0162] Specifically, the camera 201 captures an image of an area located in front of the vehicle 10 and provides the obtained captured image to the image processor 221 .
[0163] Millimeter wave radar 202 transmits a signal in the front direction of vehicle 10 and receives a received signal using a plurality of receiving antennas, which is a signal reflected from an object located in front of vehicle 10. Millimeter wave radar 202 provides signal processor 222 with millimeter wave data indicating the strength of the received signal received using each receiving antenna in time sequence.
[0164] In step S102 , the image processor 221 performs pre-processing on the captured image. Specifically, the image processor 221 performs, for example, a reduced number of processes on the captured image to generate a low-resolution image, and supplies the low-resolution image to the object recognition section 224 .
[0165] In step S103, the signal processor 222 generates a millimeter wave image. Specifically, the signal processor 222 performs the same Figure 5The signal processor 222 performs a similar process to the process performed by the signal processor 314 in step S1 to generate a signal intensity image and a velocity image based on the millimeter wave data. The signal processor 222 supplies the signal intensity image and the velocity image to the geometric transformation section 223.
[0166] In step S104, the geometric transformation section 223 performs geometric transformation on the millimeter wave image. Specifically, the geometric transformation section 223 performs the same Figure 5 The geometric transformation section 315 performs a similar process to the process performed in step S1 to transform the signal intensity image and the velocity image into a geometrically transformed signal intensity image and a geometrically transformed velocity image. The geometric transformation section 223 provides the geometrically transformed signal intensity image and the geometrically transformed velocity image to the object recognition section 224.
[0167] In step S105, the object recognition unit 224 identifies an object based on the low-resolution image and the geometrically transformed millimeter-wave image. Specifically, the object recognition unit 224 inputs input data, including the low-resolution image, the geometrically transformed signal intensity image, and the geometrically transformed velocity image, to the object recognition model 251. The object recognition model 251 identifies an object located in front of the vehicle 10 based on the input data.
[0168] Figure 12 An example of the recognition result when the object is a vehicle is shown. Figure 12 A shows an example of a captured image. Figure 12 B shows an example of the result of identifying a vehicle. Figure 12 In B, the area where the vehicle is recognized is framed.
[0169] The object recognition section 224 provides data indicating the recognition result of the object to, for example, the own position estimator 132; the map analyzer 151, traffic regulation recognition section 152 and state recognition section 153 of the state analyzer 133; and the emergency avoidance section 171 of the action controller 135.
[0170] The own position estimator 132 executes a process of estimating the position, posture, etc. of the vehicle 10 based on, for example, the recognition result of the object.
[0171] The map analyzer 151 performs a process of analyzing various maps stored in the storage section 111 based on the recognition result of the object, for example, and constructs a map including information required for the autonomous driving process.
[0172] The traffic regulation recognition section 152 performs a process of recognizing traffic regulations of the surrounding environment of the vehicle 10 based on, for example, the recognition result of the object.
[0173] The state recognition portion 153 performs a process of recognizing the state of the surrounding environment of the vehicle 10 based on the recognition result of the object, for example.
[0174] When the emergency avoidance portion 171 detects the occurrence of an emergency based on, for example, the recognition result of the object, the emergency avoidance portion 171 plans an action of the vehicle 10 , such as a sudden stop or a quick turn, to avoid the emergency.
[0175] Thereafter, the process returns to step S101 , and the processes of step S101 and thereafter are performed.
[0176] As described above, the accuracy of recognizing an object located in front of the vehicle 10 can be improved.
[0177] Specifically, when only the captured image is used to perform the process of identifying the object, the accuracy of identifying the object is reduced due to bad weather (such as rain or fog), at night, or under poor visual conditions caused by, for example, obstacles. On the other hand, when using millimeter wave radar, the accuracy of identifying the object is hardly reduced due to bad weather, at night, or under poor visual conditions caused by, for example, obstacles. Therefore, the camera 201 and the millimeter wave radar 202 (captured image and millimeter wave data) are fused to perform the process of identifying the object, which makes it possible to compensate for the defects caused when only the captured image is used. This leads to improved recognition accuracy.
[0178] In addition, if Figure 13 As shown in FIG. 1A , the captured image is represented by a coordinate system defined by an x-axis and a z-axis. The x-axis represents the lateral direction (the width direction of the vehicle 10), and the z-axis represents the height direction. Figure 13 As shown in FIG. 1B , the millimeter wave image is represented by a coordinate system defined by an x-axis and a y-axis. The x-axis is similar to the x-axis of the coordinate system of the captured image. Note that the x-axis extends in the same direction as the direction in which the transmitted signal of the millimeter wave radar 202 is spread out in a planar manner. The y-axis represents the optical axis direction of the millimeter wave radar 202 (the direction of travel of the vehicle 10, the depth direction).
[0179] When there is a difference in coordinate system between the captured image and the millimeter wave image as described above, it makes it difficult to understand the correlation between the captured image and the millimeter wave image. For example, it is difficult to match each pixel of the captured image with a reflection point (a point where the intensity of the received signal is high) in the millimeter wave image. Therefore, when the object recognition model 251 is made by using Figure 13 When deep learning of the captured images and millimeter wave images of A is used to perform learning, the difficulty of learning will increase, and this may lead to a decrease in the accuracy of learning.
[0180] On the other hand, according to the present technology, a geometric transformation is performed on the millimeter wave image (signal intensity image and velocity image) to obtain an image (a geometrically transformed signal intensity image and a geometrically transformed velocity image) whose coordinate system matches the coordinate system of the captured image, and the object recognition model 251 is caused to perform learning using the obtained image. This results in facilitating the matching of each pixel of the captured image with the reflection point in the millimeter wave image, and leads to improved learning accuracy. In addition, in actual vehicle recognition processing, the use of the geometrically transformed signal intensity image and the geometrically transformed velocity image leads to improved object recognition accuracy.
[0181] <<2. Second embodiment>>
[0182] Next, refer to Figures 14 to 16 A second embodiment of the present technology is described.
[0183] <Configuration Example of Vehicle Exterior Information Detector 141B>
[0184] Figure 14 Shown as Figure 1 An example of the configuration of the vehicle exterior information detector 141B of the second embodiment of the vehicle exterior information detector 141 of the vehicle control system 100. Figure 2 The part corresponding to the part in Figure 2 The same reference numerals are denoted, and descriptions thereof are appropriately omitted.
[0185] The vehicle exterior information detector 141B includes an information processor 401. The information processor 401 and Figure 2 Information processor 401 is similar to information processor 211 in that it includes signal processor 222 and geometric transformation section 223. On the other hand, information processor 401 differs from information processor 211 in that it includes image processor 421 and object recognition section 422 instead of image processor 221 and object recognition section 224, and in that a synthesizer 423 is added. Object recognition section 422 includes object recognition section 431a and object recognition section 431b.
[0186] The image processor 421 generates a low-resolution image based on the captured image, similarly to the case of the image processor 221. The image processor 421 supplies the low-resolution image to the object recognition section 431a.
[0187] In addition, the image processor 421 cuts out a portion of the captured image according to the image size that the object recognition section 431b can process. The image processor 421 supplies the image cut out from the captured image (hereinafter referred to as a cropped image) to the object recognition section 431b.
[0188] With Figure 2The situation of the object recognition part 224 is similar, Figure 3 The object recognition model 251 is used for the object recognition section 431a and the object recognition section 431b.
[0189] and Figure 2 Similar to the case of the object recognition section 224, the object recognition section 431a performs processing for recognizing an object located in front of the vehicle 10 based on the low-resolution image, the geometrically transformed signal intensity image, and the geometrically transformed speed image. The object recognition section 431a provides data indicating the processing result of recognizing the object to the synthesizer 423.
[0190] The object recognition section 431b performs processing for recognizing an object located in front of the vehicle 10 based on the cropped image, the geometrically transformed signal intensity image, and the geometrically transformed speed image. The object recognition section 431b supplies data indicating the processing result of recognizing the object to the synthesizer 423.
[0191] Note that, although a detailed description thereof is omitted, learning processing is performed on each of the object recognition model 251 used by the object recognition section 431a and the object recognition model 251 used by the object recognition section 431b using different training data. Specifically, the object recognition model 251 used by the object recognition section 431a is caused to perform learning using training data comprising input data including a low-resolution image, a geometrically transformed signal intensity image, and a geometrically transformed velocity image, and correct answer data generated based on the low-resolution image. On the other hand, the object recognition model 251 used by the object recognition section 431B is caused to perform learning using training data comprising input data including a cropped image, a geometrically transformed signal intensity image, and a geometrically transformed velocity image, and correct answer data generated based on the cropped image.
[0192] The synthesizer 423 synthesizes the object recognition result obtained by the object recognition section 431a and the object recognition result obtained by the object recognition section 431b. The synthesizer 423 provides data indicating the object recognition result obtained by the synthesis to, for example, the own position estimator 132; the map analyzer 151, the traffic regulation recognition section 152, and the state recognition section 153 of the state analyzer 133; and the emergency avoidance section 171 of the action controller 135.
[0193] <Object Recognition Processing>
[0194] Next, refer to Figure 15 The flowchart of the second embodiment describes the object recognition process.
[0195] This process is initiated, for example, when an operation is performed to start the vehicle 10 and start traveling, that is, when the ignition switch, power switch, starter switch, etc. of the vehicle 10 is turned on. Furthermore, this process is terminated, for example, when an operation is performed to terminate driving of the vehicle 10, that is, when the ignition switch, power switch, starter switch, etc. of the vehicle 10 is turned off.
[0196] In step S201, Figure 11 The processing of step S101 is similar to that of step S101 , and sensing is performed on the area located in front of the vehicle 10 .
[0197] In step S202, the image processor 421 performs pre-processing on the captured image. Specifically, the image processor 421 performs Figure 11 The image processor 421 performs a process similar to the process performed in step S102 to generate a low-resolution image based on the captured image. The image processor 421 supplies the low-resolution image to the object recognition section 431a.
[0198] Furthermore, the image processor 421 detects, for example, a vanishing point of a road in the captured image. From the captured image, the image processor 421 cuts out an image in a rectangular area of a specified size centered on the vanishing point. The image processor 421 supplies the cropped image obtained by cutting out to the object recognition unit 431b.
[0199] Figure 16 An example of the relationship between the captured image and the cropped image is shown. Specifically, Figure 16 A shows an example of the captured image. In addition, in the captured image, Figure 16 A rectangular area framed by a dotted line in B is cut out as a cropped image. The rectangular area has a specified size and is centered at the vanishing point of the road.
[0200] In step S203, Figure 11 The process of step S103 is similar to that of generating a millimeter wave image, ie, a signal intensity image and a velocity image.
[0201] In step S204, Figure 11 Similar to the processing of step S104, a geometric transformation is performed on the millimeter wave image. This results in the generation of a geometrically transformed signal intensity image and a geometrically transformed velocity image. The geometric transformation section 223 provides the geometrically transformed signal intensity image and the geometrically transformed velocity image to the object recognition section 431a and the object recognition section 431b.
[0202] In step S205, the object recognition section 431a performs a process of recognizing an object based on the low-resolution image and the geometrically transformed millimeter-wave image. Specifically, the object recognition section 431a performs a process of recognizing an object based on the low-resolution image and the geometrically transformed millimeter-wave image. Figure 11The object recognition section 431a performs a process similar to the process performed in step S105 to recognize an object located in front of the vehicle 10 based on the low-resolution image, the geometrically transformed signal intensity image, and the geometrically transformed speed image. The object recognition section 431a provides the synthesizer 423 with data indicating the result of the process of recognizing the object.
[0203] In step S206, object recognition section 431b identifies the vehicle based on the cropped image and the geometrically transformed millimeter-wave image. Specifically, object recognition section 224 inputs input data, including the cropped image, the geometrically transformed signal intensity image, and the geometrically transformed velocity image, to object recognition model 251. Object recognition model 251 identifies the object located in front of vehicle 10 based on the input data. Object recognition section 431b provides data indicating the object recognition result to synthesizer 423.
[0204] In step S207, the synthesizer 423 synthesizes the object recognition results. Specifically, the synthesizer 423 synthesizes the object recognition results performed by the object recognition unit 431a and the object recognition results performed by the object recognition unit 431b. The synthesizer 423 provides data indicating the object recognition results obtained through synthesis to, for example, the self-position estimator 132; the map analyzer 151, traffic regulation recognition unit 152, and state recognition unit 153 of the state analyzer 133; and the emergency avoidance unit 171 of the action controller 135.
[0205] The own position estimator 132 executes a process of estimating the position, posture, etc. of the vehicle 10 based on, for example, the recognition result of the object.
[0206] The map analyzer 151 performs a process of analyzing various maps stored in the storage section 111 based on the recognition result of the object, for example, and constructs a map including information required for the autonomous driving process.
[0207] The traffic regulation recognition section 152 performs a process of recognizing traffic regulations of the surrounding environment of the vehicle 10 based on, for example, the recognition result of the object.
[0208] The state recognition portion 153 performs a process of recognizing the state of the surrounding environment of the vehicle 10 based on the recognition result of the object, for example.
[0209] When the emergency avoidance portion 171 detects the occurrence of an emergency based on, for example, the recognition result of the object, the emergency avoidance portion 171 plans an action of the vehicle 10 , such as a sudden stop or a quick turn, to avoid the emergency.
[0210] Thereafter, the process returns to step S201 , and the processes of step S201 and thereafter are performed.
[0211] As described above, it is possible to improve the accuracy of identifying objects located in front of vehicle 10. Specifically, using a low-resolution image instead of a captured image particularly reduces the accuracy of identifying distant objects. However, when the object is identified using a high-resolution cropped image obtained by cropping the image of the area around the vanishing point of the road, this makes it possible to improve the accuracy of identifying distant vehicles located around the vanishing point when the object is a vehicle.
[0212] <<3. Third embodiment>>
[0213] Next, refer to Figure 17 A third embodiment of the present technology is described.
[0214] <Configuration Example of Data Acquisition Section 102B and Vehicle Exterior Information Detector 141C>
[0215] Figure 17 Shown as Figure 1 The data acquisition section 102B of the second embodiment of the data acquisition section 102 of the vehicle control system 100 and the data acquisition section 102B of the second embodiment of the vehicle control system 100 Figure 1 An example of the configuration of the vehicle exterior information detector 141C of the third embodiment of the vehicle exterior information detector 141 in the vehicle control system 100 of FIG. Figure 2 The part corresponding to the part in Figure 2 The same reference numerals are denoted, and descriptions thereof are appropriately omitted.
[0216] The data acquisition section 102B is similar to the data acquisition section 102 in that it includes a camera 201 and a millimeter wave radar 202. Figure 2 The data acquisition portion 102A is similar to and includes the LiDAR 501. Figure 2 The data acquisition portion 102A is different.
[0217] The vehicle exterior information detector 141C includes an information processor 511. The information processor 511 and Figure 2 The information processor 511 is similar to the information processor 211 in that it includes an image processor 221, a signal processor 222, and a geometric transformation section 223. On the other hand, the information processor 511 is different from the information processor 211 in that it includes an object recognition section 523 instead of the object recognition section 224, and in that the signal processor 521 and the geometric transformation section 522 are added.
[0218] The LiDAR 501 performs sensing with respect to the area in front of the vehicle 10, and the sensing ranges of the LiDAR 501 and the camera 201 at least partially overlap. For example, the LiDAR 501 performs scanning in the lateral and height directions relative to the area in front of the vehicle 10 using laser pulses, and receives reflected light as a reflection of the laser pulses. The LiDAR 501 calculates the distance to the object in front of the vehicle 10 based on the time it takes to receive the reflected light, and based on the calculation results, the LiDAR 501 generates three-dimensional point group data (point cloud) indicating the shape and position of the object in front of the vehicle 10. The LiDAR 501 provides this point group data to the signal processor 521.
[0219] The signal processor 521 performs prescribed signal processing (for example, interpolation processing or number reduction processing) on the point group data, and supplies the point group data on which the signal processing has been performed to the geometric transformation section 522 .
[0220] The geometric transformation section 522 performs geometric transformation on the point group data to generate a two-dimensional image (hereinafter referred to as two-dimensional point group data) whose coordinate system is the same as the coordinate system of the captured image. The geometric transformation section 522 supplies the two-dimensional point group data to the object recognition section 523.
[0221] The object recognition section 523 performs processing to recognize an object located in front of the vehicle 10 based on the low-resolution image, the geometrically transformed signal intensity image, the geometrically transformed velocity image, and the two-dimensional point cloud data. The object recognition section 523 provides data indicating the recognition result of the object to, for example, the self-position estimator 132; the map analyzer 151, the traffic regulation recognition section 152, and the state recognition section 153 of the state analyzer 133; and the emergency avoidance section 171 of the action controller 135. The data indicating the recognition result of the object also includes, for example, the position and size of the object in the captured image and the type of the object.
[0222] Note that, with e.g. Figure 3 An object recognition model with a configuration similar to that of object recognition model 251 is used in object recognition section 523, although a detailed description thereof is omitted. However, a VGG16 is added to the two-dimensional point cloud data. The object recognition model of object recognition section 523 is then trained using training data comprising input data including a low-resolution image, a geometrically transformed signal intensity image, a geometrically transformed velocity image, and two-dimensional point cloud data, and correct answer data generated based on the low-resolution image.
[0223] As described above, the addition of the LiDAR 501 leads to further improvement in the accuracy of identifying the object.
[0224] <<4. Fourth embodiment>>
[0225] Next, refer to Figure 18 A fourth embodiment of the present technology is described.
[0226] <Configuration Example of Vehicle Exterior Information Detector 141D>
[0227] Figure 18 Shown as Figure 2 An example of the configuration of the vehicle exterior information detector 141D of the fourth embodiment of the vehicle exterior information detector 141 in the vehicle control system 100. Figure 2 The part corresponding to the part in Figure 2 The same reference numerals are denoted, and descriptions thereof are appropriately omitted.
[0228] The vehicle exterior information detector 141D includes an information processor 611. The information processor 611 and Figure 2 The information processor 611 is similar to the information processor 211 in that it includes an image processor 221 and a signal processor 222. On the other hand, the information processor 611 is different from the information processor 211 in that it includes an object recognition section 622 instead of the object recognition section 224, and in that a geometric transformation section 621 is added and the geometric transformation section 223 is removed.
[0229] The geometric transformation section 621 performs a geometric transformation on the low-resolution image provided by the image processor 221 to transform the low-resolution image into an image (hereinafter referred to as a geometrically transformed low-resolution image) having the same coordinate system as the millimeter wave image output by the signal processor 222. For example, the geometric transformation section 621 transforms the low-resolution image into an image of a bird's-eye view. The geometric transformation section 621 provides the geometrically transformed low-resolution image to the object recognition section 622.
[0230] The object recognition section 622 performs processing for recognizing an object located in front of the vehicle 10 based on the geometrically transformed low-resolution image and the signal strength image and speed image provided by the signal processor 222. The object recognition section 622 provides data indicating the recognition result of the object to, for example, the own position estimator 132; the map analyzer 151, the traffic regulation recognition section 152, and the state recognition section 153 of the state analyzer 133; and the emergency avoidance section 171 of the action controller 135.
[0231] As described above, the captured image can be transformed into an image having the same coordinate system as that of the millimeter wave image to perform a process of recognizing an object.
[0232] <<5. Fifth embodiment>>
[0233] Next, refer to Figure 19 A fifth embodiment of the present technology is described.
[0234] <Configuration Example of Vehicle Exterior Information Detector 141E>
[0235] Figure 19 Shown as Figure 2 An example of the configuration of the vehicle exterior information detector 141E of the fifth embodiment of the vehicle exterior information detector 141 in the vehicle control system 100. Figure 17 The part corresponding to the part in Figure 17 The same reference numerals are denoted, and descriptions thereof are appropriately omitted.
[0236] The vehicle exterior information detector 141E includes an information processor 711. The information processor 711 and Figure 17 The information processor 711 is similar to the information processor 511 in that it includes the image processor 221, the signal processor 222, and the signal processor 521. On the other hand, the information processor 711 is different from the information processor 511 in that it includes a geometric transformation section 722 and an object recognition section 723 instead of the geometric transformation section 223 and the object recognition section 523, and in that the geometric transformation section 721 is added and the geometric transformation section 522 is removed.
[0237] The geometric transformation section 721 performs geometric transformation on the low-resolution image supplied from the image processor 221 to transform the low-resolution image into three-dimensional point group data (hereinafter referred to as point group image data) having the same coordinate system as the point group data output from the signal processor 521. The geometric transformation section 721 supplies the point group image data to the object recognition section 723.
[0238] The geometric transformation section 722 performs geometric transformation on the signal intensity image and the velocity image provided by the signal processor 222 to transform the signal intensity image and the velocity image into three-dimensional point group data (hereinafter referred to as point group signal intensity data and point group velocity data) having the same coordinate system as the point group data output by the signal processor 521. The geometric transformation section 722 provides the point group signal intensity data and the point group velocity data to the object recognition section 723.
[0239] The object recognition section 723 performs processing for recognizing an object located in front of the vehicle 10 based on the point group image data, point group signal strength data, point group velocity data, and point group data generated by the sensing performed by the LiDAR 501 and provided by the signal processor 521. The object recognition section 723 provides data indicating the recognition result of the object to, for example, the own position estimator 132; the map analyzer 151, the traffic regulation recognition section 152, and the state recognition section 153 of the state analyzer 133; and the emergency avoidance section 171 of the action controller 135.
[0240] As described above, the captured image and millimeter wave image can be transformed into a plurality of point group data to perform a process of recognizing an object.
[0241] <<6. Sixth embodiment>>
[0242] Next, refer to Figure 20 A sixth embodiment of the present technology is described.
[0243] <Configuration Example of Data Acquisition Section 102C and Vehicle Exterior Information Detector 141F>
[0244] Figure 20 Shown as Figure 1 The data acquisition section 102C of the third embodiment of the data acquisition section 102 in the vehicle control system 100 and the data acquisition section 102C of the third embodiment of the vehicle control system 100 Figure 1 An example of the configuration of the vehicle exterior information detector 141F of the sixth embodiment of the vehicle exterior information detector 141 in the vehicle control system 100 of FIG. Figure 2 The part corresponding to the part in Figure 2 The same reference numerals are denoted, and descriptions thereof are appropriately omitted.
[0245] Data acquisition section 102C and Figure 2 The data acquisition section 102A is similar to the data acquisition section 102A in that the camera 201 is included, but is different from the data acquisition section 102A in that the millimeter wave radar 811 is included instead of the millimeter wave radar 202.
[0246] The vehicle exterior information detector 141F includes an information processor 831. The information processor 831 and Figure 2 The information processor 211 is similar to the information processor 211, includes an image processor 221, a geometric transformation section 223 and an object recognition section 224, and differs from the information processor 211 in that the signal processor 222 has been removed.
[0247] Millimeter wave radar 811 includes a signal processor 821, which includes functions equivalent to those of signal processor 222. Signal processor 821 performs specified signal processing on millimeter wave data to generate two types of millimeter wave images: a signal intensity image and a velocity image indicating the sensing results performed by millimeter wave radar 811. Signal processor 821 provides the signal intensity image and the velocity image to geometric transformation section 223.
[0248] As described above, in the millimeter wave radar 811, millimeter wave data can be converted into a millimeter wave image.
[0249] <<7. Seventh embodiment>>
[0250] Next, refer to Figure 21 A seventh embodiment of the present technology is described.
[0251] <Configuration Example of Data Acquisition Section 102D and Vehicle Exterior Information Detector 141G>
[0252] Figure 21 Shown as Figure 1 The data acquisition portion 102D of the fourth embodiment of the data acquisition portion 102 in the vehicle control system 100 and the data acquisition portion 102D as Figure 1 An example of the configuration of the vehicle exterior information detector 141G of the seventh embodiment of the vehicle exterior information detector 141 in the vehicle control system 100. Figure 20 The part corresponding to the part in Figure 20 The same reference numerals are denoted, and descriptions thereof are appropriately omitted.
[0253] Data acquisition section 102D and Figure 20 The data acquisition section 102C is similar to that of in that it includes the camera 201 , but is different from the data acquisition section 102C in that it includes a millimeter wave radar 911 instead of the millimeter wave radar 811 .
[0254] The vehicle exterior information detector 141G includes an information processor 931. The information processor 931 is similar to the vehicle exterior information detector 141G in that it includes an image processor 221 and an object recognition section 224. Figure 20 The information processor 831 is similar to the information processor 831 and is different from the information processor 831 in that the geometric transformation part 223 is removed.
[0255] In addition to the signal processor 821 , the millimeter wave radar 911 includes a geometric transformation section 921 including a function equivalent to that of the geometric transformation section 223 .
[0256] The geometric transformation part 921 transforms the coordinate system of the signal intensity image and the velocity image from the coordinate system of the millimeter wave image to the coordinate system of the captured image, and provides the geometrically transformed signal intensity image and the geometrically transformed velocity image obtained by the geometric transformation to the object recognition part 224.
[0257] As described above, millimeter wave data can be transformed into a millimeter wave image, and geometric transformation can be performed on the millimeter wave image in the millimeter wave radar 911.
[0258] <<8. Amendments>>
[0259] Modifications to the embodiments of the present technology described above are described below.
[0260] The above description mainly describes an example in which a vehicle is the target of recognition. However, as mentioned above, any object other than a vehicle can be the target of recognition. For example, it is sufficient to perform a learning process on the object recognition model 251 using training data including correct answer data indicating the position of the object being the target of recognition.
[0261] In addition, the present technology is also applicable to the case of recognizing multiple types of objects. For example, it is sufficient to perform a learning process on the object recognition model 251 using training data including correct answer data indicating the position and label (type of object) of each object.
[0262] Furthermore, when the captured image has a size at which the object recognition section 251 can satisfactorily perform processing, the captured image may be directly input to the object recognition model 251 to perform processing for recognizing a target object.
[0263] The above has described an example of recognizing an object located in front of the vehicle 10. However, the present technology is also applicable to a case of recognizing an object located in another direction around the vehicle 10 when viewed from the vehicle 10.
[0264] Furthermore, this technology can also be applied to situations where objects around mobile objects other than vehicles are to be identified. For example, it is conceivable that this technology could be applied to mobile objects such as motorcycles, bicycles, personal mobile devices, aircraft, ships, construction machinery, and agricultural machinery (tractors). Furthermore, examples of mobile objects to which this technology can be applied include mobile objects such as drones and robots that are remotely operated by a user without the user having to board the mobile object.
[0265] Furthermore, the present technology can also be applied to a case where a process of recognizing an object at a fixed location such as a surveillance system is performed.
[0266] also, Figure 3 The object recognition model 251 is merely an example, and a model other than the object recognition model 251 generated by machine learning may also be used.
[0267] Furthermore, the present technology can also be applied to a case where a process of recognizing an object is performed by using a camera (image sensor) and LiDAR in combination.
[0268] Furthermore, the present technology is also applicable to the case where a sensor for detecting an object other than millimeter-wave radar and LiDAR is used.
[0269] Furthermore, for example, the second embodiment and the third to seventh embodiments may be combined.
[0270] Furthermore, for example, the coordinate systems of all images may be transformed so that the coordinate systems of the respective images match a new coordinate system that is different from the coordinate systems of the respective images.
[0271] <<9. Others>>
[0272] <Example of Computer Configuration>
[0273] The above series of processes can be performed using hardware or software. When a series of processes are performed using software, the program included in the software is installed on a computer. Here, examples of computers include computers incorporated with dedicated hardware, as well as computers such as general-purpose personal computers that can perform various functions through various programs installed thereon.
[0274] Figure 18 is a block diagram of an example of the hardware configuration of a computer that executes the above-described series of processes using a program.
[0275] In a computer 1000 , a central processing unit (CPU) 1001 , a read only memory (ROM) 1002 , and a random access memory (RAM) 1003 are connected to one another via a bus 1004 .
[0276] Furthermore, an input / output interface 1005 is connected to the bus 1004 . An input section 1006 , an output section 1007 , a recording section 1008 , a communication section 1009 , and a drive 1010 are connected to the input / output interface 1005 .
[0277] The input section 1006 includes, for example, an input switch, buttons, a microphone, and an imaging element. The output section 1007 includes, for example, a display and a speaker. The recording section 1008 includes, for example, a hard disk and non-volatile memory. The communication section 1009 includes, for example, a network interface. The drive 1010 drives a removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0278] In the computer 1000 having the above configuration, the above series of processing is performed by the CPU 1001 loading a program recorded in, for example, the recording section 1008 into the RAM 1003 and executing the program via the input / output interface 1005 and the bus 1004 .
[0279] For example, the program executed by the computer 1000 (CPU 1001) can be provided by being recorded in the removable medium 1011 serving as, for example, a package medium. In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0280] In the computer 1000, the program can be installed on the recording section 1008 via the input / output interface 1005 by the removable medium 1011 mounted on the drive 1010. In addition, the program can be received by the communication section 1009 via a wired or wireless transmission medium to be installed on the recording section 1008. In addition, the program can be pre-installed on the ROM 1002 or the recording section 1008.
[0281] Note that the program executed by the computer may be a program in which processing is performed chronologically in the order described here, or may be a program in which processing is performed in parallel or at necessary timing such as a calling timing.
[0282] In addition, as used herein, a system refers to a collection of multiple components such as devices and modules (components), and it does not matter whether all components are in a single housing. Therefore, multiple devices housed in separate housings and connected to each other via a network, as well as a single device in which multiple modules are housed in a single housing, are both systems.
[0283] Furthermore, the embodiments of the present technology are not limited to the above-described examples, and various modifications may be made thereto without departing from the scope of the present technology.
[0284] For example, the present technology may also have a configuration of cloud computing in which a single function is shared to be processed cooperatively by a plurality of devices via a network.
[0285] Furthermore, in addition to being executed by a single device, each step described using the above flowcharts may be executed shared by a plurality of devices.
[0286] Furthermore, when a single step includes a plurality of processes, in addition to being executed by a single device, the plurality of processes included in the single step may be shared to be executed by a plurality of devices.
[0287] <Example of configuration combination>
[0288] The present technology can also adopt the following configurations.
[0289] (1) An information processing device comprising:
[0290] a geometric transformation section that transforms at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, the captured image being obtained by an image sensor, the sensor image indicating a sensing result of the sensor, and a sensing range of the sensor at least partially overlapping with a sensing range of the image sensor; and
[0291] The object recognition section performs a process of recognizing a target object based on the captured image and the sensor image whose coordinate systems have been matched with each other.
[0292] (2) The information processing device according to (1), wherein
[0293] The geometric transformation portion transforms the sensor image into a geometrically transformed sensor image having a coordinate system identical to a coordinate system of the captured image, and
[0294] The object recognition section performs a process of recognizing the target object based on the captured image and the geometrically transformed sensor image.
[0295] (3) The information processing device according to (2), wherein
[0296] The object recognition section performs processing for recognizing the target object using an object recognition model obtained through machine learning.
[0297] (4) The information processing device according to (3), wherein
[0298] The object recognition model is caused to perform learning using training data comprising input data including the captured image and the geometrically transformed sensor image, and correct answer data indicating a location of an object in the captured image.
[0299] (5) The information processing device according to (4), wherein
[0300] The object recognition model is a model using a deep neural network.
[0301] (6) The information processing device according to (5), wherein
[0302] The object recognition model includes
[0303] a first convolutional neural network that extracts feature quantities of the captured image and the geometrically transformed sensor image, and
[0304] A second convolutional neural network recognizes the object based on the captured image and feature quantities of the geometrically transformed sensor image.
[0305] (7) The information processing device according to any one of (2) to (6), wherein
[0306] The sensor includes a millimeter wave radar, and
[0307] The sensor image indicates a location of an object that transmits a transmission signal from the millimeter wave radar.
[0308] (8) The information processing device according to (7), wherein
[0309] The coordinate system of the sensor image is represented by an axis indicating a direction in which a transmission signal is spread in a planar manner and an axis indicating an optical axis direction of the millimeter wave radar.
[0310] (9) The information processing device according to (7) or (8), wherein
[0311] The geometric transformation section transforms the first sensor image and the second sensor image into a geometrically transformed first sensor image and a geometrically transformed second sensor image, respectively, the first sensor image indicating the position of the object and the intensity of a signal reflected from the object, and the second sensor image indicating the position and speed of the object, and
[0312] The object recognition section performs a process of recognizing the target object based on the captured image, the geometrically transformed first sensor image, and the geometrically transformed second sensor image.
[0313] (10) The information processing device according to any one of (2) to (9), further comprising
[0314] an image processor that generates a low-resolution image obtained by reducing the resolution of a captured image and a cropped image obtained by cutting out a portion of the captured image, wherein
[0315] The object recognition section performs a process of recognizing the target object based on the low-resolution image, the cropped image, and the geometrically transformed sensor image.
[0316] (11) The information processing device according to (10), wherein
[0317] The object recognition part includes
[0318] a first object recognition section that performs processing for recognizing the object based on the low-resolution image and the geometrically transformed sensor image, and
[0319] a second object recognition section that performs processing for recognizing the object based on the cropped image and the geometrically transformed sensor image, and
[0320] The information processing apparatus further includes a synthesizer that synthesizes a result of recognition of the object performed by the first object recognition section and a result of recognition of the object performed by the second object recognition section.
[0321] (12) The information processing device according to (10) or (11), wherein
[0322] The image sensor and the sensor perform sensing of the surrounding environment of the vehicle, and
[0323] The image processor cuts out a cropped image based on a vanishing point of a road in the captured image.
[0324] (13) The information processing device according to (1), wherein
[0325] The geometric transformation section transforms the captured image into a geometrically transformed captured image having a coordinate system identical to that of the sensor image, and
[0326] The object recognition section performs a process of recognizing the target object based on the geometrically transformed captured image and the sensor image.
[0327] (14) The information processing device according to any one of (1) to (13), wherein
[0328] The sensor includes at least one of a millimeter wave radar or a laser radar LiDAR, and
[0329] The sensor image includes at least one of an image indicating a position of an object reflecting a transmission signal from a millimeter wave radar and point group data obtained by LiDAR.
[0330] (15) The information processing device according to any one of (1) to (14), wherein
[0331] The image sensor and the sensor perform sensing of the surrounding environment of the moving object, and
[0332] The object recognition section performs a process of recognizing the target object in the surrounding environment of the mobile body.
[0333] (16) An information processing method executed by an information processing device, the information processing method comprising:
[0334] transforming at least one of a captured image and a sensor image and matching coordinate systems of the captured image and the sensor image, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor; and
[0335] A process of recognizing an object is performed based on the captured image and the sensor image, whose coordinate systems have been matched with each other.
[0336] (17) A program for causing a computer to execute a process comprising the following steps:
[0337] transforming at least one of a captured image and a sensor image and matching coordinate systems of the captured image and the sensor image, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor; and
[0338] A process of recognizing an object is performed based on the captured image and the sensor image, whose coordinate systems have been matched with each other.
[0339] (18) A mobile body control device comprising:
[0340] a geometric transformation portion that transforms at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor;
[0341] an object recognition section that performs a process of recognizing a target object based on the captured image and the sensor image whose coordinate systems have been matched with each other; and
[0342] The motion controller controls the motion of the moving object based on the recognition result of the object.
[0343] (19) A mobile object comprising:
[0344] Image sensor;
[0345] a sensor, wherein a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor;
[0346] a geometric transformation part that transforms at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, the captured image being obtained by an image sensor, the sensor image indicating a sensing result of the sensor;
[0347] an object recognition section that performs a process of recognizing a target object based on the captured image and the sensor image whose coordinate systems have been matched with each other; and
[0348] The motion controller controls the motion of the moving object based on the recognition result of the object.
[0349] Note that the effects described here are not limiting but merely illustrative, and other effects may be provided.
[0350] Reference Signs List
[0351] 10 vehicles
[0352] 100 Vehicle Control Systems
[0353] 102, 102A to 102D data acquisition section
[0354] 107 Powertrain Controller
[0355] 108 Drivetrain System
[0356] 135 Motion Controller
[0357] 141, 141A to 141G External vehicle information detector
[0358] 201 Camera
[0359] 201A Image Sensor
[0360] 202 millimeter-wave radar
[0361] 211 Information Processor
[0362] 221 Image Processor
[0363] 222 Signal Processor
[0364] 223 Geometric Transformation Section
[0365] 224 Object Recognition
[0366] 251 Object Recognition Model
[0367] 261 Feature Extraction
[0368] 262 Identification section
[0369] 301 Learning System
[0370] 316 Training Data Generator
[0371] 317 Learning Section
[0372] 401 Message Processor
[0373] 421 Image Processor
[0374] 422 Object Recognition
[0375] 423 Synthesizer
[0376] 431a, 431b Object recognition part
[0377] 501 LiDAR
[0378] 511 Information Processor
[0379] 521 Signal Processor
[0380] 522 Geometric Transformation Section
[0381] 523 Object Recognition
[0382] 611 Information Processor
[0383] 621 Geometric Transformation Section
[0384] 622 Object Recognition Part
[0385] 711 Information Processor
[0386] 721, 722 Geometric Transformation
[0387] 723 Object Recognition
[0388] 811 millimeter-wave radar
[0389] 821 Signal Processor
[0390] 831 Information Processor
[0391] 911 millimeter-wave radar
[0392] 921 Geometric Transformation Part
Claims
1. An information processing device, comprising: a geometric transformation portion that transforms at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor; as well as an object recognition section that performs processing for recognizing an object based on the captured image and the sensor image whose coordinate systems have been matched with each other, wherein the object recognition section includes: a first convolutional neural network that extracts features of the captured image and the geometrically transformed sensor image, and a second convolutional neural network that recognizes the object based on a feature quantity of the captured image and the geometrically transformed sensor image; in: The geometric transformation section transforms the first sensor image and the second sensor image into a geometrically transformed first sensor image and a geometrically transformed second sensor image, respectively, so that the captured image matches the coordinate systems of the first sensor image and the second sensor image, the first sensor image indicating the position of the object and the intensity of the signal reflected from the object, and the second sensor image indicating the position and speed of the object, and The object recognition section performs a process of recognizing the target object based on the captured image, the geometrically transformed first sensor image, and the geometrically transformed second sensor image.
2. The information processing apparatus according to claim 1, wherein The geometric transformation portion transforms the sensor image into a geometrically transformed sensor image having a coordinate system identical to a coordinate system of the captured image, and The object recognition section performs a process of recognizing the target object based on the captured image and the geometrically transformed sensor image.
3. The information processing apparatus according to claim 1, wherein The object recognition model is caused to perform learning using training data comprising input data including the captured image and the geometrically transformed sensor image, and correct answer data indicating a location of an object in the captured image.
4. The information processing apparatus according to claim 2, wherein The sensor includes a millimeter wave radar, and The sensor image indicates a location of an object that transmits a transmission signal from the millimeter wave radar. The information processing apparatus according to claim 4 , wherein The coordinate system of the sensor image is represented by an axis indicating a direction in which a transmission signal is spread in a planar manner and an axis indicating an optical axis direction of the millimeter wave radar.
6. The information processing apparatus according to claim 2, further comprising an image processor that generates a low-resolution image obtained by reducing the resolution of a captured image and a cropped image obtained by cutting out a portion of the captured image, wherein The object recognition section performs a process of recognizing the target object based on the low-resolution image, the cropped image, and the geometrically transformed sensor image.
7. The information processing apparatus according to claim 6, wherein The object recognition part includes a first object recognition section that performs processing for recognizing the object based on the low-resolution image and the geometrically transformed sensor image, and a second object recognition section that performs processing for recognizing the object based on the cropped image and the geometrically transformed sensor image, and The information processing apparatus further includes a synthesizer that synthesizes a result of recognition of the object performed by the first object recognition section and a result of recognition of the object performed by the second object recognition section. The information processing apparatus according to claim 7 , wherein The image sensor and the sensor perform sensing of the surrounding environment of the vehicle, and The image processor cuts out a cropped image based on a vanishing point of a road in the captured image.
9. The information processing apparatus according to claim 1, wherein The geometric transformation section transforms the captured image into a geometrically transformed captured image having a coordinate system identical to that of the sensor image, and The object recognition section performs a process of recognizing the target object based on the geometrically transformed captured image and the sensor image.
10. The information processing apparatus according to claim 1, wherein The sensor includes at least one of a millimeter wave radar or a laser radar LiDAR, and The sensor image includes at least one of an image indicating a position of an object reflecting a transmission signal from a millimeter wave radar and point group data obtained by LiDAR. The information processing apparatus according to claim 1 , wherein The image sensor and the sensor perform sensing of the surrounding environment of the moving object, and The object recognition section performs a process of recognizing the target object in the surrounding environment of the mobile body.
12. An information processing method performed by an information processing device, the information processing method comprising: transforming at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor, including transforming a first sensor image and a second sensor image into a geometrically transformed first sensor image and a geometrically transformed second sensor image, respectively, so that coordinate systems of the captured image and the first and second sensor images match, the first sensor image indicating a position of an object and an intensity of a signal reflected from the object, and the second sensor image indicating a position and a speed of the object; as well as Performing a process of recognizing an object based on the captured image and the sensor image whose coordinate systems have been matched with each other, wherein performing the process of recognizing the object includes: extracting feature quantities of the captured image and the geometrically transformed sensor image using a first convolutional neural network, and The object is recognized using a second convolutional neural network based on feature quantities of the captured image and the geometrically transformed sensor image.
13. A program for causing a computer to execute a process comprising the following steps: transforming at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor, including transforming a first sensor image and a second sensor image into a geometrically transformed first sensor image and a geometrically transformed second sensor image, respectively, so that coordinate systems of the captured image and the first and second sensor images match, the first sensor image indicating a position of an object and an intensity of a signal reflected from the object, and the second sensor image indicating a position and a speed of the object; and Performing a process of recognizing an object based on the captured image and the sensor image whose coordinate systems have been matched with each other, wherein performing the process of recognizing the object includes: extracting feature quantities of the captured image and the geometrically transformed sensor image using a first convolutional neural network, and The object is identified based on feature quantities of the captured image and the geometrically transformed sensor image using a second convolutional neural network, including performing processing to identify the object based on the captured image, the geometrically transformed first sensor image, and the geometrically transformed second sensor image.
14. A mobile object control device, comprising: a geometric transformation portion that transforms at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, wherein the captured image is obtained by an image sensor, the sensor image indicates a sensing result of the sensor, and a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor; an object recognition section that performs processing for recognizing an object based on the captured image and the sensor image whose coordinate systems have been matched with each other, wherein the object recognition section includes: a first convolutional neural network that extracts features of the captured image and the geometrically transformed sensor image, and a second convolutional neural network that recognizes the object based on a feature quantity of the captured image and the geometrically transformed sensor image; as well as a motion controller for controlling the motion of the mobile body based on the recognition result of the object; in: The geometric transformation section transforms the first sensor image and the second sensor image into a geometrically transformed first sensor image and a geometrically transformed second sensor image, respectively, so that the coordinate systems of the captured image and the first sensor image and the second sensor image match, the first sensor image indicating the position of the object and the intensity of the signal reflected from the object, and the second sensor image indicating the position and speed of the object, and The object recognition section performs a process of recognizing the target object based on the captured image, the geometrically transformed first sensor image, and the geometrically transformed second sensor image.
15. A mobile object, comprising: Image sensor; a sensor, wherein a sensing range of the sensor at least partially overlaps with a sensing range of the image sensor; a geometric transformation part that transforms at least one of a captured image and a sensor image so that coordinate systems of the captured image and the sensor image match, the captured image being obtained by an image sensor, the sensor image indicating a sensing result of the sensor; an object recognition section that performs processing for recognizing an object based on the captured image and the sensor image whose coordinate systems have been matched with each other, wherein the object recognition section includes: a first convolutional neural network that extracts features of the captured image and the geometrically transformed sensor image, and a second convolutional neural network that recognizes the object based on a feature quantity of the captured image and the geometrically transformed sensor image; as well as an action controller, controlling an action based on a recognition result of the object; in: The geometric transformation section transforms the first sensor image and the second sensor image into a geometrically transformed first sensor image and a geometrically transformed second sensor image, respectively, so that the coordinate systems of the captured image and the first sensor image and the second sensor image match, the first sensor image indicating the position of the object and the intensity of the signal reflected from the object, and the second sensor image indicating the position and speed of the object, and The object recognition section performs a process of recognizing the target object based on the captured image, the geometrically transformed first sensor image, and the geometrically transformed second sensor image.
Citation Information
Patent Citations
Method and system for displaying obstacle using radar
JP2005175603A
Database construction system for article recognition algorism machine-learning
JP2017102838A
Apparatus and method for detecting an object from input image data in vehicle
KR101920281B1
Image recognition imaging apparatus
WO2018101247A1