Information processing device, information processing method, program, mobile object control device, and mobile object
By generating estimated position images and combining image processing and object recognition technology, the problem of insufficient accuracy in identifying target objects in the prior art is solved, and target object recognition and moving object control are achieved with higher accuracy.
Patent Information
- Application Number
- CN201980079022.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-07
- Filing Date
- 2019-11-22
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2039-11-22
AI Technical Summary
The prior art has failed to effectively improve the accuracy of using cameras and millimeter wave radar to identify target objects (such as vehicles).
By generating an estimated position image, the sensor sensing result in which the sensor sensing range at least partially overlaps the image sensor sensing range by using the first coordinate system, and in combination with image processing and object recognition technology, the target object is identified and the movement of the moving object is controlled.
Improve the accuracy of target object recognition and enhance the accuracy of moving object control.
Smart Images

Figure CN113168692B_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to an information processing device, an information processing method, a program, a mobile object control device, and a mobile object, and in particular to an information processing device, an information processing method, a program, a mobile object control device, and a mobile object designed to improve the accuracy of identifying a target object. Background Art
[0002] It has been proposed in the past that position information on an obstacle detected by a millimeter wave radar is superimposed to be displayed on a camera image using projective transformation performed with respect to a radar plane and a camera image plane (for example, refer to Patent Document 1).
[0003] Citation List
[0004] Patent Literature
[0005] Patent Document 1: Japanese Patent Application Laid-Open No. 2005-175603 Summary of the Invention
[0006] Technical issues
[0007] However, Patent Document 1 does not discuss improving the accuracy of recognizing a target object (such as a vehicle) using a camera and a millimeter wave radar.
[0008] The present technology has been made in view of the above-described circumstances, and aims to improve the accuracy of recognizing a target object.
[0009] Solution to the problem
[0010] An information processing device according to a first aspect of the present technology includes: an image processor that generates an estimated position image based on a sensor image, the sensor image indicating a sensing result of a sensor whose sensing range at least partially overlaps with the sensing range of the image sensor using a first coordinate system, the estimated position image indicating an estimated position of a target object in a second coordinate system that is the same as the coordinate system of a captured image obtained by the image sensor; and an object recognition part that performs processing for recognizing the target object based on the captured image and the estimated position image.
[0011] According to the first aspect of the present technology, the information processing method includes: generating an estimated position image by an information processing device based on a sensor image, wherein the sensor image indicates a sensing result of a sensor whose sensing range at least partially overlaps with the sensing range of an image sensor using a first coordinate system, and the estimated position image indicates an estimated position of the target object in a second coordinate system that is the same as the coordinate system of the captured image obtained by the image sensor; and performing processing for identifying the target object by the information processing device based on the captured image and the estimated position image.
[0012] According to a first aspect of the present technology, a program causes a computer to perform processing, including: generating an estimated position image based on a sensor image, wherein the sensor image indicates a sensing result of a sensor whose sensing range at least partially overlaps with the sensing range of an image sensor using a first coordinate system, and the estimated position image indicates an estimated position of a target object in a second coordinate system that is the same as the coordinate system of a captured image obtained by the image sensor; and performing processing for identifying the target object based on the captured image and the estimated position image.
[0013] According to the second aspect of the present technology, a mobile object control device includes: an image processor, which generates an estimated position image based on a sensor image, wherein the sensor image indicates a sensing result of a sensor whose sensing range at least partially overlaps with the sensing range of an image sensor that captures an image around the mobile object using a first coordinate system, and the estimated position image indicates an estimated position of the target object in a second coordinate system that is the same as the coordinate system of the captured image obtained by the image sensor; an object recognition part, which performs processing to recognize the target object based on the captured image and the estimated position image; and a motion controller, which controls the movement of the mobile object based on the recognition result of the target object.
[0014] According to the third aspect of the present technology, a mobile object control device includes: an image sensor; a sensor, the sensing range of which at least partially overlaps with the sensing range of the image sensor; an image processor, the image processor generating an estimated position image based on the sensor image, the sensor image indicating the sensing result of the sensor using a first coordinate system, the estimated position image indicating the estimated position of the target object in a second coordinate system that is the same as the coordinate system of the captured image obtained by the image sensor; an object recognition part, the object recognition part performing processing to identify the target object based on the captured image and the estimated position image; and a motion controller, the motion controller controlling the movement of the mobile object based on the recognition result of the target object.
[0015] In a first aspect of the present technology, an estimated position image is generated based on a sensor image, the estimated position image indicating an estimated position of a target object in a second coordinate system, the second coordinate system being the same as a coordinate system of a captured image obtained by an image sensor whose sensing range at least partially overlaps with the sensing range of the sensor, the sensor image indicating a sensing result of the sensor in a first coordinate system; and processing for identifying the target object is performed based on the captured image and the estimated position image.
[0016] In a second aspect of the present technology, an estimated position image is generated based on a sensor image, the estimated position image indicating an estimated position of a target object in a second coordinate system, the second coordinate system being the same as a coordinate system of a captured image obtained by an image sensor that captures an image around a moving object and whose sensing range at least partially overlaps with the sensing range of the sensor, the sensor image indicating a sensing result of the sensor in a first coordinate system; processing for identifying the target object is performed based on the captured image and the estimated position image; and the movement of the moving object is controlled based on the recognition result of the target object.
[0017] In a third aspect of the present technology, an estimated position image is generated based on a sensor image, the estimated position image indicating an estimated position of a target object in a second coordinate system, the second coordinate system being the same as a coordinate system of a captured image obtained by an image sensor whose sensing range at least partially overlaps with the sensing range of the sensor, the sensor image indicating a sensing result of the sensor in a first coordinate system; processing for identifying the target object is performed based on the captured image and the estimated position image; and movement is controlled based on the recognition result of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a block diagram illustrating a configuration example of a vehicle control system to which the present technology is applied.
[0019] Figure 2 is a block diagram illustrating a first embodiment of a data acquisition section and a first embodiment of a vehicle exterior information detector.
[0020] Figure 3 Illustrate a configuration example of an image processing model.
[0021] Figure 4 Figure 1 shows an example configuration of an object recognition model.
[0022] Figure 5 Illustrated is a configuration example of a learning system for an image processing model.
[0023] Figure 6 An example configuration of a learning system for an object recognition model is shown.
[0024] Figure 7 is a flowchart for describing the learning process performed on the image processing model.
[0025] Figure 8 is a diagram for describing the learning process performed on the image processing model.
[0026] Figure 9 is a flowchart for describing the learning process performed on the object recognition model.
[0027] Figure 10is a diagram for describing the learning process performed on the object recognition model.
[0028] Figure 11 is a flowchart for describing the target object recognition process.
[0029] Figure 12 is a diagram for describing the effects provided by the present technology.
[0030] Figure 13 is a diagram for describing the effects provided by the present technology.
[0031] Figure 14 is a diagram for describing the effects provided by the present technology.
[0032] Figure 15 is a block diagram illustrating a second embodiment of a data acquisition portion and a second embodiment of a vehicle exterior information detector.
[0033] Figure 16 is a diagram for describing processing performed when the millimeter wave radar has resolution in the height direction.
[0034] Figure 17 Illustrated is a modified example of a millimeter wave image.
[0035] Figure 18 The figure shows an example of a computer configuration. DETAILED DESCRIPTION
[0036] The following describes an embodiment for implementing the present technology. The description is given in the following order.
[0037] 1. First Embodiment (Example Using Camera and Millimeter Wave Radar)
[0038] 2. Second embodiment (example with LiDAR added)
[0039] 3. Modification
[0040] 4. Others
[0041] <<1. First embodiment>>
[0042] First, refer to Figures 1 to 14 A first embodiment of the present technology is described.
[0043] <Configuration Example of Vehicle Control System 100>
[0044] Figure 1 is a block diagram illustrating a schematic functional configuration example of a vehicle control system 100 , which is an example of a mobile object control system to which the present technology can be applied.
[0045] Note that when the vehicle 10 provided with the vehicle control system 100 is to be distinguished from other vehicles, the vehicle provided with the vehicle control system 100 will be referred to as a host vehicle or a host vehicle hereinafter.
[0046] The vehicle control system 100 includes an input unit 101, a data acquisition unit 102, a communication unit 103, an onboard device 104, an output controller 105, an output unit 106, a powertrain controller 107, a powertrain system 108, a vehicle body-related controller 109, a vehicle body-related system 110, a storage device 111, and an autonomous driving controller 112. The input unit 101, the data acquisition unit 102, the communication unit 103, the output controller 105, the powertrain controller 107, the vehicle body-related controller 109, the storage device 111, and the autonomous driving controller 112 are connected to one another via a communication network 121. For example, the communication network 121 includes a bus or an onboard communication network conforming to any standard, such as a controller area network (CAN), a local interconnect network (LIN), a local area network (LAN), or FlexRay (registered trademark). Note that the respective structural elements of the vehicle control system 100 can be directly connected to one another without using the communication network 121.
[0047] Note that when the respective structural elements of the vehicle control system 100 communicate with each other via the communication network 121, the description of the communication network 121 will be omitted below. For example, when the input portion 101 and the automatic driving controller 112 communicate with each other via the communication network 121, it will be simply stated that the input portion 101 and the automatic driving controller 112 communicate with each other.
[0048] The input section 101 includes devices used by onboard personnel to input various data, instructions, and the like. For example, the input section 101 includes operating devices such as a touch panel, buttons, microphones, switches, and joysticks; operating devices that can perform input by methods other than manual operation (such as voice or gestures); and the like. Alternatively, for example, the input section 101 may be externally connected equipment, such as a remote control device using infrared or another type of radio wave, or mobile equipment or wearable equipment compatible with the operation of the vehicle control system 100. The input section 101 generates input signals based on the data, instructions, and the like input by the onboard personnel, and supplies the generated input signals to the corresponding structural elements of the vehicle control system 100.
[0049] The data acquisition portion 102 includes various sensors and the like to acquire data used for processing performed by the vehicle control system 100 and supplies the acquired data to respective structural elements of the vehicle control system 100 .
[0050] For example, the data acquisition section 102 includes various sensors for detecting, for example, the status of the vehicle. Specifically, for example, the data acquisition section 102 includes a gyroscope; an acceleration sensor; an inertial measurement unit (IMU); and sensors for detecting the amount of accelerator pedal operation, the amount of brake pedal operation, the steering angle of the steering wheel, the number of engine revolutions, the number of motor revolutions, the rotational speed of the wheels, and the like.
[0051] In addition, for example, the data acquisition section 102 includes various sensors for detecting information about the exterior of the vehicle. Specifically, for example, the data acquisition section 102 includes an image capture device such as a time-of-flight (ToF) camera, a stereo camera, a monocular camera, an infrared camera, and other cameras. In addition, for example, the data acquisition section 102 includes environmental sensors for detecting weather, meteorological phenomena, and the like, and surrounding information detection sensors for detecting objects around the vehicle. For example, environmental sensors include raindrop sensors, fog sensors, sunlight sensors, snow sensors, and the like. Surrounding information detection sensors include ultrasonic sensors, radars, LiDAR (light detection and ranging, laser imaging detection and ranging), sonars, and the like.
[0052] Furthermore, for example, the data acquisition section 102 includes various sensors for detecting the current position of the vehicle. Specifically, for example, the data acquisition section 102 includes, for example, a Global Navigation Satellite System (GNSS) receiver that receives GNSS signals from GNSS satellites.
[0053] Furthermore, for example, data acquisition section 102 includes various sensors for detecting information about the interior of the vehicle. Specifically, for example, data acquisition section 102 includes an image capture device for capturing an image of the driver, a biometric sensor for detecting the driver's biological information, and a microphone for collecting sounds from the interior of the vehicle. For example, the biometric sensor is provided on a seat surface, a steering wheel, or the like, and detects biological information of a vehicle occupant seated in the seat or the driver holding the steering wheel.
[0054] The communication section 103 communicates with the in-vehicle equipment 104, various external equipment, servers, base stations, and the like, transmitting data supplied by the corresponding components of the vehicle control system 100 and supplying received data to the corresponding components of the vehicle control system 100. It should be noted that the communication protocols supported by the communication section 103 are not particularly limited. The communication section 103 may also support multiple types of communication protocols.
[0055] For example, the communication section 103 wirelessly communicates with the in-vehicle equipment 104 using wireless LAN, Bluetooth (registered trademark), near field communication (NFC), wireless USB (WUSB), etc. In addition, for example, the communication section 103 communicates with the in-vehicle equipment 104 using a universal serial bus (USB), a high-definition multimedia interface (HDMI) (registered trademark), a mobile high-definition link (MHL), etc. using electric wires through a connection terminal (not shown) (or a cable if necessary).
[0056] In addition, for example, the communication section 103 communicates with equipment (e.g., an application server or a control server) located on an external network (e.g., the Internet, a cloud network, or a network specific to an operator) through a base station or an access point. In addition, for example, the communication section 103 communicates with a terminal (e.g., a terminal of a pedestrian or a store, or a machine type communication (MTC) terminal) located near the vehicle using a peer-to-peer (P2P) technology. Moreover, for example, the communication section 103 performs V2X communication such as vehicle-to-vehicle communication, vehicle-to-infrastructure communication, vehicle-to-home communication between the vehicle and a home, and vehicle-to-pedestrian communication. In addition, for example, the communication section 103 includes a beacon receiver that receives radio waves or electromagnetic waves transmitted from, for example, a radio station installed on a road, and obtains information about, for example, the current position, traffic congestion, traffic control, or necessary time.
[0057] Examples of the in-vehicle equipment 104 include mobile equipment or wearable equipment of people on board the vehicle, information equipment brought into or attached to the vehicle, and a navigation device that searches for a route to any destination.
[0058] The output controller 105 controls the output of various information to occupants of the vehicle or to the exterior of the vehicle. For example, the output controller 105 generates an output signal including at least one of visual information (such as image data) or audio information (such as sound data), supplies the output signal to the output portion 106, and thereby controls the output of the visual and audio information from the output portion 106. Specifically, for example, the output controller 105 combines multiple pieces of image data captured by different image capture devices of the data acquisition portion 102 to generate a bird's-eye view image, a panoramic image, or the like, and supplies the output signal including the generated image to the output portion 106. Furthermore, for example, the output controller 105 generates sound data including, for example, a warning buzzer or a warning message warning of dangers such as collision, contact, or entering a danger zone, and supplies the output signal including the generated sound data to the output portion 106.
[0059] The output portion 106 includes a device capable of outputting visual or audio information to occupants of the vehicle or to the exterior of the vehicle. For example, the output portion 106 includes a display device, an instrument panel, audio speakers, headphones, a wearable device (such as a glasses-type display for occupants), a projector, a light, and the like. Instead of a device including a conventional display, the display device included in the output portion 106 may be a device that displays visual information in the driver's field of view, such as a head-up display, a transparent display, or a device including augmented reality (AR) display functionality.
[0060] The powertrain controller 107 generates various control signals and supplies them to the powertrain system 108, thereby controlling the powertrain system 108. In addition, the powertrain controller 107 also supplies control signals to structural elements other than the powertrain system 108 as needed, for example, to notify them of the conditions of controlling the powertrain system 108.
[0061] The powertrain system 108 includes various devices related to the powertrain of the vehicle. For example, the powertrain system 108 includes a driving force generating device (such as an internal combustion engine and a drive motor) that generates driving force, a driving force transmission mechanism for transmitting driving force to the wheels, a steering mechanism for adjusting the steering angle, a braking device that generates braking force, an anti-lock braking system (ABS), an electronic stability control (ESC) system, an electric power steering device, and the like.
[0062] The vehicle body controller 109 generates various control signals and supplies them to the vehicle body system 110, thereby controlling the vehicle body system 110. Furthermore, the vehicle body controller 109 supplies control signals to components other than the vehicle body system 110 as necessary, for example, to notify them of the status of controlling the vehicle body system 110.
[0063] The vehicle body-related system 110 includes various vehicle body-related devices provided to the vehicle body. For example, the vehicle body-related system 110 includes a keyless entry system, a smart key system, power windows, power seats, a steering wheel, an air conditioner, and various lights (such as headlights, taillights, brake lights, turn signals, and fog lights).
[0064] For example, the storage device 111 includes a read-only memory (ROM), a random access memory (RAM), a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, a magneto-optical storage device, etc. The storage device 111 stores therein various programs, data, etc. used by the corresponding structural elements of the vehicle control system 100. For example, the storage device 111 stores therein map data such as a three-dimensional high-precision map, a global map, and a local map. The high-precision map is a dynamic map, etc. The accuracy of the global map is lower than that of the high-precision map, but the coverage is wider than that of the high-precision map. The local map includes information about the surroundings of the vehicle.
[0065] The autonomous driving controller 112 performs control related to autonomous driving, such as autonomous driving or driving assistance. Specifically, for example, the autonomous driving controller 112 performs cooperative control to implement the functions of an advanced driver assistance system (ADAS), including collision avoidance or shock absorption for the host vehicle, driving behind a preceding vehicle based on inter-vehicle distance, driving while maintaining vehicle speed, warnings for host vehicle collisions, warnings for host vehicle lane departure, and the like. Furthermore, for example, the autonomous driving controller 112 performs cooperative control to implement autonomous driving (which is autonomous driving without requiring driver input). The autonomous driving controller 112 includes a detector 131, a self-position estimator 132, a situation analyzer 133, a planning unit 134, and a motion controller 135.
[0066] The detector 131 detects various information required for controlling the automatic driving, and includes a vehicle exterior information detector 141 , a vehicle interior information detector 142 , and a vehicle condition detector 143 .
[0067] The vehicle exterior information detector 141 performs processing for detecting information about the exterior of the vehicle based on data or signals from each structural element of the vehicle control system 100. For example, the vehicle exterior information detector 141 performs processing for detecting, identifying, and tracking objects around the vehicle, as well as processing for detecting the distance to the objects. Examples of detection target objects include vehicles, people, obstacles, structures, roads, traffic lights, traffic signs, and road signs. Furthermore, for example, the vehicle exterior information detector 141 performs processing for detecting the environment surrounding the vehicle. Examples of detection target surrounding environments include weather, temperature, humidity, brightness, and road surface conditions. The vehicle exterior information detector 141 supplies data indicating the results of the detection processing to, for example, the own position estimator 132; the map analyzer 151, traffic rule recognition section 152, and situation recognition section 153 of the situation analyzer 133; and the emergency avoidance section 171 of the motion controller 135.
[0068] The vehicle interior information detector 142 performs processing for detecting information about the vehicle interior based on data or signals from each structural element of the vehicle control system 100. For example, the vehicle interior information detector 142 performs processing for authenticating and identifying the driver, processing for detecting the driver's condition, processing for detecting occupants, and processing for detecting the vehicle interior environment. Examples of the driver's detection target conditions include physical state, arousal level, concentration level, fatigue level, and line of sight. Examples of detection target vehicle interior environments include temperature, humidity, brightness, and odor. The vehicle interior information detector 142 supplies data indicating the results of the detection processing to, for example, the situation identification portion 153 of the situation analyzer 133 and the emergency avoidance portion 171 of the motion controller 135.
[0069] The vehicle condition detector 143 performs processing to detect the condition of the vehicle based on data or signals from each structural element of the vehicle control system 100. Examples of detection target conditions of the vehicle include speed, acceleration, steering angle, the presence or absence of an abnormality and its details, driving operation conditions, the position and inclination of the power seat, the condition of the door locks, and the conditions of other in-vehicle equipment. The vehicle condition detector 143 supplies data indicating the results of the detection processing to, for example, the condition identification section 153 of the condition analyzer 133 and the emergency avoidance section 171 of the motion controller 135.
[0070] The own-position estimator 132 performs processing to estimate the position, posture, and other aspects of the vehicle based on data or signals from corresponding structural elements of the vehicle control system 100 (such as the vehicle exterior information detector 141 and the situation identification section 153 of the situation analyzer 133). Furthermore, the own-position estimator 132 generates a local map (hereinafter referred to as the own-position estimation map) for estimating the own-position, as needed. For example, the own-position estimation map is a high-precision map using techniques such as simultaneous localization and mapping (SLAM). The own-position estimator 132 supplies data indicating the results of the estimation processing to, for example, the map analyzer 151, traffic regulation identification section 152, and situation identification section 153 of the situation analyzer 133. Furthermore, the own-position estimator 132 stores the own-position estimation map in the storage device 111.
[0071] The situation analyzer 133 performs processing for analyzing the situation of the vehicle and its surroundings. The situation analyzer 133 includes a map analyzer 151 , a traffic regulation recognition section 152 , a situation recognition section 153 , and a situation prediction section 154 .
[0072] As needed, the map analyzer 151 uses data or signals from corresponding structural elements of the vehicle control system 100 (such as the own position estimator 132 and the vehicle external information detector 141) to analyze various maps stored in the storage device 111 and construct a map including information required for autonomous driving processing. The map analyzer 151 supplies the constructed map to, for example, the traffic regulation recognition section 152, the situation recognition section 153, and the situation prediction section 154, as well as the route planning section 161, the behavior planning section 162, and the motion planning section 163 of the planning section 134.
[0073] The traffic regulation recognition section 152 performs processing to identify traffic regulations surrounding the vehicle based on data or signals from the corresponding structural elements of the vehicle control system 100 (such as the own position estimator 132, the vehicle external information detector 141, and the map analyzer 151). This recognition process enables the identification of the position and status of traffic lights surrounding the vehicle, the details of traffic control implemented around the vehicle, and drivable lanes. The traffic regulation recognition section 152 supplies data indicating the results of the recognition process to, for example, the situation prediction section 154.
[0074] The situation recognition section 153 performs processing to identify the situation related to the vehicle based on data or signals from the corresponding structural elements of the vehicle control system 100 (such as the vehicle position estimator 132, the vehicle exterior information detector 141, the vehicle interior information detector 142, the vehicle condition detector 143, and the map analyzer 151). For example, the situation recognition section 153 performs processing to identify the situation of the vehicle, the situation around the vehicle, the situation of the driver of the vehicle, and so on. Furthermore, as needed, the situation recognition section 153 generates a local map (hereinafter referred to as a situation recognition map) for identifying the situation around the vehicle. The situation recognition map is, for example, an occupancy grid map.
[0075] Examples of the vehicle's target conditions include its position, posture, and motion (such as speed, acceleration, and direction of movement), as well as the presence and details of any abnormalities. Examples of the vehicle's surrounding conditions include the type and position of stationary objects; the type, position, and motion (such as speed, acceleration, and direction of movement) of moving objects; the structure and surface condition of the road surrounding the vehicle; and the weather, temperature, humidity, and brightness surrounding the vehicle. Examples of the driver's target conditions include their physical state, arousal level, concentration level, fatigue level, eye movement, and driving operation.
[0076] The situation recognition section 153 supplies data indicating the result of the recognition processing (including the situation recognition map as needed) to, for example, the own position estimator 132 and the situation prediction section 154. In addition, the situation recognition section 153 stores the situation recognition map in the storage device 111.
[0077] The situation prediction section 154 performs processing for predicting a situation related to the vehicle based on data or signals from corresponding structural elements of the vehicle control system 100, such as the map analyzer 151, the traffic regulation recognition section 152, and the situation recognition section 153. For example, the situation prediction section 154 performs processing for predicting the situation of the vehicle, the situation around the vehicle, the situation of the driver, and the like.
[0078] Examples of predicted target conditions for the vehicle include the vehicle's behavior, the occurrence of an abnormality in the vehicle, and the vehicle's drivable distance. Examples of predicted target conditions around the vehicle include the behavior of moving objects, changes in traffic light conditions, and changes in the environment around the vehicle (such as weather). Examples of predicted target conditions for the driver include the driver's behavior and physical condition.
[0079] The situation prediction section 154 supplies data indicating the result of the prediction process together with data from the traffic regulation recognition section 152 and the situation recognition section 153 to, for example, the route planning section 161 , the behavior planning section 162 and the movement planning section 163 of the planning section 134 .
[0080] The route planning section 161 plans a route to a destination based on data or signals from corresponding structural elements of the vehicle control system 100 (such as the map analyzer 151 and the situation prediction section 154). For example, the route planning section 161 sets a route from the current location to the designated destination based on a global map. Furthermore, the route planning section 161 appropriately changes the route based on conditions such as traffic congestion, accidents, traffic regulations, and construction, as well as the driver's physical condition. The route planning section 161 supplies data indicating the planned route to, for example, the behavior planning section 162.
[0081] Based on data or signals from corresponding structural elements of the vehicle control system 100 (such as the map analyzer 151 and the situation prediction section 154), the behavior planning section 162 plans the behavior of the host vehicle so that the host vehicle can safely travel on the route planned by the route planning section 161 within the time planned by the route planning section 161. For example, the behavior planning section 162 formulates plans regarding, for example, starting movement, stopping, travel direction (such as forward movement, backward movement, left turn, right turn, and direction change), travel lane, travel speed, and overtaking. The behavior planning section 162 supplies data indicating the planned behavior of the host vehicle to, for example, the movement planning section 163.
[0082] Based on data or signals from corresponding structural elements of the vehicle control system 100 (such as the map analyzer 151 and the situation prediction section 154), the motion planning section 163 plans the motion of the host vehicle so as to achieve the behavior planned by the behavior planning section 162. For example, the motion planning section 163 formulates plans regarding, for example, acceleration, deceleration, and the driving route. The motion planning section 163 supplies data indicating the planned motion of the host vehicle to, for example, the acceleration / deceleration controller 172 and the direction controller 173 of the motion controller 135.
[0083] The motion controller 135 controls the motion of the vehicle. The motion controller 135 includes an emergency avoidance section 171, an acceleration / deceleration controller 172, and a direction controller 173.
[0084] Based on the results of the detection performed by the vehicle exterior information detector 141, the vehicle interior information detector 142, and the vehicle condition detector 143, the emergency avoidance portion 171 performs processing for detecting emergency events (such as collision, contact, entry into a dangerous area, an unusual condition of the driver, and an abnormality in the vehicle). When the emergency avoidance portion 171 detects the occurrence of an emergency, the emergency avoidance portion 171 plans the movement of the vehicle, such as a sudden stop or a quick turn, to avoid the emergency. The emergency avoidance portion 171 supplies data indicating the planned movement of the vehicle to, for example, the acceleration / deceleration controller 172 and the direction controller 173.
[0085] The acceleration / deceleration controller 172 controls acceleration / deceleration to realize the motion of the host vehicle planned by the motion planning section 163 or the emergency avoidance section 171. For example, the acceleration / deceleration controller 172 calculates a control target value of a driving force generating device or a braking device for realizing the planned acceleration, planned deceleration, or planned sudden stop, and supplies a control instruction indicating the calculated control target value to the powertrain controller 107.
[0086] The direction controller 173 controls the direction to realize the movement of the host vehicle planned by the movement planning section 163 or the emergency avoidance section 171. For example, the direction controller 173 calculates a control target value of a steering mechanism for realizing the travel route planned by the movement planning section 163 or the quick turn planned by the emergency avoidance section 171, and supplies a control instruction indicating the calculated control target value to the powertrain controller 107.
[0087] <Configuration Example of Data Acquisition Section 102A and Vehicle Exterior Information Detector 141A>
[0088] Figure 2 As shown in the figure Figure 1The data acquisition section 102A of the first embodiment of the data acquisition section 102 in the vehicle control system 100 and the data acquisition section 102A of the first embodiment of the vehicle control system 100 Figure 1 FIG. 1 is a portion of a configuration example of a vehicle exterior information detector 141A of the first embodiment of the vehicle exterior information detector 141 in the vehicle control system 100 .
[0089] The data acquisition section 102A includes a camera 201 and a millimeter wave radar 202. The vehicle exterior information detector 141A includes an information processor 211. The information processor 211 includes an image processor 221, a signal processor 222, an image processor 223, and an object recognition section 224.
[0090] The camera 201 includes an image sensor 201A. Any type of image sensor, such as a CMOS image sensor or a CCD image sensor, can be used as the image sensor 201A. The camera 201 (image sensor 201A) captures an image of the area in front of the vehicle 10 and supplies the obtained image (hereinafter referred to as a captured image) to the image processor 221.
[0091] The millimeter-wave radar 202 performs sensing relative to an area located in front of the vehicle 10, and the sensing ranges of the millimeter-wave radar 202 and the camera 201 at least partially overlap. For example, the millimeter-wave radar 202 transmits a transmission signal including millimeter waves in the forward direction of the vehicle 10, and receives a reception signal that is a signal reflected from an object (reflector) located in front of the vehicle 10 using a receiving antenna. For example, a plurality of receiving antennas are arranged at specified intervals in the lateral direction (width direction) of the vehicle 10. In addition, a plurality of receiving antennas can also be arranged in the height direction. The millimeter-wave radar 202 supplies data (hereinafter referred to as millimeter-wave data) indicating the strength of the reception signal received using each receiving antenna in chronological order to the signal processor 222.
[0092] The image processor 221 performs specified image processing on the captured image. For example, the image processor 221 performs processing to interpolate the red (R), green (G), and blue (B) components for each pixel of the captured image to generate an R image composed of the R component of the captured image, a G image composed of the G component of the captured image, and a B image composed of the B component of the captured image. The image processor 221 supplies the R image, G image, and B image to the object recognition unit 224.
[0093] The signal processor 222 performs specified signal processing on the millimeter wave data to generate a millimeter wave image, which is an image indicating a result of sensing performed by the millimeter wave radar 202. The signal processor 222 supplies the millimeter wave image to the image processor 223.
[0094] The image processor 223 performs specified image processing on the millimeter wave image to generate an estimated position image indicating the estimated position of the target object whose coordinate system is exactly the same as the coordinate system of the captured image. The image processor 223 supplies the estimated position image to the object recognition section 224.
[0095] The object recognition section 224 performs processing for recognizing a target object located in front of the vehicle 10 based on the R image, the G image, the B image, and the estimated position image. The object recognition section 224 supplies data indicating the result of recognizing the target object to, for example, the own position estimator 132; the map analyzer 151, the traffic regulation recognition section 152, and the situation recognition section 153 of the situation analyzer 133; and the emergency avoidance section 171 of the motion controller 135.
[0096] Note that the target object is an object to be identified by object identification section 224, and any object can be set as the target object. However, it is advantageous to set as the target object an object that includes a portion with a high reflectivity for the signal transmitted by millimeter-wave radar 202. The following description will be made using the case where the target object is a vehicle as an example.
[0097] <Configuration Example of Image Processing Model 301>
[0098] Figure 3 A configuration example of an image processing model 301 used for the image processor 223 is illustrated.
[0099] Image processing model 301 is a model obtained through machine learning. Specifically, image processing model 301 is a model obtained through deep learning, which is a type of machine learning and uses a deep neural network. Image processing model 301 includes a feature extraction section 311, a geometric transformation section 312, and a deconvolution section 313.
[0100] The feature extraction section 311 includes a convolutional neural network. Specifically, the feature extraction section 311 includes convolutional layers 321a to 321c. The convolutional layers 321a to 321c perform convolution operations to extract the feature quantities of the millimeter wave image, generate a feature map indicating the distribution of the feature quantities in a coordinate system identical to that of the millimeter wave image, and supply the feature map to the geometric transformation section 312.
[0101] The geometric transformation section 312 includes geometric transformation layers 322a and 322b. The geometric transformation layers 322a and 322b perform geometric transformation on the feature map to transform the coordinate system of the feature map from the coordinate system of the millimeter wave image to the coordinate system of the captured image. The geometric transformation section 312 supplies the feature map on which the geometric transformation has been performed to the deconvolution section 313.
[0102] The deconvolution section 313 includes deconvolution layers 323a to 323c, which deconvolution layers 323a to 323c deconvolve the feature map on which the geometric transformation has been performed to generate and output an estimated position image.
[0103] <Configuration Example of Object Recognition Model 351>
[0104] Figure 4 An example of the configuration of the object recognition model 351 used in the object recognition section 224 is illustrated.
[0105] Object recognition model 351 is a model obtained through machine learning. Specifically, object recognition model 351 is a model obtained through deep learning, a type of machine learning that uses a deep neural network. More specifically, object recognition model 351 is composed of a single-shot multi-box detector (SSD), which is one type of object recognition model that uses a deep neural network. Object recognition model 351 includes a feature extraction unit 361 and a recognition unit 362.
[0106] The feature extraction section 361 includes a VGG16 371, which is a convolutional layer using a convolutional neural network. Four-channel image data P, including an R image, a G image, a B image, and an estimated position image, is input to the VGG16 371. The VGG16 371 extracts feature quantities from each of the R image, the G image, the B image, and the estimated position image, and generates a combined feature map that two-dimensionally represents the distribution of feature quantities obtained by combining the feature quantities extracted from the respective images. The combined feature map represents the distribution of feature quantities in a coordinate system identical to the coordinate system of the captured image. The VGG16 371 supplies the combined feature map to the recognition section 362.
[0107] The recognition part 362 includes a convolutional neural network. Specifically, the recognition part 362 includes convolutional layers 372a to 372f.
[0108] The convolution layer 372a performs a convolution operation on the combined feature map. The convolution layer 372a performs a process of recognizing the target object based on the combined feature map on which the convolution operation has been performed. The convolution layer 372a supplies the combined feature map on which the convolution operation has been performed to the convolution layer 372b.
[0109] Convolution layer 372b performs a convolution operation on the combined feature map supplied by convolution layer 372a. Convolution layer 372b performs a process of recognizing a target object based on the combined feature map on which the convolution operation has been performed. Convolution layer 372b supplies the combined feature map on which the convolution operation has been performed to convolution layer 372c.
[0110] The convolution layer 372c performs a convolution operation on the combined feature map supplied by the convolution layer 372b. The convolution layer 372c performs a process of recognizing the target object based on the combined feature map on which the convolution operation has been performed. The convolution layer 372c supplies the combined feature map on which the convolution operation has been performed to the convolution layer 372d.
[0111] The convolution layer 372d performs a convolution operation on the combined feature map supplied by the convolution layer 372c. The convolution layer 372d performs a process of recognizing the target object based on the combined feature map on which the convolution operation has been performed. The convolution layer 372d supplies the combined feature map on which the convolution operation has been performed to the convolution layer 372e.
[0112] The convolution layer 372e performs a convolution operation on the combined feature map supplied by the convolution layer 372d. The convolution layer 372e performs a process of recognizing the target object based on the combined feature map on which the convolution operation has been performed. The convolution layer 372e supplies the combined feature map on which the convolution operation has been performed to the convolution layer 372f.
[0113] The convolution layer 372f performs a convolution operation on the combined feature map supplied by the convolution layer 372e. The convolution layer 372f performs a process of recognizing a target object based on the combined feature map on which the convolution operation has been performed.
[0114] The object recognition model 351 outputs data indicating the recognition result of the target object performed by the convolutional layers 372 a to 372 f.
[0115] Note that the size (number of pixels) of the combined feature map becomes smaller in order from the convolution layer 372a, and is smallest in the convolution layer 372f. In addition, if the combined feature map has a larger size, then when viewed from the vehicle 10, a target object with a small size is recognized with higher accuracy, and if the combined feature map has a smaller size, then when viewed from the vehicle 10, a target object with a large size is recognized with higher accuracy. Therefore, for example, when the target object is a vehicle, a small distant vehicle is easily recognized in a combined feature map with a large size, and a large nearby vehicle is easily recognized in a combined feature map with a small size.
[0116] <Configuration Example of Learning System 401>
[0117] Figure 5 A configuration example of the learning system 401 is illustrated.
[0118] Learning System 401 Figure 3 The learning system 401 performs a learning process on the image processing model 301. The learning system 401 includes an input section 411, a correct answer data generator 412, a signal processor 413, a training data generator 414, and a learning section 415.
[0119] Input section 411 includes various input devices and is used to input, for example, data required for generating training data and operations performed by the user. For example, when a captured image is input, input section 411 supplies the captured image to correct answer data generator 412. For example, when millimeter wave data is input, input section 411 supplies the millimeter wave data to signal processor 413. For example, input section 411 supplies data indicating user instructions input through user operations to correct answer data generator 412 and training data generator 414.
[0120] Correct answer data generator 412 generates correct answer data based on the captured image. For example, a user specifies the location of a vehicle in the captured image via input section 411. Correct answer data generator 412 generates correct answer data indicating the location of the vehicle in the captured image based on the user's specified location. Correct answer data generator 412 supplies the correct answer data to training data generator 414.
[0121] The signal processor 413 performs the same Figure 2 The signal processor 413 performs a process similar to that performed by the signal processor 222. In other words, the signal processor 413 performs a specified signal process on the millimeter wave data to generate a millimeter wave image. The signal processor 413 supplies the millimeter wave image to the training data generator 414.
[0122] The training data generator 414 generates training data including input data including a millimeter wave image and correct answer data, and supplies the training data to the learning section 415 .
[0123] The learning section 415 uses the training data to perform a learning process on the image processing model 301. The learning section 415 outputs the image processing model 301 on which the learning has been performed.
[0124] <Configuration Example of Learning System 451>
[0125] Figure 6 A configuration example of the learning system 451 is illustrated.
[0126] Learning System 451 Figure 4 The object recognition model 351 performs learning processing. The learning system 451 includes an input part 461, an image processor 462, a correct answer data generator 463, a signal processor 464, an image processor 465, a training data generator 466 and a learning part 467.
[0127] Input section 461 includes various input devices and is used to input, for example, data required to generate training data and operations performed by the user. For example, when a captured image is input, input section 461 supplies the captured image to image processor 462 and correct answer data generator 463. For example, when millimeter wave data is input, input section 461 supplies the millimeter wave data to signal processor 464. For example, input section 461 supplies data indicating user instructions input through user operations to correct answer data generator 463 and training data generator 466.
[0128] Image processor 462 performs the same Figure 2 The image processor 462 performs a process similar to that performed by the image processor 221. In other words, the image processor 462 performs a specified image process on the captured image to generate an R image, a G image, and a B image. The image processor 462 supplies the R image, the G image, and the B image to the training data generator 466.
[0129] Correct answer data generator 463 generates correct answer data based on the captured image. For example, a user specifies the location of a vehicle in the captured image via input section 461. Based on the user-specified vehicle location, correct answer data generator 463 generates correct answer data indicating the vehicle's location in the captured image. Correct answer data generator 463 supplies the correct answer data to training data generator 466.
[0130] Signal processor 464 performs the same Figure 2 The signal processor 464 performs a process similar to that performed by the signal processor 222. In other words, the signal processor 464 performs a specified signal process on the millimeter wave data to generate a millimeter wave image. The signal processor 464 supplies the millimeter wave image to the image processor 465.
[0131] Image processor 465 performs the same Figure 2 The image processor 465 performs a process similar to that performed by the image processor 223. In other words, the image processor 465 generates an estimated position image based on the millimeter wave image. The image processor 465 supplies the estimated position image to the training data generator 466.
[0132] Note that the image processing model 301 on which learning has been performed is used for the image processor 465 .
[0133] The training data generator 466 generates training data including input data and correct answer data, wherein the input data includes four-channel image data including R image, G image, B image and estimated position image, and supplies the training data to the learning part 467.
[0134] The learning section 467 uses the training data to perform a learning process on the object recognition model 351. The learning section 467 outputs the object recognition model 351 on which learning has been performed.
[0135] <Learning Processing Performed on Image Processing Model>
[0136] Next, refer to Figure 7 The flowchart of describes the learning process of the image processing model performed by the learning system 401.
[0137] Note that before starting this process, data for generating training data is collected. For example, under actual driving conditions of vehicle 10, camera 201 and millimeter-wave radar 202 provided to vehicle 10 sense the area in front of vehicle 10. Specifically, camera 201 captures an image of the area in front of vehicle 10 and stores the captured image in storage device 111. Millimeter-wave radar 202 detects objects in front of vehicle 10 and stores the acquired millimeter-wave data in storage device 111. Training data is generated based on the captured image and millimeter-wave data accumulated in storage device 111.
[0138] In step S1 , the learning system 401 generates training data.
[0139] For example, the user inputs a captured image and millimeter wave data acquired at substantially the same time into the learning system 401 and inputs them through the input section 411. In other words, the captured image and millimeter wave data obtained by performing sensing at substantially the same time point are input into the learning system 401. The captured image is supplied to the correct answer data generator 412, and the millimeter wave data is supplied to the signal processor 413.
[0140] In addition, the user specifies an area in the captured image in which the target object exists. The correct answer data generator 412 generates correct answer data including a binary image indicating the area in which the target object specified by the user exists.
[0141] For example, through the input portion 411, the user frames an area in which a vehicle exists. Figure 8 The target object in the captured image 502. The correct answer data generator 412 generates correct answer data 503, which is an image binarized by filling the square portion with solid white and filling the other portions with solid black.
[0142] The correct answer data generator 412 supplies the correct answer data to the training data generator 414 .
[0143] The signal processor 413 performs specified signal processing on the millimeter wave data to estimate the position and speed of the object that has reflected the transmission signal in the area in front of the vehicle 10. The position of the object is represented by, for example, the distance from the vehicle 10 to the object and the direction (angle) of the object relative to the optical axis direction of the millimeter wave radar 202 (the driving direction of the vehicle 10). Note that, for example, when the transmission signal is radially transmitted, the optical axis direction of the millimeter wave radar 202 is the same as the direction of the center of the range in which the radial transmission is performed, and when scanning is performed with the transmission signal, the optical axis direction of the millimeter wave radar 202 is the same as the direction of the center of the range in which the scanning is performed. The speed of the object is represented by, for example, the relative speed of the object relative to the vehicle 10. The signal processor 413 generates a millimeter wave image based on the result of estimating the position of the object.
[0144] For example, generate Figure 8 Millimeter wave image 501 is shown. The x-axis of millimeter wave image 501 represents the angle of an object relative to the optical axis direction of millimeter wave radar 202 (the direction of travel of vehicle 10), and the y-axis of millimeter wave image 501 represents the distance to the object. In addition, in millimeter wave image 501, the intensity of the signal (received signal) reflected from the object at the position defined by the x-axis and y-axis is indicated by color or density.
[0145] The signal processor 413 supplies the millimeter wave image to the training data generator 414 .
[0146] Training data generator 414 generates training data including input data and correct answer data, the input data including millimeter wave images. For example, training data including input data and correct answer data 503 is generated, the input data including millimeter wave images 501. Training data generator 414 supplies the generated training data to learning section 415.
[0147] In step S2, the learning section 415 causes the image processing model to perform learning. Specifically, the learning section 415 inputs input data to the image processing model 301. The image processing model 301 generates an estimated position image based on the millimeter wave image included in the input data.
[0148] For example, based on the millimeter wave image 501, Figure 8 Estimated position image 504 is captured. Estimated position image 504 is a grayscale image whose coordinate system is exactly the same as that of captured image 502. Captured image 502 and estimated position image 504 are images of the area in front of vehicle 10 when viewed from the same viewpoint. In estimated position image 504, pixels that are more likely to be included in the area where the target object exists are brighter, while pixels that are less likely to be included in the area where the target object exists are darker.
[0149] The learning section 415 compares the estimated position image with the correct answer data and, based on the comparison result, adjusts, for example, the parameters of the image processing model 301. For example, the learning section 415 compares the estimated position image 504 with the correct answer data 503 and, for example, adjusts the parameters of the image processing model 301 so that the error is reduced.
[0150] In step S3, the learning section 415 determines whether learning is to be performed continuously. For example, when the learning performed by the image processing model 301 has not yet ended, the learning section 415 determines that learning is to be performed continuously, and the process returns to step S1.
[0151] Thereafter, the processes of steps S1 to S3 are repeatedly performed until it is determined in step S3 that the learning is to be terminated.
[0152] On the other hand, the learning section 415 determines in step S3 that the learning performed by the image processing model 301 is to be terminated when the learning has ended, for example, and terminates the learning process performed on the image processing model.
[0153] As described above, the image processing model 301 on which learning has been performed is generated.
[0154] <Learning Processing Performed on Object Recognition Model>
[0155] Next, refer to Figure 9 The flowchart of describes the learning process performed by the learning system 451 regarding the object recognition model.
[0156] Note that, just as before starting the learning process for the image processing model, data for generating training data is collected before starting this process. Note that the same captured images and the same millimeter wave data can be used for both the learning process for the image processing model and the learning process for the object recognition model.
[0157] In step S51 , the learning system 451 generates training data.
[0158] For example, the user inputs a captured image and millimeter wave data acquired at substantially the same time into the learning system 451 and inputs them through the input section 461. In other words, the captured image and millimeter wave data obtained by performing sensing at substantially the same time point are input into the learning system 451. The captured image is supplied to the image processor 462 and the correct answer data generator 463, and the millimeter wave data is supplied to the signal processor 464.
[0159] The image processor 462 performs a process of interpolating the R component, the G component, and the B component in each pixel of the captured image to generate an R image composed of the R component of the captured image, a G image composed of the G component of the captured image, and a B image composed of the B component of the captured image. Figure 10 The image processor 462 generates an R image 552R, a G image 552G, and a B image 552B from the captured image 551. The image processor 462 supplies the R image, the G image, and the B image to the training data generator 466.
[0160] Signal processor 464 performs the same operation as signal processor 413 in Figure 7 The millimeter wave image is generated based on the millimeter wave data by performing a similar process to that performed in step S1. Figure 10 The signal processor 464 supplies the millimeter wave image 553 to the image processor 465.
[0161] The image processor 465 inputs the millimeter wave image to the image processing model 301 to generate an estimated position image. For example, Figure 10 The image processor 462 supplies the estimated position image 554 to the training data generator 466.
[0162] In addition, the user specifies the position where the target object exists in the captured image through the input part 461. The correct answer data generator 463 generates correct answer data indicating the position of the vehicle in the captured image based on the position of the target object specified by the user. For example, the correct answer data generated from the captured image 551 Figure 10 The correct answer data 555 includes the framed vehicle as the target object in the captured image 551 . The correct answer data generator 463 supplies the correct answer data to the training data generator 466 .
[0163] Training data generator 466 generates training data including input data and correct answer data, the input data including four-channel image data of R image, G image, B image, and estimated position image. For example, training data including input data including four-channel image data of R image 552R, G image 552G, B image 552B, and estimated position image 554, and correct answer data is generated. Training data generator 466 supplies the training data to learning section 467.
[0164] In step S52, the learning section 467 causes the object recognition model 351 to perform learning. Specifically, the learning section 467 inputs the input data included in the training data to the object recognition model 351. The object recognition model 351 recognizes the target object in the captured image 551 based on the R image, G image, B image, and estimated position image included in the input data, and generates recognition result data indicating the result of the recognition. For example, Figure 10 Recognition result data 556. In the recognition result data 556, the vehicle as the recognized target object is framed.
[0165] The learning section 467 compares the recognition result data with the correct answer data and, based on the comparison result, adjusts, for example, the parameters of the object recognition model 351. For example, the learning section 467 compares the recognition result data 556 with the correct answer data 555 and, for example, adjusts the parameters of the object recognition model 351 so that the error is reduced.
[0166] In step S53, the learning section 467 determines whether learning is to be performed continuously. For example, when the learning performed by the object recognition model 351 has not yet ended, the learning section 467 determines that learning is to be performed continuously, and the process returns to step S51.
[0167] Thereafter, the processing of steps S51 to S53 is repeatedly performed until it is determined in step S53 that the learning is to be terminated.
[0168] On the other hand, the learning section 467 determines in step S53 that the learning performed by the object recognition model 351 is to be terminated, for example, when the learning has ended, and terminates the learning process performed on the object recognition model.
[0169] As described above, the object recognition model 351 on which learning has been performed is generated.
[0170] <Target Object Recognition Processing>
[0171] Next, refer to Figure 11 The flowchart of FIG. 1 describes the target object recognition process performed by the vehicle 10 .
[0172] This process is started when, for example, an operation for activating the vehicle 10 to start driving is performed (i.e., when the ignition switch, power switch, start switch, etc. of the vehicle 10 are turned on). This process is terminated when, for example, an operation for terminating driving of the vehicle 10 is performed (i.e., when the ignition switch, power switch, start switch, etc. of the vehicle 10 are turned off).
[0173] In step S101 , the camera 201 and the millimeter wave radar 202 perform sensing of an area located in front of the vehicle 10 .
[0174] Specifically, the camera 201 captures an image of an area located in front of the vehicle 10 and supplies the obtained captured image to the image processor 221 .
[0175] Millimeter wave radar 202 transmits a transmission signal in the forward direction of vehicle 10 and receives a reception signal using a plurality of reception antennas, which is a signal reflected from an object located in front of vehicle 10. Millimeter wave radar 202 supplies millimeter wave data, which indicates the strength of the reception signal received using each reception antenna in time sequence, to signal processor 222.
[0176] In step S102, the image processor 221 performs pre-processing on the captured image. Specifically, the image processor 221 performs Figure 9 The image processor 221 performs similar processing to the processing performed by the image processor 462 in step S51 to generate an R image, a G image, and a B image based on the captured image. The image processor 221 supplies the R image, the G image, and the B image to the object recognition section 224.
[0177] In step S103, the signal processor 222 generates a millimeter wave image. Specifically, the signal processor 222 performs the same operation as in Figure 7 The signal processor 222 performs a process similar to the process performed by the signal processor 413 in step S1 to generate a millimeter wave image based on the millimeter wave data. The signal processor 222 supplies the millimeter wave image to the image processor 223.
[0178] In step S104, the image processor 223 generates an estimated position image based on the millimeter wave image. Specifically, the image processor 223 performs the same Figure 9 The image processor 223 performs processing similar to that performed by the image processor 465 in step S51 to generate an estimated position image based on the millimeter wave image. The image processor 223 supplies the estimated position image to the object recognition section 224.
[0179] In step S105, the object recognition section 224 performs processing to identify the target object based on the captured image and the estimated position image. Specifically, the object recognition section 224 inputs input data comprising four channels of image data (R image, G image, B image, and estimated position image) to the object recognition model 351. The object recognition model 351 performs processing to identify the target object located in front of the vehicle 10 based on the input data.
[0180] The object recognition section 224 supplies data indicating the result of recognizing the target object to, for example, the own position estimator 132; the map analyzer 151, traffic rule recognition section 152 and situation recognition section 153 of the situation analyzer 133; and the emergency avoidance section 171 of the motion controller 135.
[0181] The own position estimator 132 performs processing of estimating the position, posture, and the like of the vehicle 10 based on, for example, the result of recognizing the target object.
[0182] For example, based on the result of recognizing the target object, the map analyzer 151 performs a process of analyzing various maps stored in the storage device 111 and constructs a map including information required for the autonomous driving process.
[0183] For example, based on the result of recognizing the target object, the traffic regulation recognition portion 152 performs a process of recognizing traffic regulations around the vehicle 10 .
[0184] For example, based on the result of recognizing the target object, the situation recognition portion 153 performs a process of recognizing the situation of the surroundings of the vehicle 10 .
[0185] When the emergency avoiding portion 171 detects the occurrence of an emergency based on, for example, a result of recognizing a target object, the emergency avoiding portion 171 plans the movement of the vehicle 10 , such as sudden stopping or quick turning, to avoid the emergency.
[0186] Thereafter, the process returns to step S101 , and the processes of step S101 and thereafter are performed.
[0187] As described above, the accuracy of recognizing a target object located in front of the vehicle 10 can be improved.
[0188] Figure 12 These are radar charts comparing the properties of target objects identified when only the camera 201 (image sensor 201A) is used, when only the millimeter-wave radar 202 is used, and when both the camera 201 and the millimeter-wave radar 202 are used. Graph 601 illustrates the characteristics of identification performed when only the camera 201 is used. Graph 602 illustrates the characteristics of identification performed when only the millimeter-wave radar 202 is used. Graph 603 illustrates the characteristics of identification performed when both the camera 201 and the millimeter-wave radar 202 are used.
[0189] The radar map is defined by six axes: range accuracy, interference-free performance, material independence, bad weather, night driving, and horizontal angular resolution.
[0190] The distance accuracy axis indicates the accuracy of the distance of the object being detected. If the accuracy of the distance of the object being detected is high, this axis indicates a larger value, whereas if the accuracy of the distance of the object being detected is low, this axis indicates a smaller value.
[0191] The axis of interference-free performance represents a condition that is less susceptible to interference from other electromagnetic waves. This axis shows a larger value in a condition that is less susceptible to interference from other electromagnetic waves, and a smaller value in a condition that is more susceptible to interference from other electromagnetic waves.
[0192] The axis of material independence indicates whether the recognition accuracy is less affected by the material type. If the recognition accuracy is less affected by the material type, this axis shows a larger value, while if the recognition accuracy is greatly affected by the material type, this axis shows a smaller value.
[0193] The bad weather axis represents the accuracy of object recognition during bad weather. If the accuracy of object recognition during bad weather is high, this axis shows a large value, and if the accuracy of object recognition during bad weather is low, this axis shows a small value.
[0194] The night driving axis represents the accuracy of object recognition during night driving. If the accuracy of object recognition during night driving is high, this axis shows a large value, and if the accuracy of object recognition during night driving is low, it shows a small value.
[0195] The axis of horizontal angular resolution represents the horizontal (lateral) angular resolution in the position of the recognized object. If the horizontal angular resolution is high, this axis shows a large value, and if the horizontal angular resolution is low, this axis shows a small value.
[0196] Camera 201 outperforms millimeter-wave radar 202 in terms of interference-free performance, material independence, and horizontal angular resolution. Meanwhile, millimeter-wave radar 202 outperforms camera 201 in terms of range accuracy, recognition accuracy during bad weather, and recognition accuracy during nighttime driving. Therefore, when both camera 201 and millimeter-wave radar 202 are used to fuse recognition results, they can compensate for each other's weaknesses, resulting in improved accuracy in identifying target objects.
[0197] For example, Figure 13 A illustrates an example of a recognition result obtained when processing for recognizing a vehicle is performed using only the camera 201, and Figure 13 B illustrates an example of a recognition result obtained when the process of recognizing a vehicle is performed using both the camera 201 and the millimeter wave radar 202 .
[0198] In both cases, vehicles 621 to 623 were identified. When only camera 201 was used, this resulted in failure to identify vehicle 624, which was partially hidden behind vehicles 622 and 623. On the other hand, when both camera 201 and millimeter wave radar 202 were used simultaneously, vehicle 624 was successfully identified.
[0199] For example, Figure 14 A illustrates an example of a recognition result obtained when processing for recognizing a vehicle is performed using only the camera 201, and Figure 14 B illustrates an example of a recognition result obtained when the process of recognizing a vehicle is performed using both the camera 201 and the millimeter wave radar 202 .
[0200] In both cases, vehicle 641 was recognized. When only camera 201 was used, vehicle 642, which had a unique shape and color, could not be recognized. On the other hand, when both camera 201 and millimeter wave radar 202 were used, vehicle 642 was successfully recognized.
[0201] In addition, when the process of recognizing the target object is performed using the estimated position image instead of the millimeter wave image, this results in improved accuracy in recognizing the target object.
[0202] Specifically, a geometric transformation is performed on the millimeter wave image to obtain an estimated position image whose coordinate system matches the coordinate system of the captured image, and the object recognition model 351 is learned using the estimated position image. This makes it easier to match each pixel of the captured image with a reflection point (a point where the intensity of the received signal is high) in the estimated position image and improves the accuracy of learning. In addition, in the estimated position image, the component of the received signal that is included in the millimeter wave image and reflected from objects other than the target object located in front of the vehicle 10 (i.e., the component that is not required for performing the process of identifying the target object) is reduced. Therefore, using the estimated position image can improve the accuracy of identifying the target object.
[0203] <<2. Second embodiment>>
[0204] Next, refer to Figure 15 A second embodiment of the present technology is described.
[0205] <Configuration Example of Data Acquisition Section 102B and Vehicle Exterior Information Detector 141B>
[0206] Figure 15 As shown in the figure Figure 1 The data acquisition section 102B of the second embodiment of the data acquisition section 102 of the vehicle control system 100 and the data acquisition section 102B of the second embodiment of the vehicle control system 100 Figure 1 An example of the configuration of the vehicle exterior information detector 141B of the second embodiment of the exterior information detector 141 in the vehicle control system 100. Figure 2 The part corresponding to the part in Figure 2 The same reference numerals are denoted, and descriptions thereof are appropriately omitted.
[0207] The data acquisition section 102B is similar to the data acquisition section 102A in including the camera 201 and the millimeter wave radar 202 , and is different from the data acquisition section 102A in including the LiDAR 701 .
[0208] The vehicle exterior information detector 141B differs from the vehicle exterior information detector 141A in that it includes an information processor 711 instead of the information processor 211. The information processor 711 is similar to the information processor 211 in that it includes an image processor 221, a signal processor 222, and an image processor 223. On the other hand, the information processor 711 differs from the information processor 211 in that it includes an object recognition section 723 instead of the object recognition section 224, and in that a signal processor 721 and an image processor 722 are added.
[0209] The LiDAR 701 senses the area in front of the vehicle 10, and the sensing ranges of the LiDAR 701 and the camera 201 at least partially overlap. For example, the LiDAR 701 scans the area in front of the vehicle 10 in the lateral and height directions using laser pulses, and receives reflected light as a reflection of the laser pulses. The LiDAR 701 calculates the distance to the object in front of the vehicle 10 based on the time it takes to receive the reflected light, and based on the calculated results, the LiDAR 701 generates three-dimensional point group data (point cloud) indicating the shape and position of the object in front of the vehicle 10. The LiDAR 701 supplies the point group data to the signal processor 721.
[0210] The signal processor 721 performs designated signal processing (for example, interpolation processing or number reduction processing) on the point group data, and supplies the point group data on which the signal processing has been performed to the image processor 722 .
[0211] The image processor 722 performs specified image processing on the point group data to generate an estimated position image indicating the estimated position of the target object in exactly the same coordinate system as that of the captured image, as in the case of the image processor 223. The image processor 722 supplies the estimated position image to the object recognition section 723.
[0212] Note that, similar to Figure 3 The image processing model 301 of the image processing model is used for the image processor 722, but its detailed description is omitted. The image processing model for the image processor 722 is learned using training data including input data and correct answer data, the input data including point group data, and the correct answer data generated based on the captured image.
[0213] Note that the estimated position image generated by the image processor 223 based on the millimeter wave image is referred to as a millimeter wave-based estimated position image, and the estimated position image generated by the image processor 722 based on the point group data is referred to as a point group-based estimated position image.
[0214] The object recognition section 723 performs processing for recognizing a target object located in front of the vehicle 10 based on the R image, the G image, the B image, the millimeter wave-based estimated position image, and the point group-based estimated position image. The object recognition section 723 supplies data indicating the result of recognizing the target object to, for example, the own position estimator 132; the map analyzer 151, the traffic regulation recognition section 152, and the situation recognition section 153 of the situation analyzer 133; and the emergency avoidance section 171 of the motion controller 135.
[0215] Note that, unlike e.g. Figure 4 An object recognition model similar to the object recognition model 351 is used for the object recognition portion 723, but its detailed description is omitted. The object recognition model used for the object recognition portion 723 is learned using training data including input data and correct answer data, the input data including five-channel image data, namely, an R image, a G image, a B image, an estimated position image based on millimeter waves, and an estimated position image based on point groups, and the correct answer data is generated based on the captured image.
[0216] As described above, the addition of the LiDAR 701 leads to further improvement in the accuracy of identifying target objects.
[0217] <<3. Modification>>>
[0218] Modifications of the embodiments of the present technology described above are described below.
[0219] The above description primarily describes an example in which a vehicle is the target of recognition. However, as described above, any object other than a vehicle can be the target of recognition. For example, it is sufficient if the learning process is performed on the image processing model 301 and the object recognition model 351 using training data including correct answer data indicating the position of the target object to be recognized.
[0220] In addition, the present technology is also applicable to cases where multiple types of objects are to be recognized. For example, it is sufficient if the learning process is performed on the image processing model 301 and the object recognition model 351 using training data including correct answer data indicating the position and label (type of the target object) of each target object.
[0221] Already referenced Figures 7 to 9 The example in which the learning process is executed while the training data is generated has been described. However, for example, the learning process may be executed after necessary training data is generated in advance.
[0222] The above has described an example of recognizing a target object located in front of the vehicle 10. However, the present technology is also applicable to a case of recognizing a target object located around the vehicle 10 in another direction when viewed from the vehicle 10.
[0223] In addition, the present technology is also applicable to cases where target objects are identified around mobile objects other than vehicles. For example, it is conceivable that the present technology can be applied to mobile objects such as motorcycles, bicycles, personal vehicles, airplanes, ships, construction machinery, and agricultural machinery (tractors). In addition, examples of mobile objects to which the present technology is applicable also include mobile objects that are remotely operated by a user (such as drones and robots) without the user having to board the mobile object.
[0224] Furthermore, the present technology is also applicable to the case where a process of recognizing a target object at a fixed position, such as a surveillance system, is performed.
[0225] and, Figure 3 Image processing model 301 and Figure 4 The object recognition model 351 is merely an example, and models other than the image processing model 301 and the object recognition model 351 generated by machine learning may also be used.
[0226] In addition, the present technology is also applicable to a case where a process of recognizing a target object is performed by using a camera (image sensor) and LiDAR in combination.
[0227] Furthermore, the present technology is also applicable to cases where sensors other than millimeter-wave radar or LiDAR are used to detect objects.
[0228] In addition, the present technology is also applicable when the millimeter wave radar has resolution in the height direction (that is, when the millimeter wave radar can detect the position (angle) of an object in the height direction).
[0229] For example, when the resolution of the millimeter wave radar in the height direction is 6, millimeter wave images 801a to 801f corresponding to different heights are generated based on the millimeter wave data, as shown in FIG. Figure 16 In this case, for example, it is sufficient if the image processing model 301 can be learned using training data including input data including six-channel image data 802 as millimeter wave images 801a to 801f.
[0230] Alternatively, for example, the millimeter wave images 801a to 801f may be combined to generate a single millimeter wave image, and the image processing model 301 may be learned using training data including input data including the generated millimeter wave image.
[0231] Alternatively, for example, you can use Figure 17 The millimeter wave image of B 822 instead Figure 17 Millimeter wave image 821 of A.
[0232] In the millimeter wave image 821, as shown in Figure 8 As in the case of the millimeter wave image 501 , the x-axis represents the angle of the object relative to the optical axis direction of the millimeter wave radar 202 , and the y-axis represents the distance to the object.
[0233] On the other hand, in millimeter wave image 822, the x-axis represents the lateral direction (the width direction of vehicle 10), and the y-axis represents the optical axis direction of millimeter wave radar 202 (the traveling direction of vehicle 10). In millimeter wave image 822, the positions of objects located in front of vehicle 10 and the distribution of reflection intensity of each object (i.e., the distribution of the intensity of the received signal reflected from the objects located in front of vehicle 10) are given in a bird's-eye view.
[0234] Millimeter-wave image 822 is generated based on millimeter-wave image 821, and the use of millimeter-wave image 822 makes it easier to visually grasp the position of an object located in front of vehicle 10 than when millimeter-wave image 821 is used. However, when millimeter-wave image 821 is converted to millimeter-wave image 822, some information is lost. Therefore, when millimeter-wave image 821 is used without any change, the accuracy of identifying the target object is higher.
[0235] <<4. Others>>
[0236] <Computer Configuration Example>
[0237] The above series of processes can be performed using hardware or software. When the series of processes are performed using software, the program included in the software is installed on a computer. Here, examples of computers include computers incorporated into dedicated hardware, as well as computers such as general-purpose personal computers that can perform various functions through various programs installed thereon.
[0238] Figure 18 is a block diagram of a hardware configuration example of a computer that executes the above-described series of processes using a program.
[0239] In a computer 1000 , a central processing unit (CPU) 1001 , a read only memory (ROM) 1002 , and a random access memory (RAM) 1003 are connected to one another via a bus 1004 .
[0240] In addition, an input / output interface 1005 is connected to the bus 1004 . An input section 1006 , an output section 1007 , a recording section 1008 , a communication section 1009 , and a drive 1010 are connected to the input / output interface 1005 .
[0241] The input section 1006 includes, for example, an input switch, buttons, a microphone, and an imaging element. The output section 1007 includes, for example, a display and a speaker. The recording section 1008 includes, for example, a hard disk and non-volatile memory. The communication section 1009 includes, for example, a network interface. The drive 1010 drives a removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0242] In the computer 1000 having the above configuration, the above series of processing is performed by the CPU 1001 , for example, loading a program recorded in the recording section 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing the program.
[0243] For example, the program can be provided by recording the program executed by computer 1000 (CPU 1001) in removable medium 1011 serving as, for example, a package medium. In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0244] In the computer 1000, by mounting the removable medium 1011 on the drive 1010, the program can be installed on the recording section 1008 via the input / output interface 1005. In addition, the program can be received by the communication section 1009 via a wired or wireless transmission medium to be installed on the recording section 1008. Moreover, the program can be installed in advance on the ROM 1002 or the recording section 1008.
[0245] Note that the program executed by the computer may be one in which processing is performed time-sequentially in the order described herein, or may be one in which processing is performed in parallel or at necessary timing such as timing of calling.
[0246] In addition, the system used in this article refers to a collection of multiple components such as devices and modules (parts), and it does not matter whether all components are in a single housing. Therefore, multiple devices housed in separate housings and connected to each other via a network, as well as a single device in which multiple modules are housed in a single housing, are both systems.
[0247] Furthermore, the embodiments of the present technology are not limited to the above-described examples, and various modifications may be made thereto without departing from the scope of the present technology.
[0248] For example, the present technology may also have a configuration of cloud computing in which a single function is shared to be commonly processed by a plurality of devices via a network.
[0249] Additionally, in addition to being performed by a single device, the corresponding steps described using the above flowcharts may be shared so as to be performed by a plurality of devices.
[0250] Furthermore, when a single step includes a plurality of processes, in addition to being executed by a single device, the plurality of processes included in the single step may be shared so as to be executed by a plurality of devices.
[0251] <Example of configuration combination>
[0252] The present technology can also adopt the following configurations.
[0253] (1) An information processing device comprising:
[0254] an image processor that generates an estimated position image based on a sensor image, the sensor image indicating, using a first coordinate system, a sensing result of a sensor whose sensing range at least partially overlaps with a sensing range of the image sensor, the estimated position image indicating an estimated position of the target object in a second coordinate system that is the same as a coordinate system of a captured image obtained by the image sensor; and
[0255] An object recognition section performs a process of recognizing a target object based on the captured image and the estimated position image.
[0256] (2) The information processing device according to (1), wherein
[0257] The image processor generates an estimated position image using an image processing model obtained through machine learning.
[0258] (3) The information processing device according to (2), wherein
[0259] The image processing model is learned using training data comprising input data including sensor images and correct answer data indicating the location of a target object in the captured image.
[0260] (4) The information processing device according to (3), wherein
[0261] The correct answer data is a binary image indicating an area where the target object exists in the captured image.
[0262] (5) The information processing device according to (3) or (4), wherein
[0263] The image processing model is a model that uses a deep neural network.
[0264] (6) The information processing device according to (5), wherein
[0265] Image processing models include
[0266] a feature quantity extraction section that extracts feature quantities of a sensor image to generate a feature map indicating distribution of the feature quantities in a first coordinate system, a geometric transformation section that transforms the feature map in the first coordinate system into a feature map in a second coordinate system, and
[0267] A deconvolution part deconvolves the feature map in the second coordinate system to generate an estimated position image.
[0268] (7) The information processing device according to any one of (1) to (6), wherein
[0269] The object recognition section performs processing for recognizing a target object using an object recognition model obtained through machine learning.
[0270] (8) The information processing device according to (7), wherein
[0271] The object recognition model is learned using training data including input data including a captured image and an estimated position image, and correct answer data indicating a position of a target object in the captured image.
[0272] (9) The information processing device according to (8), wherein
[0273] The object recognition model is a model that uses a deep neural network.
[0274] (10) The information processing device according to (9), wherein
[0275] Object recognition models include
[0276] a first convolutional neural network that extracts feature quantities of the captured image and the estimated position image, and
[0277] A second convolutional neural network is provided, which recognizes the target object based on feature quantities of the captured image and the estimated position image.
[0278] (11) The information processing device according to any one of (1) to (10), wherein
[0279] The image sensor and the sensor perform sensing on the surroundings of the moving object, and
[0280] The object recognition section performs processing for recognizing a target object around a moving object.
[0281] (12) The information processing device according to any one of (1) to (11), wherein
[0282] Sensors include millimeter-wave radar, and
[0283] The sensor image indicates the position of the object from which the transmission signal from the millimeter wave radar was reflected.
[0284] (13) The information processing device according to (12), wherein
[0285] The first coordinate system is defined by an axis representing an angle to the optical axis direction of the millimeter wave radar and an axis representing a distance to an object.
[0286] (14) The information processing device according to (12), wherein
[0287] Millimeter wave radar has high resolution and
[0288] The image processor generates an estimated position image based on a plurality of sensor images corresponding to different heights.
[0289] (15) The information processing device according to any one of (1) to (14), wherein
[0290] Sensors include LiDAR (Light Detection and Ranging), and
[0291] The sensor image is the point group data obtained by LiDAR.
[0292] (16) An information processing method comprising:
[0293] generating, by the information processing device, an estimated position image based on a sensor image, the sensor image indicating, using a first coordinate system, a sensing result of a sensor whose sensing range at least partially overlaps with a sensing range of the image sensor, the estimated position image indicating an estimated position of the target object in a second coordinate system that is the same as a coordinate system of a captured image obtained by the image sensor; and
[0294] The process of recognizing the target object is performed by the information processing device based on the captured image and the estimated position image.
[0295] (17) A program for causing a computer to execute a process, comprising:
[0296] generating an estimated position image based on the sensor image, the sensor image indicating, using a first coordinate system, a sensing result of a sensor whose sensing range at least partially overlaps with a sensing range of the image sensor, the estimated position image indicating an estimated position of the target object in a second coordinate system that is the same as a coordinate system of a captured image obtained by the image sensor; and
[0297] A process of recognizing a target object is performed based on the captured image and the estimated position image.
[0298] (18) A mobile object control device comprising:
[0299] an image processor that generates an estimated position image based on a sensor image, the sensor image indicating, using a first coordinate system, a sensing result of a sensor whose sensing range at least partially overlaps with a sensing range of an image sensor that captures an image around the moving object, the estimated position image indicating an estimated position of the target object in a second coordinate system that is the same as the coordinate system of the captured image obtained by the image sensor;
[0300] an object recognition section that performs processing for recognizing a target object based on the captured image and the estimated position image; and
[0301] A motion controller controls the motion of the mobile object based on the recognition result of the target object.
[0302] (19) A moving object comprising:
[0303] Image sensor;
[0304] a sensor having a sensing range that at least partially overlaps with a sensing range of the image sensor;
[0305] an image processor that generates an estimated position image based on a sensor image, the sensor image indicating a sensing result of the sensor using a first coordinate system, the estimated position image indicating an estimated position of the target object in a second coordinate system that is the same as a coordinate system of a captured image obtained by the image sensor;
[0306] an object recognition section that performs processing for recognizing a target object based on the captured image and the estimated position image; and
[0307] A motion controller controls the motion of the mobile object based on the recognition result of the target object.
[0308] Note that the effects described herein are not limiting but merely illustrative, and other effects may be provided.
[0309] Reference Signs List
[0310] 10 vehicles
[0311] 100 Vehicle Control Systems
[0312] 102, 102A, 102B data acquisition section
[0313] 107 Powertrain Controller
[0314] 108 Powertrain System
[0315] 135 Motion Controller
[0316] 141, 141A, 141B Vehicle external information detector
[0317] 201 Camera
[0318] 201A Image Sensor
[0319] 202 millimeter-wave radar
[0320] 211 Information Processor
[0321] 221 Image Processor
[0322] 222 Signal Processor
[0323] 223 Image Processor
[0324] 224 Object Recognition
[0325] 301 Image Processing Model
[0326] 311 Feature Extraction
[0327] 312 Geometric Transformation Section
[0328] 313 Deconvolution part
[0329] 351 Object Recognition Model
[0330] 361 Feature Extraction
[0331] 362 Identification Section
[0332] 401 Learning System
[0333] 414 Training Data Generator
[0334] 415 Learning Section
[0335] 451 Learning System
[0336] 466 Training Data Generator
[0337] 467 Learning Section
[0338] 701 LiDAR
[0339] 711 Information Processor
[0340] 721 Signal Processor
[0341] 722 Image Processor
[0342] 723 Object Recognition
Claims
1. An information processing device, comprising: an image sensor having a sensing range and adapted to provide a captured image; a radar sensor and a signal processor portion adapted to sense a range within a sensing range and provide a sensor image, the sensor image indicating a sensing result of the radar sensor in a first coordinate system, the sensing range of the radar sensor at least partially overlapping with the sensing range of the image sensor; an image processor adapted to receive the sensor image and perform a geometric transformation to generate an estimated position image based on the sensor image, the estimated position image indicating an estimated position of the target object in a second coordinate system that is the same as the coordinate system of the captured image; as well as an object recognition section adapted to perform a process of recognizing a target object based on the captured image and the estimated position image; wherein the image processor is adapted to generate the estimated position image using an image processing model obtained through machine learning; wherein the image processing model is caused to perform learning using training data comprising input data and correct answer data, the input data comprising the sensor image, and the correct answer data indicating a location of a target object in the captured image; and In which, the object recognition part is suitable for using an object recognition model obtained through machine learning to perform the processing of identifying the target object, wherein the object recognition model is used to perform learning using training data including input data and correct answer data, the input data including the captured image and the estimated position image, and the correct answer data indicating the position of the target object in the captured image.
2. The information processing apparatus according to claim 1, wherein The correct answer data is a binary image indicating an area where a target object exists in the captured image.
3. The information processing apparatus according to claim 1, wherein The image processing model is a model using a deep neural network, and the image processing model includes a feature amount extraction section that extracts a feature amount of the sensor image to generate a feature map indicating a distribution of the feature amount in a first coordinate system, a geometric transformation part, which transforms the feature map in the first coordinate system into a feature map in the second coordinate system, and A deconvolution part is used to deconvolve the feature map in the second coordinate system to generate the estimated position image. The information processing apparatus according to claim 1 , wherein The object recognition model is a model using a deep neural network, and the object recognition model includes a first convolutional neural network that extracts feature quantities of the captured image and the estimated position image, and A second convolutional neural network is configured to recognize a target object based on feature quantities of the captured image and the estimated position image.
5. The information processing device according to any one of claims 1 to 4, wherein The image sensor and the sensor perform sensing on the surroundings of a moving object, and The object recognition section performs a process of recognizing a target object around the moving object. The information processing apparatus according to claim 1 , wherein The sensor includes a millimeter wave radar having a resolution in a height direction, wherein the image processor generates the estimated position image based on a plurality of sensor images corresponding to different heights, and The sensor image indicates a position of an object from which a transmission signal from the millimeter wave radar is reflected.
7. The information processing apparatus according to claim 6, wherein A first coordinate system is defined by an axis representing an angle to an optical axis direction of the millimeter wave radar and an axis representing a distance to the object. The information processing apparatus according to claim 1 , wherein The image sensor includes a light detection and ranging LiDAR, and The sensor image is point group data obtained by LiDAR.
9. An information processing method comprising: providing a captured image by an image sensor having a sensing range; providing a sensor image by a sensor and a signal processor portion adapted to sense a range within a sensing range, the sensor image indicating a sensing result of the sensor in a first coordinate system, the sensing range of the sensor at least partially overlapping with a sensing range of the image sensor; generating, by the information processing device, an estimated position image based on the sensor image, the estimated position image indicating an estimated position of the target object in a second coordinate system that is the same as the coordinate system of the captured image; as well as executing, by the information processing device, a process of recognizing a target object based on the captured image and the estimated position image; wherein generating the estimated position image comprises using an image processing model obtained through machine learning; wherein the image processing model is caused to perform learning using training data comprising input data and correct answer data, the input data comprising the sensor image, and the correct answer data indicating a location of a target object in the captured image; and The object recognition part uses an object recognition model obtained through machine learning to perform processing for recognizing the target object, wherein the object recognition model is caused to perform learning using training data including input data and correct answer data, the input data including the captured image and the estimated position image, and the correct answer data indicating the position of the target object in the captured image.
10. A computer-readable storage medium comprising instructions that, when executed by a processor, cause the processor to perform the method of claim 9.
11. A mobile object control device, comprising: The information processing device according to any one of claims 1 to 8; as well as A motion controller controls the motion of the mobile object based on the recognition result of the target object.
12. A mobile object, comprising: The mobile object control device according to claim 11.
Citation Information
Patent Citations
Information processing device, information processing method and program
CN108139476A