Object recognition devices and storage media
By combining image information from external cameras and ranging sensors, and utilizing the intensity difference between reflected light and background light, a neural network is used for plant identification, solving the problem of high-precision plant identification around vehicles and improving the accuracy and efficiency of autonomous driving.
Patent Information
- Application Number
- CN202080066071.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-23
- Filing Date
- 2020-09-08
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2040-09-08
AI Technical Summary
In existing technologies, it is difficult to identify plants around vehicles with high precision, especially in visible light images where plants are difficult to distinguish from other obstacles, affecting the accuracy of autonomous driving.
By combining image information from external cameras and range sensors mounted on the vehicle, and utilizing the intensity differences between visible light information from the camera images and reflected light and background light images from the range sensors, plant identification is performed through neural networks. In particular, plant discriminant algorithms are used to calculate discriminant values to distinguish plant regions.
It achieves high-precision plant identification, reduces misclassification, improves the accuracy of autonomous driving, and reduces the burden of processing volume and speed.
Smart Images

Figure CN114514565B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is based on Japanese Patent Application No. 2019-172398, filed in Japan on September 23, 2019, and is incorporated herein by reference in its entirety. Technical Field
[0003] This specification discloses an object recognition device and an object recognition program. Background Technology
[0004] Patent document 1 discloses an apparatus for recognizing objects in an image of a road scene, that is, an object recognition apparatus for recognizing objects around a vehicle.
[0005] Patent Document 1: Japanese Patent Application Publication No. 2017-162456
[0006] In addition, when vehicles are in motion, they often encounter scenarios where vegetation such as grass and trees grow too large and encroaches on the roadside. In object recognition, the impact on vehicle operation (e.g., the actions selected in autonomous driving) varies depending on whether the intruding object is simply vegetation or other obstacles that should be avoided.
[0007] In Patent Document 1, the image used for object recognition during actual driving is considered to be a camera image generated by a camera element sensing visible light from the outside world. There are concerns that the intensity of visible light represented by such a camera image may not allow for highly accurate identification of plants and other obstacles that should be avoided. Summary of the Invention
[0008] One of the purposes of this specification is to provide an object recognition device and object recognition program capable of identifying plants with high precision.
[0009] One disclosed method provides an object recognition device that uses image information from both an external camera mounted on a vehicle and a range sensor mounted on the vehicle to identify objects around the vehicle. The external camera image information is image information including a camera image generated by a camera element sensing visible light from the outside environment. The range sensor image information is image information including a reflected light image and a background light image. The reflected light image is generated by a light-receiving element sensing near-infrared reflected light from an object after illumination. The background light image is generated by a light-receiving element sensing near-infrared background light relative to the reflected light. The object recognition device comprises:
[0010] The image information acquisition unit acquires image information from the ranging sensor and the external camera; and
[0011] The recognition unit identifies plants around the vehicle by considering the difference between the intensity of the camera image and the intensity of the sensor image.
[0012] Another disclosed method provides an object recognition program that uses image information from both an external camera mounted on the vehicle and a range sensor mounted on the vehicle to identify objects around the vehicle. The external camera image information is image information containing a camera image generated by a camera element sensing visible light from the outside world. The range sensor image information is image information containing a reflected light image and a background light image. The reflected light image is generated by a light-receiving element sensing near-infrared reflected light from an object after illumination. The background light image is generated by a light-receiving element sensing near-infrared background light relative to the reflected light. This object recognition program causes at least one processing unit to perform the following processing:
[0013] Processing of image information acquired from ranging sensors and external cameras; and
[0014] The process of identifying vegetation around a vehicle by considering the difference in intensity between the camera image and the sensor image.
[0015] Based on these methods, in the identification of plant regions that reflect plants, the difference between the intensity of the camera image and the intensity of the sensor image is considered. That is, the difference between the information perceived in the camera image and the information perceived in the near-infrared light in the range sensor image. Therefore, it is possible to easily distinguish plants with significantly different spectroscopic characteristics in the visible and near-infrared regions, as well as other objects that should be avoided as they do not exhibit the same degree of difference between the visible and near-infrared regions. Thus, object recognition capable of identifying plants with high accuracy can be achieved.
[0016] Furthermore, the reference numerals in parentheses in the claims and the like illustratively indicate the correspondence between parts of the embodiments described below, and are not intended to limit the scope of the technology. Attached Figure Description
[0017] Figure 1 This is a diagram showing the overall image of the object recognition system and the driving assistance ECU according to the first embodiment.
[0018] Figure 2 This diagram shows the mounting status of the ranging sensor and the external camera in the vehicle according to the first embodiment.
[0019] Figure 3This is a structural diagram showing the structure of the object recognition ECU in the first embodiment.
[0020] Figure 4 This is a diagram illustrating an example of an intermediate image for object recognition in the first embodiment and the classification status of objects in that image.
[0021] Figure 5 This is an example of an image after object recognition in the first embodiment and the classification status of the objects in that image, and is related to... Figure 4 The corresponding examples are illustrated in the diagram.
[0022] Figure 6 This is a diagram schematically illustrating an example of the wavelength characteristics of the transmittance of a color filter that can be used in an external camera according to the first embodiment.
[0023] Figure 7 This is a diagram that schematically illustrates another example of the wavelength characteristics of the transmittance of a color filter that can be used in an external camera according to the first embodiment.
[0024] Figure 8 This is a diagram schematically representing the wavelength characteristics of the reflectance of four plant species.
[0025] Figure 9 This is a flowchart for explaining the processing of the object recognition ECU in the first embodiment.
[0026] Figure 10 This is a structural diagram showing the structure of the object recognition ECU in the second embodiment.
[0027] Figure 11 This is a flowchart for explaining the processing of the object recognition ECU in the second embodiment. Detailed Implementation
[0028] Hereinafter, several embodiments will be described based on the accompanying drawings. Furthermore, sometimes repeated descriptions are omitted by using the same reference numerals for corresponding components in each embodiment. Where only a portion of the structure is described in each embodiment, the structures of other embodiments described above can be applied to the other parts of that structure. In addition, not only combinations of structures explicitly shown in the descriptions of each embodiment, but also structures of multiple embodiments can be partially combined with each other, even if not explicitly shown, as long as there is no particular obstacle to the combination.
[0029] (First Implementation)
[0030] like Figure 1As shown, the object recognition device of the first embodiment of this disclosure is used for object recognition around the vehicle 1, and is configured as an object recognition ECU (Electronic Control Unit) 30 mounted on the vehicle 1. The object recognition ECU 30, together with the ranging sensor 10 and the external camera 20, constitutes an object recognition system 100. The object recognition system 100 of this embodiment can recognize objects based on the image information of the ranging sensor 10 and the image information of the external camera 20, and provide the object recognition result to the driving assistance ECU 50 and the like.
[0031] The object recognition ECU 30 is communicatively connected to the communication bus of the vehicle network installed in vehicle 1. The object recognition ECU 30 is one of the nodes set in the vehicle network. In addition to the ranging sensor 10 and the external camera 20, the driver assistance ECU 50 and others are also connected to the vehicle network as nodes.
[0032] The driver assistance ECU 50 is a structure primarily comprising a computer, including a processor, RAM (Random Access Memory), a storage unit, input / output interfaces, and a bus connecting them. The driver assistance ECU 50 has at least one of two functions: a driver assistance function that assists the driver in driving operations within the vehicle 1, and a driver assistance function that can perform driving operations on behalf of the driver. The driver assistance ECU 50 executes a program stored in the storage unit via the processor. Thus, the driver assistance ECU 50 enables autonomous driving or advanced driver assistance for the vehicle 1, corresponding to the object recognition results of the vehicle 1's surroundings provided by the object recognition system 100. For example, in implementing autonomous driving or advanced driver assistance for the vehicle 1 corresponding to the object recognition results, collision avoidance is prioritized for objects identified as pedestrians or other vehicles, compared to objects identified as plants.
[0033] Next, the details of the ranging sensor 10, the external camera 20, and the object recognition ECU 30 included in the object recognition system 100 will be described in turn.
[0034] The ranging sensor 10 is, for example, a SPADRiDAR (Single Photon Avalanche Diode Light Detection and Ranging) disposed in front of the vehicle 1 or on the roof of the vehicle 1. The ranging sensor 10 is capable of measuring at least a measuring range MA1 in front of the vehicle 1.
[0035] The ranging sensor 10 includes a light-emitting unit 11, a light-receiving unit 12, a control unit 13, etc. The light-emitting unit 11 scans the light beam emitted from the light source using a movable optical component (e.g., a multifaceted mirror) to move the light towards the light source. Figure 2 The measurement range MA1 is shown to be illuminated. The light source is, for example, a semiconductor laser (laser diode), which emits a beam of light in the near-infrared region that cannot be visually detected by the occupants (driver, etc.) and outsiders, based on an electrical signal from the control unit 13.
[0036] The light-receiving part 12, for example, uses a condenser lens to focus the reflected light or the background light relative to the reflected light reflected by an object within the measurement range MA1 of the irradiated light beam, and directs it toward the light-receiving element 12a.
[0037] The light-receiving element 12a is a device that converts light into an electrical signal through photoelectric conversion, and it is a SPAD light-receiving element that achieves high sensitivity by amplifying the detection voltage. In the light-receiving element 12a, for example, to detect reflected light in the near-infrared region, a CMOS sensor with a high sensitivity set relative to the visible area in the near-infrared region is used. This sensitivity can also be adjusted by providing an optical filter in the light-receiving section 12. The light-receiving element 12a has multiple light-receiving pixels arranged in an array in a one-dimensional or two-dimensional direction.
[0038] The control unit 13 is a unit that controls the light-emitting part 11 and the light-receiving part 12. The control unit 13 is disposed on a substrate common to the light-receiving element 12a, and is configured primarily as a processor, such as a microcomputer or a FPGA (Field-Programmable Gate Array). The control unit 13 performs scanning control functions, reflected light measurement functions, and background light measurement functions.
[0039] The scanning control function is the function of controlling the scanning of the light beam. The control unit 13, based on the timing of the operating clock of the clock oscillator set on the ranging sensor 10, causes the light beam to oscillate in a pulse pattern multiple times from the light source and causes the movable optical components to move.
[0040] The reflected light measurement function is a function that matches the timing of beam scanning, for example, by using a rolling shutter method to read out the voltage value of the reflected light received by each light-receiving pixel, and then measuring the intensity of the reflected light. In the reflected light measurement, the distance from the distance sensor 10 to the object reflecting the reflected light can be determined by detecting the time difference between the timing of the beam's emission and the timing of the reflected light's reception. Through the reflected light measurement, the control unit 13 can generate a reflected light image, which is image-like data that correlates the intensity of the reflected light, the distance information of the object reflecting the reflected light, and the two-dimensional coordinates corresponding to the measurement range MA1.
[0041] The background light measurement function is a function that reads the voltage value of the background light received by each light-receiving pixel and measures the intensity of the background light immediately before measuring the reflected light. Here, background light refers to incident light that is incident on the light-receiving element 12a from the measurement range MA1 outside the environment and does not substantially contain reflected light. Incident light includes natural light, display light emitted from external displays, etc. By measuring the background light, the control unit 13 can generate a background light image, which is image-like data relating the intensity of the background light to the two-dimensional coordinates corresponding to the measurement range MA1.
[0042] The reflected light image and the background light image are sensed by a common light-receiving element 12a and acquired from a common optical system including the light-receiving element 12a. Therefore, the reflected light image and the background light image are images based primarily on the measurement results in a common wavelength domain, namely the near-infrared region. Moreover, the coordinate systems of the reflected light image and the background light image can be considered as mutually consistent coordinate systems. Furthermore, it can be said that there is almost no deviation in measurement timing between the reflected light image and the background light image (e.g., less than 1 ns). Therefore, the reflected light image and the background light image can also be considered as being acquired synchronously.
[0043] For example, in this embodiment, image data that stores three channels of data corresponding to each pixel, including the intensity of reflected light, the distance to the object, and the intensity of background light, is sequentially output to the object recognition ECU30 as a sensor image.
[0044] The external camera 20 is, for example, a camera disposed inside the passenger compartment of the windshield of the vehicle 1. The external camera 20 is capable of measuring at least a measurement range MA2 in front of the vehicle 1 in the outside world, and more specifically, a measurement range MA2 that at least partially overlaps with the measurement range MA1 of the ranging sensor 10.
[0045] The external camera 20 is a structure that includes a light-receiving unit 22 and a control unit 23. The light-receiving unit 22, for example, uses a light-receiving lens to focus the incident light (background light) that is incident from the measurement range MA2 outside the camera and directs it toward the camera element 22a.
[0046] Camera element 22a is a device that converts light into electrical signals through photoelectric conversion, and can be, for example, a CCD sensor or a CMOS sensor. In camera element 22a, the sensitivity of the visible area is set high relative to the near-infrared area in order to efficiently receive natural light from the visible area. Camera element 22a has multiple light-receiving pixels (equivalent to so-called sub-pixels) arranged in an array in a two-dimensional direction. Color filters of different colors, such as red, green, and blue, are arranged among adjacent light-receiving pixels. These color filters adjust the wavelength characteristics of the sensitivity of the camera element 22a as a whole and of each light-receiving pixel. Each light-receiving pixel receives visible light of the color corresponding to the arranged color filter. By independently measuring the intensity of red light, green light, and blue light, the camera image captured by the external camera 20 is a high-resolution image compared to reflected light images and background light images, and can become a color image of the visible area.
[0047] The control unit 23 is a unit that controls the light-receiving part 22. The control unit 23 is, for example, disposed on a substrate shared with the camera element 22a, and is primarily configured as a processor such as a microcomputer or FPGA. The control unit 23 performs the shooting function.
[0048] The shooting function is the function of capturing the aforementioned color images. The control unit 23, based on the timing of the clock oscillator of the external camera 20, reads the voltage value of the incident light received by each light-receiving pixel using a global shutter method, and senses and measures the intensity of the incident light. The control unit 23 can generate a camera image, which is image-like data relating the intensity of the incident light to the two-dimensional coordinates corresponding to the measurement range MA2. Such camera images are sequentially output to the object recognition ECU 30.
[0049] The object recognition ECU 30 is an electronic control device that uses image information from the ranging sensor 10 and the external camera 20 to identify objects around the vehicle 1. For example... Figure 1As shown, the object recognition ECU 30 is structured primarily as a computer, including a processing unit 31, RAM 32, storage unit 33, input / output interface 34, and a bus connecting them. The processing unit 31 is hardware for computational processing in conjunction with RAM 32. The processing unit 31 is a structure that includes at least one processing core such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or RISC (Reduced Instruction Set Computer). The processing unit 31 may also include an FPGA and IP cores with other dedicated functions. The processing unit 31 executes various processes to implement the functions of the functional units described later by accessing RAM 32. The storage unit 33 is a structure that includes non-volatile storage media. Various programs (image registration programs, object recognition programs, etc.) executed by the processing unit 31 are stored in the storage unit 33.
[0050] The object recognition ECU 30 has multiple functional units by executing a program stored in the storage unit 33 through the processing unit 31. Specifically, such as... Figure 3 As shown, functional units such as image information acquisition unit 41, image registration unit 42, and region recognition unit 43 are constructed.
[0051] The image information acquisition unit 41 acquires image information including reflected light image and background light image from the ranging sensor 10, and sequentially acquires image information including camera image from the external camera 20. As image information, the image information acquisition unit 41 also acquires information related to the wavelength characteristics of the sensitivity of the camera element 22a (hereinafter, wavelength characteristic information). The wavelength characteristic information does not need to be acquired from the camera image every time; for example, it can be stored in the storage unit 33 during initial setup or when acquired from the external camera 20, and then retrieved by accessing the storage unit 33. The image information acquisition unit 41 provides the latest set of reflected light image, background light image, and camera image to the image registration unit 42.
[0052] The image registration unit 42 performs image registration by determining the correspondence between the coordinate systems of the reflected light image, the background light image, and the camera image. Furthermore, the image registration unit 42 provides the image-registered reflected light image, background light image, and camera image to the region recognition unit 43.
[0053] The region recognition unit 43 identifies a spatial region containing objects surrounding the vehicle 1. In particular, the region recognition unit 43 can identify plants reflected in the camera image by considering the difference between the intensity of the camera image and the intensity of the background light image. The region recognition unit 43 implements visible object recognition, surrounding space map generation, and plant identification functions.
[0054] The visible object recognition function identifies objects at the pixel level using semantic segmentation from camera images that contain visible light intensity distribution information. In the storage unit 33, an object recognition model 44, primarily based on a neural network, is constructed as a component of the object recognition program. The object recognition model 44 is a trained AI (artificial intelligence) model; if input from a camera image, it outputs an intermediate object recognition image Im1 (see reference). Figure 4 ).
[0055] In detail, neural networks can employ fully convolutional neural networks (FCNs), encoder / decoder networks that combine encoders and decoders (such as SegNet, U-Net, PSPNet), and so on.
[0056] As an example, let's illustrate the case of an encoder / decoder network. The encoder performs convolution and pooling operations on the input camera image. The encoder downsamples the camera image and extracts its features. The encoder outputs, for example, a feature map or the probability of classifying an object to the decoder.
[0057] The decoder uses the input data from the encoder to perform inverse pooling and deconvolution operations. The decoder can then output the object recognition mid-process image Im1 (see reference) by upsampling the camera image, which has been downsampled by the encoder. Figure 4 In the object recognition process, during image Im1, classification information is added to each region of the camera image, i.e., each region or pixel unit distinguished according to the object being represented. Furthermore, in... Figure 4 The text only indicates the classification status of a portion of the plant classification information.
[0058] In the first embodiment, classes are defined, such as pedestrians, other vehicles, and roads, which are highly necessary for classification during the movement of vehicle 1, but no class equivalent to plants such as grass and trees is defined. In semantic segmentation, plants are classified together with other objects that are difficult to determine and objects with low classification necessity, for example, as "stationary objects (StO)" (see reference). Figure 4 ")" and "obstacles" are higher-level conceptual classes.
[0059] Initially, in semantic segmentation of color images (camera images) using visible light, it was difficult to distinguish plants with high accuracy. For example, green plants could essentially be distinguished based on the intensity of green light in the color image. Therefore, it was easy to misclassify green plants as objects painted green or green lights. To ensure the accuracy of the output of the object recognition model 44, i.e., to suppress misclassification, green plants, objects painted green, and green lights were classified into a common class.
[0060] The learning method for the object recognition model 44 in the first embodiment can, for example, employ supervised learning. The teacher data in this machine learning consists of camera images as input data and the object recognition completed image Im2 as output data (positive resolution data). Figure 5 The dataset.
[0061] The surrounding space map generation function is a function that inputs reflected light images and background light images into the OGM generator 45 and generates a map in the surrounding space of the vehicle 1 with additional information on the occupancy status of objects. The area recognition unit 43 generates a map of the surrounding space (hereinafter referred to as OGM) for example using an OGM (Occupancy Grid Map) generation method. This OGM is preferably represented in a coordinate system viewed from above the vehicle 1.
[0062] Specifically, in the OGM, the area from the position of the ranging sensor 10 to the position of the object reflected in the reflected light image, determined based on distance information, is designated as a non-occupied area, not occupied by the object. The position of the object reflected in the reflected light image is designated as an occupied area, occupied by the object. Areas further inland than the position of the object reflected in the reflected light image are designated as undetermined areas, where occupancy cannot be determined.
[0063] The plant discrimination function identifies objects in the intermediate image Im1 by adding the aforementioned common class information to the object recognition process, such as the discrimination object regions TA1 to TA3. Figure 5 The region identification unit 43 classifies the region into plant regions (PLTs) that reflect plants and non-plant regions (NPTs) that reflect non-plants, as shown. The region identification unit 43 identifies the plant region (PLT) by considering at least one of the differences between the intensity of the camera image and the intensity of the reflected light image, and the differences between the intensity of the camera image and the intensity of the background light image. Specifically, in this embodiment, the plant region (PLT) is identified by considering the difference between the intensity of the camera image and the intensity of the background light image.
[0064] The region identification unit 43 uses plant discrimination formula 46 to calculate a discrimination value based on the following plant discrimination formula 46 for each discrimination target region TA1 to TA3, based on the intensity of corresponding pixels in the camera image (which establishes a correspondence by using coordinates representing the same position through image registration) and the intensity of corresponding pixels in the background light image. The region identification unit 43 changes the method of plant discrimination formula 46 according to the wavelength characteristics of the sensitivity of the camera element 22a obtained by the image information acquisition unit 41.
[0065] exist Figure 6 When the red color filter in the light-receiving pixel of the camera element 22a, as shown, substantially blocks near-infrared light, one of the plant discrimination formulas 46 is selected. This formula is represented by Q = Ib / (1-Ic). Here, Q is the discrimination value, Ib is the intensity of the corresponding pixel in the background light image, and Ic is the intensity of the corresponding pixel in the camera image, which is the sum of the intensities of red light, green light, and blue light.
[0066] exist Figure 7 In the case where the red filter in the light-receiving pixel of the camera element 22a, as shown, partially transmits near-infrared light, i.e., without sufficiently blocking near-infrared light, another method in plant discrimination formula 46 is selected. This formula is represented by Q = Ib / (1-Igb). Here, Igb is the intensity of the corresponding pixel in the camera image and is the sum of the intensities of green and blue light. Furthermore, when the camera image that becomes the object of semantic segmentation consists of two channels, an R image (an image based on the intensity of red) and a GB image (an image based on the intensity of green and blue), Igb can also be the intensity of the corresponding pixel in the camera image and the intensity of the GB image.
[0067] Furthermore, Ic and Igb are intensities normalized to a range of values greater than 0 and less than 1. In the normalized intensity, 0 means the minimum intensity (camera element 22a does not sense incident light), and 1 means the maximum intensity (camera element 22a senses the maximum amount of incident light). Additionally, Ib can also be normalized to a range of values greater than 0 and less than 1, similar to Ic and Igb.
[0068] The discriminant value Q is a discriminant value that reflects the difference between the intensity of the camera image and the intensity of the background light image. It is calculated using the plant discriminant formula 46, which is based on the intensity ratio of the camera image to the background light image.
[0069] An example of plant identification is illustrated. When the average calculated discrimination value Q in a target area is greater than or equal to a benchmark value indicating a probability (e.g., 50%) of the presence of a plant, the area identification unit 43 identifies the target area (referring to...) Figure 5TA2 and TA3 are set as plant regions PLT. When the average value is less than the baseline value, the region identification unit 43 will identify the target region (refer to...). Figure 5 TA1) is set as a non-plant region NPT.
[0070] That is, such as Figure 8 As shown, in the spectral characteristics of plants, there is a tendency for the reflectance of the near-infrared region RNi to increase significantly relative to the reflectance of the visible region RV. On the other hand, in the spectral characteristics of objects colored green by paint or other coatings, or green lights, there is a tendency for the reflectance of the visible region RV and the reflectance of the near-infrared region RNi to remain almost unchanged. Utilizing this difference in tendency, it is possible to classify plant regions (PLT) and non-plant regions (NPT). Specifically, in plants where the reflectance of the near-infrared region RNi is significantly increased relative to the reflectance of the visible region RV, the intensity of the background light image perceiving infrared light is significantly increased relative to the intensity of the camera image perceiving visible light. Therefore, when the object projected into the camera image is a plant, the discrimination value Q becomes a significantly high value based on the ratio of the reflectance of the visible region RV to the reflectance of the near-infrared region RNi.
[0071] Thus, the region recognition unit 43 outputs the object recognition completed image Im2, which includes the discrimination information of the plant region PLT in the intermediate object recognition image Im1. Furthermore, the region recognition unit 43 can also reflect the result of the discrimination of the plant region PLT in the OGM. For example, the region recognition unit 43 can set the occupied area on the OGM corresponding to the plant region PLT in the camera image as an intrusive area that the vehicle 1 can enter, and set the occupied area on the OGM corresponding to the non-plant region NPT in the camera image as an intrusive area that the vehicle 1 cannot enter.
[0072] Based on the above, the region recognition unit 43, through the visible object recognition function, classifies objects reflected in the camera image into multiple categories based on the camera image. In the classification into multiple categories, plants are grouped with other objects into a common category that is conceptualized as higher than plants. Then, the region recognition unit 43, through the plant discrimination function, considers the difference between the intensity of the camera image and the intensity of the background light image, and determines whether the object included in this common category is a plant. The resulting object recognition information, including the object recognition images Im2 and OGM, is then provided to the driver assistance ECU 50.
[0073] Next, use Figure 9 The flowchart illustrates an object recognition method for recognizing objects around vehicle 1 based on the object recognition program of the first embodiment. The processing of this flowchart, consisting of each step, is repeated, for example, at predetermined time intervals.
[0074] First, in S11, the image information acquisition unit 41 acquires image information including reflected light image and background light image from the ranging sensor 10, and acquires image information including camera image from the external camera 20. After processing in S11, the process is transferred to S12.
[0075] In S12, the image registration unit 42 performs image registration between the reflected light image and the background light image and the camera image. After processing in S12, the process moves to S13.
[0076] In S13, the region recognition unit 43 performs semantic segmentation on the camera image. Green plants, objects colored green by paint or other coatings, green lights, etc., are classified into a common category. Simultaneously or before and after this, the region recognition unit 43 generates an OGM based on the reflected light image and the background light image. After processing in S13, the process moves to S14.
[0077] In S14, the region recognition unit 43 determines the discrimination object regions from the camera image that have been supplemented with information about a common class. These discrimination object regions can be zero, one, or multiple. After processing in S14, the process moves to S15.
[0078] In S15, the region identification unit 43 uses the plant discriminant formula 46 to determine whether each target region is a plant region (PLT) or a non-plant region (NPT). S15 concludes the series of processes.
[0079] Furthermore, in the first embodiment, the area identification unit 43 is equivalent to an "identification unit" for identifying plants around the vehicle 1.
[0080] (Effects)
[0081] The effects of the first embodiment described above will be explained again below.
[0082] According to the first embodiment, in the identification of the plant region PLT that reflects the plant, the difference between the intensity of the camera image and the intensity of the sensor image is taken into account. That is, the difference between the information of perceived visible light in the camera image of the external camera 20 and the information of perceived near-infrared light in the image of the ranging sensor 10 is considered. As a result, it is easy to distinguish plants with significantly different spectroscopic characteristics in the visible region RV and the near-infrared region RNi, as well as other objects that should be avoided, which do not have the same degree of difference between the visible region RV and the near-infrared region RNi. Therefore, it is possible to provide an object recognition ECU 30 as an object recognition device capable of identifying plants with high accuracy.
[0083] Furthermore, according to the first embodiment, objects reflected in the camera image are classified into multiple categories based on the camera image. In this classification based on the camera image generated by sensing visible light by the camera element 22a, misclassification is suppressed because the classification is performed to include plants and other objects in a common category that is conceptualized higher than plants. Then, considering the difference between the intensity of the camera image and the intensity of the sensor image, it is determined whether the object included in the common category is a plant. That is, misclassification is suppressed, and objects that can be classified are classified, and then plants are accurately identified from objects included in the common category. Therefore, by suppressing the application of near-infrared light information sensed by the ranging sensor 10 to all objects, the burden on the processing volume or processing speed for object recognition can be reduced, and plants can be identified with high accuracy.
[0084] Especially in situations where real-time object recognition is required, such as when vehicle 1 is in motion, it is extremely useful for object recognition that simultaneously reduces the burden on processing volume or processing speed and achieves high-precision identification of plants.
[0085] Furthermore, according to the first embodiment, classification into multiple classes is performed through semantic segmentation in the object recognition model 44 with a neural network. Through semantic segmentation, objects can be classified according to each pixel of the camera image, thus enabling high-precision identification of the spatial region where the object is located.
[0086] Furthermore, according to the first embodiment, plants are identified based on a discriminant value calculated using a plant discriminant formula 46 that is based on the intensity ratio of the camera image to the sensor image. In this plant discriminant formula 46, the spectroscopic characteristics of the plant are clearly represented, making the presence of the plant obvious, thus significantly improving the accuracy of plant identification.
[0087] Furthermore, according to the first embodiment, the plant discrimination formula 46 is modified based on the wavelength characteristic information of the sensitivity of the camera element 22a. Therefore, the discrimination value calculated by the plant discrimination formula 46 can be modified to a discrimination value that more significantly represents the spectroscopic characteristic tendency of the plant, in a way that matches the wavelength characteristic information of the sensitivity of the camera element 22a, such as the sensitivity of the near-infrared region RNi.
[0088] (Second Implementation)
[0089] like Figure 10 , 11 As shown, the second embodiment is a variation of the first embodiment. The second embodiment will be described focusing on its differences from the first embodiment.
[0090] Figure 10The region identification unit 243 shown in the second embodiment has a plant identification image acquisition function and an object recognition function.
[0091] The plant identification image acquisition function is a function that acquires plant identification images. The region recognition unit 243 acquires the plant identification image by generating it using the same plant discrimination formula 46 as in the first embodiment. The region recognition unit 243, for example, prepares two-dimensional coordinate data in the same coordinate system as the camera image. For each coordinate in the two-dimensional coordinate data, the region recognition unit 243 calculates a discrimination value Q based on the plant discrimination formula 46, according to the intensity of the corresponding pixel in the camera image (which establishes a correspondence through image registration as coordinates representing the same position) and the intensity of the corresponding pixel in the background light image.
[0092] The result is the generation of a plant discrimination image that adds a discriminant value Q to each coordinate of the two-dimensional coordinate data. The plant discrimination image is image data in the same coordinate system as the camera image, and it visualizes the distribution of the discriminant values Q. In the plant discrimination image, based on the spectral characteristics of plants described in the first embodiment, the discriminant values Q of the portions reflecting plants tend to be higher; however, the discriminant values Q of the portions reflecting objects other than plants that have similar spectral characteristics also tend to be higher. Therefore, compared to the case where discrimination is performed solely using the plant discrimination image, the plant discrimination accuracy is significantly improved when the object recognition model 244 described later is also used.
[0093] The object recognition function uses camera images with visible light intensity distribution information and plant discrimination images with discriminant value distribution information based on visible and infrared light. It performs object recognition at the pixel level through semantic segmentation. In the storage unit 33, an object recognition model 244, primarily based on a neural network, is constructed as a component of the object recognition program. The object recognition model 244 in the second embodiment is a learned AI model. If a camera image and a plant discrimination image are input, it outputs an object recognition image Im2 with added class classification information.
[0094] In neural networks, a network with essentially the same structure as the first embodiment can be used. However, to correspond to the increase of input parameters such as the plant discrimination image, changes such as the need to increase the number of channels in the convolutional layers in the encoder are required.
[0095] In the second embodiment, a class equivalent to plants, such as grass and trees, is defined in addition to categories like pedestrians, other vehicles, and roads. In semantic segmentation, plants are classified as "plants," distinct from other objects that are difficult to identify or have low classification necessity. Therefore, the object recognition image output from the object recognition model 244 becomes an image in which plants have been identified with high precision.
[0096] The learning method for the object recognition model 244 in the second embodiment can, for example, employ supervised learning. The teacher data in this machine learning is a dataset consisting of camera images and plant discrimination images as input data, and object recognition completed images Im2 as output data (positive resolution data).
[0097] As described above, the region recognition unit 243 generates a plant discrimination image reflecting the differences between the intensity of the camera image and the intensity of the background light image using the plant discrimination image generation function. Then, the region recognition unit 243 classifies the objects reflected in the camera image into multiple classes based on the camera image and the plant discrimination image using the object recognition function. In the classification into multiple classes, plants are classified into a separate class.
[0098] Next, use Figure 11 The flowchart illustrates an object recognition method for recognizing objects around vehicle 1 based on the object recognition program of the second embodiment. For example, the process shown in the flowchart is repeated at predetermined time intervals.
[0099] First, S21 to S22 are the same as S11 to S12 in the first embodiment. After processing in S12, proceed to S23.
[0100] In S23, the region recognition unit 243 generates a plant discrimination image based on the camera image and the background light image, using the plant discrimination formula 46. After processing in S23, the process is transferred to S24.
[0101] In S24, the region recognition unit 243 inputs the camera image and the plant discrimination image into the object recognition model 244 to perform semantic segmentation. The green plants are classified into a different class from objects colored green by paint or other coatings, green lights, etc. The series of processes concludes in S24.
[0102] According to the second embodiment described above, a plant identification image reflecting the difference between the intensity of the camera image and the intensity of the sensor image is obtained based on the camera image and the sensor image. Furthermore, based on the camera image and the plant identification image, objects reflected in the camera image are classified into multiple categories. In this classification, plants are classified into a separate category. In this process, plants can be identified by using a comprehensive judgment based on both the camera image and the plant identification image, thus improving the accuracy of plant identification.
[0103] Furthermore, in the second embodiment, the area identification unit 243 is equivalent to an "identification unit" that identifies the plants around the vehicle 1.
[0104] (Other implementation methods)
[0105] The above describes several embodiments, but this disclosure should not be construed as being limited to these embodiments. Various embodiments and combinations can be applied without departing from the spirit of this disclosure.
[0106] Specifically, as a variation 1, the region recognition units 43 and 243 can also identify plants by considering the difference between the intensity of the camera image and the intensity of the reflected light image. That is, Ib in the plant discrimination formula 46 can also be the intensity of the reflected light image. Moreover, the region recognition unit 43 can also generate a composite image that incorporates both the intensity of the reflected light image and the intensity of the background light image, and identify plants by considering the difference between the intensity of the camera image and the intensity of this composite image.
[0107] As a variation 2, the region recognition units 43 and 243 may not use the discrimination value Q calculated by the plant discrimination formula 46, as a physical quantity reflecting the difference between the intensity of the camera image and the intensity of the sensor image. For example, the physical quantity reflecting the difference between the intensity of the camera image and the intensity of the sensor image may also be the difference between the intensity of the camera image and the intensity of the sensor image. In this case, the plant discrimination image used in the second embodiment may also be a difference image obtained based on the difference between the intensity of the camera image and the intensity of the sensor image.
[0108] As a variation 3 related to the second embodiment, the plant identification image may not be generated by the region recognition unit 243. For example, the plant identification image may be generated when the image registration unit 42 performs image registration. The region recognition unit 243 may also obtain the plant identification image by providing it from the image registration unit 42.
[0109] As a variation 4 related to the second embodiment, the data input to the object recognition model 244 may not be a plant discrimination image set separately from the camera image. For example, an additional channel different from RGB may be prepared in the camera image, and the calculated discrimination value Q may be stored in this channel. Moreover, the data input to the object recognition model 244 may be a camera image with the discrimination value Q appended.
[0110] As a variation 5, the region recognition units 43 and 243 can also use object recognition based on bounding boxes instead of semantic segmentation, simply by considering the difference between the intensity of the camera image and the intensity of the sensor image to identify plants.
[0111] As a variation 6, the recognition unit only needs to consider the difference between the intensity of the camera image and the intensity of the sensor image to identify plants, and it may not need to identify the spatial region containing objects surrounding the vehicle 1, which is the area recognition unit 43, 243. For example, the recognition unit may also not strictly associate plants with spatial regions for recognition, but determine whether plants are reflected in the measurement range MA2 of the camera image.
[0112] As a variation 7, the camera image may also be a grayscale image instead of a color image.
[0113] As a variation 8, at least one of the object recognition ECU 30 and the driving assistance ECU 50 may not be mounted on the vehicle 1, but may be fixedly installed on the road outside the vehicle 1, or may be mounted on other vehicles. In this case, object recognition processing, driving operation, etc., can also be remotely operated through communication such as networks, road-to-road communication, and vehicle-to-vehicle communication.
[0114] As a variation 9, the object recognition ECU 30 and the driving assistance ECU 50 can also be combined into one, for example, to form an electronic control device that implements a composite function. Furthermore, the ranging sensor 10 and the external camera 20 can also be configured as an integrated sensor unit. Moreover, an object recognition device such as the object recognition ECU 30 of the first embodiment can also be included as a structural element of this sensor unit.
[0115] As a variation 10, the object recognition ECU 30 may also omit the image registration unit 42. The object recognition ECU 30 may also acquire image information including the image of the reflected light, the background light, and the camera image after image registration is completed.
[0116] As a variation 11, the functions provided by the object recognition ECU 30 can also be provided by software and hardware executing the software, software only, hardware only, or a combination thereof. Furthermore, when such functions are provided by electronic circuits as hardware, each function can also be provided by digital circuits or analog circuits containing multiple logic circuits.
[0117] As a variation 12, the storage medium storing the object recognition program capable of implementing the above-described object recognition method can be appropriately modified. For example, the storage medium is not limited to a structure mounted on a circuit board, but can also be provided as a memory card or the like, inserted into a slot, and electrically connected to the control circuit of the object recognition ECU 30. Furthermore, the storage medium can also be an optical disc or a hard disk that serves as the basis for copying the program of the object recognition ECU 30.
[0118] The control unit and method described in this disclosure can also be implemented by a dedicated computer, which is configured as a processor programmed to perform one or more functions embodied in a computer program. Alternatively, the apparatus and method described in this disclosure can also be implemented by dedicated hardware logic circuitry. Alternatively, the apparatus and method described in this disclosure can also be implemented by one or more dedicated computers, which are configured by combining a processor that executes a computer program with one or more hardware logic circuits. Furthermore, the computer program can also be stored as instructions executed by a computer on a non-transferable tangible recording medium that can be read by a computer.
Claims
1. An object recognition device, wherein the object recognition device uses image information from an external camera mounted on a vehicle and image information from a ranging sensor mounted on the vehicle to identify objects around the vehicle, wherein the image information from the external camera is image information including a camera image generated by a camera element sensing visible light from the outside, and the image information from the ranging sensor is image information including a reflected light image and a background light image, wherein the reflected light image is generated by a light-receiving element sensing near-infrared reflected light reflected from an object by light illumination, and the background light image is generated by the light-receiving element sensing near-infrared background light relative to the reflected light, wherein... The object recognition device includes: The image information acquisition unit acquires image information from the ranging sensor and image information from the external camera; and The recognition unit identifies plants around the vehicle by considering the difference between the intensity of the camera image and the intensity of the sensor image. The identification unit identifies plants based on discrimination values that reflect the different physical quantities. These discrimination values are calculated using a plant discriminant formula based on the intensity ratio of the camera image and the sensor image. The image information acquisition unit acquires image information of the camera image, including wavelength characteristic information of the sensitivity of the camera element. The identification unit changes the plant discrimination formula based on the wavelength characteristic information.
2. The object recognition device according to claim 1, wherein, Based on the camera image, the recognition unit classifies the objects reflected in the camera image into multiple categories. Within these multiple categories, it further categorizes the objects so that plants are included alongside other objects in a common category that is conceptually higher than the plant. Then, the identification unit considers the differences to determine whether an object belonging to the common class is a plant.
3. The object recognition device according to claim 1, wherein, The identification unit acquires images reflecting the different plants based on the camera images and the sensor images. The recognition unit classifies objects appearing in the camera image into multiple categories based on the camera image and the plant discrimination image. Among the multiple categories, plants are classified into a separate category.
4. The object recognition device according to claim 2 or 3, wherein, The classification of the multiple classes is carried out through semantic segmentation in an object recognition model with a neural network.
5. The object recognition device according to claim 1, wherein, The recognition unit considers at least one of the differences between the intensity of the camera image and the intensity of the reflected light image, and the differences between the intensity of the camera image and the intensity of the background light image, as the difference between the intensity of the camera image and the intensity of the sensor image, to identify the plants around the vehicle.
6. A storage medium storing an object recognition program, the object recognition program using image information from an external camera mounted on a vehicle and image information from a ranging sensor mounted on the vehicle to identify objects around the vehicle, wherein the image information from the external camera is image information including a camera image generated by a camera element sensing visible light from the outside world, and the image information from the ranging sensor is image information including a reflected light image and a background light image, wherein the reflected light image is generated by a light-receiving element sensing near-infrared reflected light from an object under illumination, and the background light image is generated by the light-receiving element sensing near-infrared background light relative to the reflected light, wherein... The storage medium contains instructions executable by a computer, the instructions being used to cause at least one processing unit to perform the following processing: Processing of acquiring image information from the ranging sensor and image information from the external camera; and The process of identifying vegetation around the vehicle by considering the difference between the intensity of the camera image and the intensity of the sensor image. The identification process includes: identifying plants based on discriminant values that reflect the different physical quantities, the discriminant values being calculated using a plant discriminant formula based on the intensity ratio of the camera image and the sensor image. The process of acquiring the image information includes: acquiring image information of the camera image that includes wavelength characteristic information of the sensitivity of the camera element. The identification process includes: modifying the plant discrimination formula based on the wavelength characteristic information.
Citation Information
Patent Citations
Training of restricted deconvolution network for semantic segmentation of road scene
JP2017162456A
Medium conveying device
JP2019172398A
Deep convolutional neutral network and superpixel-based image semantic segmentation method
CN106709924A
Method for detecting a plant using an image sensor and method for controlling a vehicle assistance system
DE102016220560A1