Outside-vehicle environment recognition device

By mixing the ratios of stereo velocity, monocular velocity, and predicted velocity, the problem of unstable stereo velocity under raindrop interference is solved, thus achieving stability and security in vehicle environment recognition.

CN113392691BActive Publication Date: 2026-06-02SUBARU CORP

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUBARU CORP
Filing Date
2021-01-14
Publication Date
2026-06-02

Smart Images

  • Figure CN113392691B_ABST
    Figure CN113392691B_ABST
Patent Text Reader

Abstract

The present application provides a kind of outside environment recognition device, which stably exports the speed of stereoscopic object.The outside environment recognition device (120) has: stereoscopic speed export part (160), which exports the speed of predetermined stereoscopic object based on the distance image exported from the luminance image of two shooting devices (110) and is extracted;Monocular speed export part (162), which exports the speed of stereoscopic object based on the luminance image of one shooting device (110) and is extracted;Predicted speed export part (164), which exports the current speed of stereoscopic object based on the past speed of stereoscopic object and is predicted;Mixing ratio export part (166), which exports the mixing ratio of stereoscopic speed, monocular speed and predicted speed;And object speed export part (168), which exports the object speed representing the speed of stereoscopic object by mixing stereoscopic speed, monocular speed and predicted speed based on mixing ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a vehicle exterior environment recognition device capable of deriving the speed of a captured three-dimensional object. Background Technology

[0002] There are known technologies that use images captured by two cameras mounted on a vehicle to identify the speed and other external environment of three-dimensional objects (e.g., vehicles ahead) in the direction of travel of the vehicle (e.g., Patent Document 1).

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2019-34664 Summary of the Invention

[0006] Technical issues

[0007] In a technique that uses two imaging devices to identify the external environment of a vehicle, a distance image containing parallax information is derived from the brightness image captured by each imaging device, and the velocity of the stereoscopic object is derived based on the derived distance image.

[0008] However, if raindrops or similar substances are present in the detection area of ​​a camera, a large number of incorrect parallaxes may be derived in the distance image. If a large number of incorrect parallaxes are derived, the velocity of the stereoscopic object may sometimes be incorrect. Depending on the situation, the derived velocity of the stereoscopic object may sometimes change drastically. Thus, when using the derived velocity of the stereoscopic object for cruise control or similar purposes, the vehicle may be unexpectedly accelerated, or the vehicle may mistakenly approach a vehicle ahead.

[0009] Therefore, the purpose of this invention is to provide a vehicle external environment recognition device that can stably derive the speed of a three-dimensional object.

[0010] Technical solution

[0011] To address the aforementioned issues, the vehicle exterior environment recognition device of the present invention comprises: a stereo velocity derivation unit that derives a stereo velocity representing the velocity of a predetermined stereo object extracted based on a distance image derived from brightness images of two imaging devices; a monocular velocity derivation unit that derives a monocular velocity representing the velocity of a stereo object extracted based on a brightness image of one imaging device; a prediction velocity derivation unit that derives a prediction velocity representing the current velocity of a stereo object predicted based on its past velocity; a mixing ratio derivation unit that derives a mixing ratio of the stereo velocity, monocular velocity, and prediction velocity; and an object velocity derivation unit that, based on the mixing ratio, mixes the stereo velocity, monocular velocity, and prediction velocity to derive an object velocity representing the velocity of the stereo object.

[0012] In addition, the mixing ratio derivation unit may have: a stereo weight derivation unit that derives a stereo weight representing the proportion of stereo velocity; a monocular weight derivation unit that derives a monocular weight representing the proportion of monocular velocity; and a prediction weight derivation unit that derives a prediction weight representing the proportion of prediction velocity. When the stereo velocity is set to Vs, the monocular velocity is set to Vm, the prediction velocity is set to Vp, the stereo weight is set to Ws, the monocular weight is set to Wm, the prediction weight is set to Wp, and the object velocity is set to V, the object velocity derivation unit can derive the object velocity by the following formula (1).

[0013] V=(Ws×Vs+Wm×Vm+Wp×Vp) / (Ws+Wm+Wp)…(1)

[0014] In addition, the stereo weight derivation unit can derive stereo weights based on the confidence value of the distance image.

[0015] In addition, the monocular weight derivation unit can derive monocular weights based on the reliability value of a brightness image from an imaging device.

[0016] In addition, the prediction weight derivation unit can derive prediction weights based on the uncertainty value of the distance image and the uncertainty value of the brightness image of the imaging device.

[0017] Technical effect

[0018] According to the present invention, the speed of a three-dimensional object can be stably derived. Attached Figure Description

[0019] Figure 1 This is a block diagram illustrating the connection relationships of the vehicle external environment recognition system.

[0020] Figure 2 It is an explanatory diagram used to illustrate brightness and distance images. Figure 2 (A) shows an example of a brightness image. Figure 2 (B) shows an example of a distance image.

[0021] Figure 3 This is a functional block diagram illustrating the general functions of the vehicle external environment recognition device.

[0022] Figure 4 It is a diagram illustrating the three-dimensional confidence value. Figure 4 (A) shows the stereo states associated with stereo confidence values ​​in a table. Figure 4 (B) shows the three-dimensional state in a diagram.

[0023] Figure 5 This is a graph illustrating IB values. Figure 5 (A) shows an example of a brightness image. Figure 5(B) shows the basis for Figure 5 An example of an image obtained by processing a distance image derived from a brightness image of (A).

[0024] Figure 6 This is a diagram illustrating the confidence value of a single eye. Figure 6 (A) shows the vehicle's external environment during daytime. Figure 6 (B) shows the vehicle's exterior environment at night.

[0025] Figure 7 This is a graph illustrating the recognition rate and recognition score.

[0026] Figure 8 This is a diagram illustrating the scoring of the lights.

[0027] Figure 9 This is a flowchart illustrating the workflow of the central control department.

[0028] Symbol Explanation

[0029] 110 camera device

[0030] 120 Vehicle External Environment Recognition Device

[0031] 126 brightness image

[0032] 128 Distance Image

[0033] 160 Stereo Velocity Export Section

[0034] 162 Monocular Speed ​​Output Unit

[0035] 164 Predicted Velocity Derivation Section

[0036] 166 Mixing Ratio Derivation Section

[0037] 168 Object velocity derivation section

[0038] 170 Solid Weight Derivation Department

[0039] 172 Monocular Weight Derivation Unit

[0040] 174 Prediction Weight Derivation Section Detailed Implementation

[0041] Embodiments of the present invention will be described in detail below with reference to the accompanying drawings. The dimensions, materials, and other specific values ​​shown in these embodiments are merely examples for ease of understanding and are not intended to limit the invention unless specifically stated otherwise. It should be noted that in this specification and the accompanying drawings, elements with substantially the same function or structure are omitted from repeated description by using the same symbols; furthermore, elements not directly related to the present invention are omitted from illustration.

[0042] Figure 1 This is a block diagram showing the connection relationships of the vehicle external environment recognition system 100. The vehicle external environment recognition system 100 is configured to include a camera 110, a vehicle external environment recognition device 120, and a vehicle control unit (ECU: Electronic Control Unit, engine control unit) 130.

[0043] The imaging device 110 is configured to include imaging elements such as a CCD (Charge-Coupled Device) and / or a CMOS (Complementary Metal-Oxide Semiconductor), capable of capturing images of the external environment in front of the vehicle 1 and generating brightness images (color images and / or black-and-white images) that contain at least brightness information. Furthermore, the imaging devices 110 are arranged horizontally on the side of the vehicle 1's travel direction, with the optical axes of the two imaging devices 110 approximately parallel. The imaging device 110 continuously generates brightness images of three-dimensional objects present in the detection area in front of the vehicle 1 at, for example, 1 / 60th of a second (60fps). Here, the three-dimensional objects identified by the imaging device 110 include not only independently existing objects such as bicycles, pedestrians, vehicles, traffic lights, roads (roadways), road signs, guardrails, and buildings, but also objects that can be identified as part of an independently existing object, such as bicycle wheels.

[0044] Additionally, the vehicle exterior environment recognition device 120 obtains brightness images from the two imaging devices 110 respectively. It uses so-called pattern matching—retrieving from one brightness image a block corresponding to an arbitrary block extracted from the other brightness image (e.g., a horizontal array of 4 pixels × vertical array)—to derive disparity information, including disparity and image position indicating the location of any arbitrary block within the image. Here, horizontal represents the horizontal orientation of the captured image, and vertical represents the vertical orientation of the captured image. As for this pattern matching, it is possible to compare brightness (Y) between a pair of images in arbitrary block units. For example, methods include SAD (Sum of Absolute Difference) which takes the difference in brightness, SSD (Sum of Squared Intensity Difference) which takes the squared difference, and NCC (Normalized CrossCorrelation) which takes the variance obtained by subtracting the average value from the brightness of each pixel. The vehicle exterior environment recognition device 120 performs parallax export processing on all blocks appearing in the detection area (e.g., 600 pixels × 200 pixels) in such block units. Here, the block is set to 4 pixels × 4 pixels, but the number of pixels within the block can be arbitrarily set.

[0045] However, in the vehicle exterior environment recognition device 120, although it can derive the parallax for each block as a detection resolution unit, it cannot identify what kind of object that block is part of. Therefore, the parallax information is derived independently not in terms of object units, but in terms of detection resolution units (e.g., block units) at the detection area. Here, the image corresponding to the parallax information derived in this way is referred to as a distance image, distinguishing it from the brightness image described above.

[0046] Figure 2 This is an explanatory diagram used to illustrate the brightness image 126 and the distance image 128. Figure 2 (A) shows an example of a brightness image 126. Figure 2 (B) illustrates an example of a distance image 128. For example, it is assumed that an image 128 is generated for the detection region 124 using two imaging devices 110. Figure 2 Brightness image 126 as shown in (A). However, for ease of understanding, only one of the two brightness images 126 is schematically shown here. The vehicle exterior environment recognition device 120 calculates the parallax of each block based on such brightness image 126, forming... Figure 2 Distance image 128, as shown in (B). Each block in distance image 128 is associated with its disparity. Here, for clarity, black dots represent blocks from which the disparity is derived.

[0047] Furthermore, the vehicle exterior environment recognition device 120 converts the parallax information of each block in the distance image 128 into three-dimensional position information using a so-called stereo method, thereby enabling it to derive the relative distance to the vehicle 1 for each block. Here, the stereo method is a method of deriving the relative distance of an object to the imaging device 110 based on the parallax of the object using triangulation.

[0048] The external environment recognition device 120 uses brightness values ​​(color values) based on brightness image 126 and three-dimensional position information in actual space, including the relative distance to the vehicle 1, calculated based on distance image 128. It groups areas with equal color values ​​and similar three-dimensional position information as objects to determine which specific object (e.g., a vehicle in front, a bicycle) corresponds to an object in the detection area in front of the vehicle 1. Furthermore, if the external environment recognition device 120 determines a three-dimensional object in this way, it controls the vehicle 1 to avoid collisions with the object (collision avoidance control) or to maintain a safe distance from the vehicle in front (cruise control).

[0049] The external environment recognition device 120 outputs the velocity of the three-dimensional object as determined above, and uses the output velocity of the three-dimensional object to perform the aforementioned collision avoidance control or cruise control. In the following text, the velocity of the three-dimensional object output by the external environment recognition device 120 will sometimes be referred to as the object velocity.

[0050] like Figure 1 As shown, the vehicle control unit 130 receives the driver's input via the steering wheel 132, accelerator pedal 134, and brake pedal 136, and transmits it to the steering mechanism 142, drive mechanism 144, and braking mechanism 146, thereby controlling the vehicle 1. Additionally, the vehicle control unit 130 controls the steering mechanism 142, drive mechanism 144, and braking mechanism 146 according to the instructions of the external environment recognition device 120.

[0051] Figure 3 This is a functional block diagram illustrating the general functions of the vehicle exterior environment recognition device 120. For example... Figure 3 As shown, the vehicle external environment recognition device 120 is configured to include an I / F unit 150, a data holding unit 152, and a central control unit 154.

[0052] The I / F section 150 is an interface for bidirectional information exchange between the imaging device 110 and the vehicle control device 130. The data storage section 152, composed of RAM, flash memory, HDD, etc., stores various information required for the processing of each functional section as shown below.

[0053] The central control unit 154 is composed of semiconductor integrated circuits including a central processing unit (CPU), a ROM storing programs, and RAM serving as the working area. It controls the I / F unit 150, the data holding unit 152, etc., via the system bus 156. In addition, in this embodiment, the central control unit 154 also functions as a stereo velocity output unit 160, a monocular velocity output unit 162, a prediction velocity output unit 164, a mixing ratio output unit 166, and an object velocity output unit 168.

[0054] Thus, in the distance image, the image consistency between the brightness images 126 of the left and right imaging devices 110, derived for each block, is above the threshold and represents the maximum parallax. In other words, in the distance image, blocks with image consistency less than the threshold are considered as having no derived parallax.

[0055] With the left and right imaging devices 110 able to clearly capture the detection area 124, the consistency of the left and right brightness images 126 becomes higher, and parallax can actually be derived from a large number of blocks in the distance image 128.

[0056] However, if raindrops or similar objects are present in the detection area 124 of the imaging device 110, the brightness image 126 becomes a blurry image (the three-dimensional object is not clearly captured). Therefore, if raindrops are present in the detection area 124 of either of the left or right imaging devices 110, the consistency of the left and right brightness images 126 decreases, and the number of blocks in the distance image 128 from which parallax can actually be derived decreases. If the number of blocks from which parallax cannot be derived increases, it becomes difficult to determine the three-dimensional object based on the distance image 128, resulting in the inability to accurately derive the speed of objects such as moving vehicles.

[0057] Therefore, if the number of blocks in the distance image 128 from which parallax can actually be derived is less than the predetermined number of blocks, the vehicle exterior environment recognition device 120 switches from stereo control using stereo distance to monocular control using monocular distance. Stereo control is control that uses at least stereo distance to control the vehicle 1. Stereo distance is the relative distance of stereoscopic objects extracted based on the distance image 128. Monocular control is control that uses monocular distance instead of stereo distance to control the vehicle 1. Monocular distance is the relative distance of stereoscopic objects extracted based on the brightness image 126 of a (single-view) imaging device 110.

[0058] If in monocular control, the external environment recognition device 120 uses the brightness image 126 of the left and right shooting devices 110 that can take relatively clear pictures (e.g., the one with the higher overall brightness value of brightness image 126).

[0059] For example, the vehicle exterior environment recognition device 120 derives an object frame representing the area occupied by a stereoscopic object (e.g., a moving vehicle) in the brightness image 126 of the monocular controlled object, indicating the object's velocity. Then, the vehicle exterior environment recognition device 120 derives the object's velocity based on the difference between the lateral width of the object frame relative to a predetermined frame (e.g., 100 frames) before the current frame and the lateral width of the object frame at that time (the stereoscopic distance before the initial switch to monocular control is the distance before the switch), treating the change in lateral width as a change in relative distance.

[0060] Thus, in the vehicle external environment recognition device 120, when the brightness image 126 of one of the two imaging devices 110 is blurry, the object speed can be derived by using only the brightness image 126 of the other imaging device 110, thereby suppressing the reduction in the accuracy of the derived object speed.

[0061] However, before switching from stereoscopic control to monocular control, the number of blocks from which parallax can be derived is greater than the predetermined number of blocks. However, due to the presence of raindrops, etc., the brightness image 126 is blurred, so sometimes the parallax value is derived incorrectly for each block. If a large number of incorrect parallax values ​​are derived, the object velocity is sometimes derived incorrectly when deriving the object velocity based on the distance image 128.

[0062] If an incorrect object speed is derived, the external environment recognition device 120 may interpret it as a sudden change in object speed. For example, although a three-dimensional object (e.g., a vehicle in front) is actually moving at a constant speed of about 80 km / h, the external environment recognition device 120 may interpret it as the object rapidly accelerating from about 80 km / h to about 90 km / h. In such a case, if cruise control is performed based on the derived object speed, the vehicle 1 may accelerate unintentionally, and the vehicle 1 may mistakenly approach a vehicle in front that is actually moving at a constant speed.

[0063] Therefore, in stereoscopic control, the vehicle exterior environment recognition device 120 of this embodiment not only outputs the speed of the stereoscopic object extracted based on the distance image 128, but also outputs the speed of the stereoscopic object extracted based on the brightness image 126 of an imaging device 110, and the current speed of the stereoscopic object predicted based on the past speed of the stereoscopic object.

[0064] In the following text, the velocity of the stereoscopic object extracted based on the distance image 128 is sometimes referred to as stereoscopic velocity. Additionally, the velocity of the stereoscopic object extracted based on the brightness image 126 of an imaging device 110 is sometimes referred to as monocular velocity. Furthermore, the current velocity of the stereoscopic object predicted based on its past velocity is sometimes referred to as predicted velocity.

[0065] Additionally, the predicted velocity is, for example, the object velocity derived at the time of the last velocity derived. That is, for the predicted velocity, the past velocity is used as the current velocity.

[0066] It should be noted that if, in the past up to the predetermined time, the absolute value of the object's deceleration was set to a predetermined value (e.g., 0.1G) or higher, this deceleration can also be taken into account to derive the predicted velocity. That is, in this case, the predicted velocity can also be set to a value that is smaller than the object's velocity at the last derived time.

[0067] In this embodiment, the vehicle external environment recognition device 120, during stereo control, mixes stereo velocity, monocular velocity, and predicted velocity at an appropriate ratio to derive the object velocity of a stereoscopic object. In the following text, the process of mixing stereo velocity, monocular velocity, and predicted velocity to derive the object velocity is sometimes referred to as velocity correction processing.

[0068] The stereo velocity derivation unit 160 in the vehicle external environment recognition device 120 derives the stereo velocity. The monocular velocity derivation unit 162 derives the monocular velocity. The prediction velocity derivation unit 164 derives the prediction velocity. The mixing ratio derivation unit 166 derives the mixing ratio of the stereo velocity, monocular velocity, and prediction velocity. The object velocity derivation unit 168, based on the derived mixing ratio, mixes the stereo velocity, monocular velocity, and prediction velocity to derive the object velocity of the stereoscopic object.

[0069] More specifically, the mixing ratio derivation unit 166 includes a stereo weight derivation unit 170, a monocular weight derivation unit 172, and a prediction weight derivation unit 174. The stereo weight derivation unit 170 derives stereo weights representing the proportion of stereo velocity. The monocular weight derivation unit 172 derives monocular weights representing the proportion of monocular velocity. The prediction weight derivation unit 174 derives prediction weights representing the proportion of prediction velocity.

[0070] Here, the stereo velocity is sometimes denoted as Vs, the monocular velocity as Vm, the predicted velocity as Vp, the stereo weight as Ws, the monocular weight as Wm, the predicted weight as Wp, and the object velocity as V. More specifically, the object velocity derivation unit 168 derives the object velocity in the velocity correction process using the following equation (1).

[0071] V=(Ws×Vs+Wm×Vm+Wp×Vp) / (Ws+Wm+Wp)…(1)

[0072] The stereo weight derivation unit 170 derives stereo weights based on the confidence value of the distance image 128. Specifically, the stereo weight derivation unit 170 derives stereo weights using the following equation (2).

[0073] Ws = (3D confidence value) 2…(2)

[0074] Figure 4 It is a diagram illustrating the three-dimensional confidence value. Figure 4 (A) shows the stereo states associated with stereo confidence values ​​in a table. Figure 4 (B) shows the three-dimensional state in a diagram.

[0075] like Figure 4 As shown in (A), the stereo confidence value is set to four levels: "0", "1", "2", and "3". A higher stereo confidence value indicates greater reliability at a distance of 128 from the image. Furthermore, the stereo confidence value is associated with the stereo state.

[0076] The stereo state is an indicator of the reliability of a distance of 128 from the image. There are five levels of stereo state: "Super Trust," "Trust," "Stable," "Maybe," and "None." The stereo state "Trust" corresponds to a stereo confidence value of "3." The stereo state "Stable" corresponds to a stereo confidence value of "2." The stereo state "Maybe" corresponds to a stereo confidence value of "1." The stereo state "None" corresponds to a stereo confidence value of "0." That is, "Trust" has higher reliability than "Stable," "Stable" has higher reliability than "Maybe," and "Maybe" has higher reliability than "None."

[0077] Furthermore, the reliability setting for "Super Trust" is higher than that for "Trust". "Super Trust" does not have a corresponding stereo reliability value. However, as an exception, in the case of "Super Trust", a mixed ratio of 100% stereo velocity, 0% monocular velocity, and 0% prediction velocity is used. That is, as an exception, in the case of "Super Trust", the object velocity derivation unit 168 directly uses the stereo velocity of the stereo object as the object velocity.

[0078] Figure 4 The solid line 180 in (B) represents the threshold distinguishing between "super trust" and "trust". The single-dash line 182 represents the threshold distinguishing between "trust" and "safety". The double-dash line 184 represents the threshold distinguishing between "safety" and "possibility". The dashed line 186 represents the threshold distinguishing between "possibility" and "untrustworthy".

[0079] Stereoscopic state is determined based on monocular distance and IB value. Figure 5 This is a graph illustrating IB values. Figure 5 (A) shows an example of a brightness image 126. Figure 5 (B) shows the basis for Figure 5An example of an image obtained by processing the distance image 128 derived from the brightness image 126 of (A). Figure 5 Image (B) shows multiple pillars 190 extending along the vertical direction (height direction) of the image. These pillars 190 are formed by connecting consecutive blocks of approximately equal parallax (distance) from bottom to top as the block extends upwards from the bottom of the three-dimensional object. Furthermore, regarding pillars 190, blocks can be connected if the number of consecutive blocks with approximately equal parallax (distance) is greater than a predetermined number, and pillars 190 are not formed if the number of consecutive blocks is less than the predetermined number. That is, the area of ​​a pillar 190 represents the size of a region with approximately equal parallax (distance).

[0080] The IB value represents the number of columns in the 190° column. If there are more blocks in the distance image 128 that derive parallax, the IB value tends to increase. Therefore, it can be inferred that a larger IB value indicates higher reliability in the distance image 128. In other words, the IB value is an indicator of the detection accuracy of three-dimensional objects in the distance image 128.

[0081] like Figure 4 (A) and Figure 4 As shown in (B), the level of the stereo state is determined based on the IB value, and it is set that the larger the IB value, the greater the stereo confidence value.

[0082] For example, when the monocular distance is less than 50m, the following determinations are made: If the IB value exceeds 23, the stereo weighting derivation unit 170 sets the stereo state to "Super Trust". If the IB value exceeds 17 but is less than 23, the stereo weighting derivation unit 170 sets the stereo state to "Trust" and sets the stereo confidence value to "3". If the IB value exceeds 14 but is less than 17, the stereo weighting derivation unit 170 sets the stereo state to "Safe" and sets the stereo confidence value to "2". If the IB value exceeds 6 but is less than 14, the stereo weighting derivation unit 170 sets the stereo state to "Possible" and sets the stereo confidence value to "1". If the IB value is less than 6, the stereo weighting derivation unit 170 sets the stereo state to "Untrustworthy" and sets the stereo confidence value to "0".

[0083] Additionally, for example, when the monocular distance exceeds 70m, the following determinations are made: If the IB value exceeds 18, the stereo weighting derivation unit 170 sets the stereo state to "Super Trust". If the IB value exceeds 12 but is less than 18, the stereo weighting derivation unit 170 sets the stereo state to "Trust" and sets the stereo confidence value to "3". If the IB value exceeds 9 but is less than 12, the stereo weighting derivation unit 170 sets the stereo state to "Safe" and sets the stereo confidence value to "2". If the IB value exceeds 2 but is less than 9, the stereo weighting derivation unit 170 sets the stereo state to "Possible" and sets the stereo confidence value to "1". If the IB value is less than 2, the stereo weighting derivation unit 170 sets the stereo state to "Untrustworthy" and sets the stereo confidence value to "0".

[0084] Additionally, within a range where the distance between one eye exceeds 50m but is less than 70m, such as Figure 4 As shown in (B), the linear interpolation value obtained by linear interpolation between the threshold value of IB at 50m and the threshold value of IB at 70m is used as the threshold to distinguish the different levels of the three-dimensional state.

[0085] It should be noted that it is presumed that the closer the stereoscopic object, i.e., the closer the monocular distance, the more easily the parallax in the distance image 128 can be derived across a large number of blocks. That is, it is presumed that the closer the monocular distance, the easier it is for the IB value to increase. Therefore, as... Figure 4 (A) and Figure 4 As shown in (B), the threshold for the IB value is set such that the closer the distance between the two eyes, the larger the threshold for the IB value.

[0086] In addition, if the stereo velocity is less than 0.8 times the predicted velocity or more than 1.3 times the predicted velocity, the stereo weight derivation unit 170 can reduce the stereo state by one level and reduce the stereo confidence value by 1.

[0087] Next, the monocular weight will be explained. The monocular weight derivation unit 172 derives the monocular weight based on the confidence value of the brightness image 126 of an imaging device. Specifically, the monocular weight derivation unit 172 derives the monocular weight using the following formula (3).

[0088] Wm = (Mono-eye confidence value) 2 …(3)

[0089] Figure 6 This is a diagram illustrating the confidence value of a single eye. Figure 6 (A) shows the situation where the external environment of this vehicle 1 is daytime. Figure 6 (B) shows the external environment of this vehicle 1 at night.

[0090] like Figure 6 (A) and Figure 6As shown in (B), the monocular confidence value is set to four levels: "0", "1", "2", and "3". A higher monocular confidence value indicates greater reliability of the luminance image 126. Furthermore, the monocular confidence value is related to the monocular state.

[0091] The monocular state is an indicator of the reliability of a luminance image 126. The monocular state can be categorized into four levels: "Trust," "Safe," "Possible," and "Unreliable." The "Trust" state corresponds to a monocular confidence value of "3." The "Safe" state corresponds to a monocular confidence value of "2." The "Possible" state corresponds to a monocular confidence value of "1." The "Unreliable" state corresponds to a monocular confidence value of "0." In other words, "Trust" has higher reliability than "Safe," "Safe" has higher reliability than "Possible," and "Possible" has higher reliability than "Unreliable."

[0092] like Figure 6 As shown in (A), in the daytime situation, the monocular state is determined based on the recognition rate and recognition score. In contrast, as... Figure 6 As shown in (B), in the case of nighttime, the monocular state is determined based on the lamp score and the lamp detection duration flag.

[0093] Figure 7 This is a graph illustrating the recognition rate and recognition score. In the case of determining an object using a brightness image 126 from an imaging device 110, such as... Figure 7 As shown, at least a portion of an object is surrounded by a frame 200 of a predetermined size, and identification is performed to determine whether the object within the frame 200 is a three-dimensional object to be identified (e.g., a moving vehicle). In this identification, the frame 200 is randomly moved relative to the object to identify whether the object is a three-dimensional object to be identified in each of the multiple frames 200. It should be noted that the frame 200 can also be formed to match the size of the detected three-dimensional object.

[0094] In the recognition process within each bounding box 200, machine learning is used to determine the numerical similarity between the object and the identified solid object. The recognition score is the numerical similarity score. The recognition score is standardized to a value between 0 and 1. A score closer to 1 indicates a stronger resemblance to the identified solid object (e.g., a moving vehicle), while a score closer to 0 indicates a weaker resemblance. The average recognition score is derived by averaging the scores across the number of bounding boxes 200 (in other words, the number of times the recognition score is derived).

[0095] In addition, regarding each box 200, if the recognition score is higher than 0.51, the object within the box 200 is identified as a three-dimensional object that should be determined; if the recognition score is lower than 0.51, the object within the box 200 is identified as a three-dimensional object that should be determined.

[0096] The recognition rate represents the ratio of the number of times a solid is identified as a definite object to the number of frames 200 (in other words, the number of times it is identified). It can be inferred that a higher recognition rate indicates a higher probability of identifying a definite object.

[0097] like Figure 6 As shown in (A), during daytime conditions, the higher the recognition rate and the average recognition score, the higher the monocular confidence value. For example, when the recognition rate exceeds 0.74 and the average recognition score exceeds 0.70, the monocular weight derivation unit 172 sets the monocular state to "trust" and the monocular confidence value to "3". Furthermore, when the recognition rate exceeds 0.70 but is below 0.74, and the average recognition score exceeds 0.61 but is below 0.70, the monocular weight derivation unit 172 sets the monocular state to "safe" and the monocular confidence value to "2". Additionally, when the recognition rate exceeds 0.63 but is below 0.70, and the average recognition score exceeds 0.55 but is below 0.61, the monocular weight derivation unit 172 sets the monocular state to "possible" and the monocular confidence value to "1". In addition, when the recognition rate is below 0.63 and the average recognition score is below 0.55, the monocular weight derivation unit 172 sets the monocular state to "untrusted" and sets the monocular trust value to "0".

[0098] It should be noted that during the daytime, if a three-dimensional object cannot be recognized for a certain period of time (e.g., 0.5 seconds to 1.5 seconds), the monocular weight derivation unit 172 can reduce the monocular state to "untrustworthy" and reduce the monocular trust value to "0".

[0099] Furthermore, when the recognition rate and the average recognition score drop sharply, the loss count, which represents the number of times the recognition is considered unrecognizable, is counted. If the loss count is counted more than a predetermined number of times during the day, the monocular weight derivation unit 172 can reduce the monocular state by one level and reduce the monocular confidence value by 1.

[0100] Figure 8 This is a diagram illustrating the taillight score. At night, the left and right taillights of the vehicle in front are illuminated. The taillight score is a rating based on the characteristics of the left and right taillights of the vehicle in front.

[0101] Figure 8 Arrow 210 indicates the height position of the center of each taillight in brightness image 126. Figure 8 The shaded area 212 represents the area of ​​each taillight. Figure 8Arrow 214 indicates the lateral width of each taillight. The taillight score is calculated by adding the consistency of the center height, area, and lateral width of the taillights in the left and right taillight group. The taillight score is adjusted to a value between 0 and 100. For example, if the center height, area, and lateral width are all consistent for both left and right taillights, the taillight score is 100.

[0102] Additionally, the continuous taillight detection indicator indicates whether the taillights of the preceding vehicle are continuously detected. If taillights are detected for a predetermined number of consecutive frames (e.g., 20 frames), the continuous taillight detection indicator is activated and remains activated as long as taillights are continuously detected. It should be noted that the continuous taillight detection indicator is deactivated if taillight detection is interrupted.

[0103] like Figure 6 As shown in (B), at night, the higher the light score and the continuous light detection flag being on, the higher the monocular confidence value. For example, when the light score exceeds 80 and the continuous light detection flag is on, the monocular weight derivation unit 172 sets the monocular state to "trusted" and the monocular confidence value to "3". Furthermore, regardless of the continuous light detection flag, when the light score is between 70 and 80, the monocular weight derivation unit 172 sets the monocular state to "safe" and the monocular confidence value to "2". Additionally, regardless of the continuous light detection flag, when the light score is below 70, the monocular weight derivation unit 172 sets the monocular state to "untrusted" and the monocular confidence value to "0". The monocular state "possible" corresponds to a monocular confidence value of "1". For example, the monocular state "possible" can be achieved when the monocular confidence value decreases by 1 from "2" or increases by 1 from "0".

[0104] It should be noted that when the monocular distance to a three-dimensional object (a vehicle in front) exceeds 40m at night, the monocular weight derivation unit 172 can reduce the monocular state by one level and decrease the monocular confidence value by 1. This is to suppress the monocular weight from becoming erroneously too high due to the light score of a taillight located in the distance.

[0105] Alternatively, whether during the day or at night, if the monocular distance to the stereoscopic object exceeds 60m, the monocular weight derivation unit 172 can reduce the monocular state by one level and decrease the monocular confidence value by 1. It should be noted that at night, the monocular confidence value can also be reduced by a total of 2 using the aforementioned conditions of 40m and 60m.

[0106] Alternatively, in cases where the change in monocular velocity of a stereoscopic object indicates acceleration during both day and night, the monocular weight derivation unit 172 may reduce the monocular state by one level and decrease the monocular confidence value by 1. This is to reduce the monocular confidence value in order to account for the possibility of incorrectly judging that the stereoscopic object is accelerating, thereby suppressing erroneous acceleration of the vehicle based on the increase in monocular velocity.

[0107] Alternatively, during both day and night, if the change in the monocular speed of the stereoscopic object indicates acceleration, and the brake indicator of the brake light recognition module is activated, the monocular weight derivation unit 172 will reduce the monocular state by one level and decrease the monocular confidence value by 1. It should be noted that the brake light recognition module determines whether the brake lights of the vehicle ahead are on, and activates the brake indicator when the brake lights are on.

[0108] Furthermore, regarding the aforementioned stereo confidence value, exceptions can be set for the brightness image 126 using one of the imaging devices 110. For example, the stereo weight derivation unit 170 can also obtain the ratio of the number of times an object within the frame 200 is identified as a stereoscopic object in the brightness image 126 on the right (right-side recognition count) to the number of times an object within the frame 200 is identified as a stereoscopic object in the brightness image 126 on the left (left-side recognition count). Moreover, if the stereo state level is below "trust" and the ratio of the right-side recognition count to the left-side recognition count is less than 0.9 (the change in monocular speed is less than 0.94 when accelerating), the stereo weight derivation unit 170 can lower the stereo state level by one and reduce the stereo confidence value by 1. This is because when there is a significant discrepancy between the right-side and left-side recognition counts, it is inferred that there is a high probability of raindrops or the like in the detection area 124 of one of the left and right imaging devices 110.

[0109] Furthermore, when both the stereo state and the monocular state are "unreliable," the stereo weight derivation unit 170 can increase the stereo confidence value by 1. This is to avoid deriving the object's velocity solely based on the predicted velocity.

[0110] Next, the prediction weights will be explained. The prediction weight derivation unit 174 derives the prediction weights based on the uncertainty value of the distance image and the uncertainty value of the brightness image of the imaging device. Specifically, the prediction weight derivation unit 174 derives the prediction weights using the following equation (4).

[0111] Wp = (stereo disbelief value) × (monocular disbelief value) ... (4)

[0112] Here, the stereo confidence value is sometimes denoted as Rs, the maximum value of the stereo confidence value is denoted as Rsmax, and the stereo deviation is denoted as Bs. The prediction weight derivation unit 174 derives the unconfidential value by the following equation (5).

[0113] (Stereounreliable value) = Rsmax - min(Rsmax, (Rs + Bs)) ... (5)

[0114] In equation (5), min(Rsmax, (Rs+Bs)) indicates that the smaller of Rsmax and (Rs+Bs) is used. Furthermore, the stereo confidence value (Rs) is the value derived by the stereo weight derivation unit 170. The maximum value of the stereo confidence value (Rsmax) is the stereo confidence value when the stereo state is "trusted", specifically "3".

[0115] Stereo bias (Bs) is a stereo confidence value used to increase the stereo unconfidence value derived from the stereo unconfidence value, thereby reducing the stereo unconfidence value. For example, if the change in the stereo velocity of the stereo object indicates deceleration and the stereo object is close (stereo distance is less than a predetermined value), the stereo bias is set to "3". Conversely, if the change in the stereo velocity of the stereo object indicates deceleration and the stereo object is not close (stereo distance exceeds a predetermined value), the stereo bias is set to "2". Furthermore, if the stereo bias is not in the condition of "3" or "2", it is set to "1".

[0116] In addition, sometimes the monocular confidence value is labeled as Rm, the maximum value of the monocular confidence value is labeled as Rmmax, and the monocular bias is labeled as Bm. The prediction weight derivation unit 174 derives the monocular unconfidence value by the following equation (6).

[0117] (Unreliable value for one eye) = Rmmax - min(Rmmax, (Rm + Bm)) ... (6)

[0118] In equation (6), min(Rmmax, (Rm+Bm)) indicates that the smaller of Rmmax and (Rm+Bm) is used. Furthermore, the monocular confidence value (Rm) uses the value derived from the monocular weight derivation unit 172. The maximum monocular confidence value (Rmmax) is the monocular confidence value when the monocular state is "trusted," specifically "3".

[0119] Monocular bias (Bm) is a monocular confidence value used to increase the derived monocular unconfidence value, which in turn reduces the monocular unconfidence value. For example, if the change in the monocular velocity of a stereoscopic object represents deceleration, the monocular bias is set to "1", and otherwise, the monocular bias is set to "0".

[0120] Setting the stereo bias or monocular bias to "1" or higher reduces the prediction weight ratio. As a result, it prevents the predicted velocity ratio in the calculation of object velocity from decreasing and prevents the responsiveness of object velocity over time from decreasing excessively.

[0121] The object velocity derivation unit 168 applies the stereo weight derived by equation (2), the monocular weight derived by equation (3), and the prediction weight derived by equation (4) to equation (1) to derive the object velocity.

[0122] It should be noted that if the predetermined exception conditions are met, the object velocity derivation unit 168 may not perform the mixing of stereo velocity, monocular velocity, and predicted velocity (no velocity correction processing), and instead use the stereo velocity of the stereo object as the object velocity. The predetermined exception conditions may be set based on factors such as the ambient darkness, stereo distance, stereo velocity, or stereo state of the vehicle 1.

[0123] For example, the object velocity derivation unit 168 can detect the darkness around the vehicle 1. If the darkness is greater than or equal to nighttime, it is considered to meet an exception condition, and the three-dimensional velocity of the three-dimensional object is taken as the object velocity. This is to distinguish between cases where the brightness image 126 is blurred due to raindrops, etc., and cases where the brightness image 126 is darkened due to nighttime.

[0124] Furthermore, the object velocity derivation unit 168 can also consider exception conditions to be met when the three-dimensional distance is less than 20m or the three-dimensional velocity is less than 45km / h, and thus use the three-dimensional velocity of the three-dimensional object as the object velocity. If velocity correction processing is performed when a three-dimensional object is present at close range, the change of the object velocity over time may cause a response delay relative to the behavior of the three-dimensional object. The above-mentioned use of the three-dimensional velocity of the three-dimensional object as the object velocity is to avoid such a response delay.

[0125] Additionally, the object velocity derivation unit 168 can also consider the three-dimensional velocity of the three-dimensional object as the object velocity if the three-dimensional distance is 20m or more but less than 25m (17m or more but less than 25m when the three-dimensional object is determined to be accelerating) and the three-dimensional state level is "stable" or higher. This is also to avoid response delay in the derived object velocity.

[0126] Furthermore, the object velocity derivation unit 168 can also consider the exception condition to be met when the three-dimensional distance is 70m or more, and thus use the three-dimensional velocity of the three-dimensional object as the object velocity. This is because if the three-dimensional object is located at a distance, the values ​​of the monocular velocity and the predicted velocity are prone to deviation, causing a decrease in the accuracy of the object velocity. The above-mentioned use of the three-dimensional velocity of the three-dimensional object as the object velocity is to suppress the decrease in the accuracy of the object velocity.

[0127] Figure 9 This is a flowchart illustrating the operation of the central control unit 154. The central control unit 154 repeats the process at each predetermined interruption time that occurs according to a predetermined control cycle. Figure 9 A series of processes. In Figure 9The flowchart describes the process related to deriving the object's velocity, while omitting descriptions of processes unrelated to deriving the object's velocity.

[0128] First, the stereo velocity export unit 160 acquires brightness images 126 from the left and right imaging devices 110 respectively (S100). Next, the stereo velocity export unit 160 matches the left and right brightness images 126 for each block and exports a distance image 128 containing parallax information (S110).

[0129] Next, the stereo velocity derivation unit 160 determines the stereo object (e.g., a moving vehicle) based on the distance image 128 and derives the stereo distance of the stereo object (S120). Then, the stereo velocity derivation unit 160 derives the stereo velocity of the stereo object based on the stereo distance (S130).

[0130] Next, the monocular speed derivation unit 162 uses the brightness image 126 of the left and right brightness images 126 obtained in step S100 that captures the stereoscopic object more clearly (e.g., a brightness image 126 with generally high brightness values) to derive the monocular speed of the stereoscopic object (S140). Specifically, the monocular speed derivation unit 162 derives the monocular speed based on the change in the lateral width of the stereoscopic object in the brightness images 126 from a predetermined frame ago (e.g., 100 frames ago) to the present. It should be noted that the brightness image used to derive the monocular speed can also be fixed to either the brightness image of the left-side imaging device 110 or the brightness image of the right-side imaging device 110.

[0131] In addition, the monocular velocity deriving unit 162 integrates the derived monocular velocity over time to derive the monocular distance of the stereoscopic object (S150).

[0132] Next, the prediction velocity derivation unit 164 derives the predicted velocity of the solid object based on the past velocity of the solid object (S160). Specifically, the prediction velocity derivation unit 164 uses the object velocity derived at the last interruption time (an interruption time that arrives before the current interruption time, relative to the current interruption time) as the current predicted velocity.

[0133] Next, the object velocity derivation unit 168 determines whether the exception condition is met (S170). If the exception condition is met ("Yes" in S170), the object velocity derivation unit 168 uses the three-dimensional velocity of the three-dimensional object derived in step S130 as the object velocity for this time (S180) and ends the series of processes.

[0134] Furthermore, if the exception condition is not met ("No" in S170), the processing after step S200 is performed. Specifically, the stereo weight derivation unit 170 derives the stereo weights (S200). Next, the monocular weight derivation unit 172 derives the monocular weights (S210). Next, the prediction weight derivation unit 174 derives the prediction weights (S220).

[0135] Then, the object velocity derivation unit 168 derives the object velocity based on the derived stereo weights, monocular weights, and prediction weights (S230), and ends a series of processes. Specifically, the object velocity derivation unit 168 applies the stereo weights, stereo velocity, monocular weights, monocular velocity, prediction weights, and prediction velocity to Equation (1) to derive the object velocity.

[0136] It should be noted that, although in Figure 9 The flowchart is omitted, but if the exception conditions are met during the process of deriving each weight, the object velocity deriving unit 168 can also enter the processing of step S180 and use the three-dimensional velocity as the object velocity.

[0137] And, although in Figure 9 The flowchart is omitted, but if the number of blocks from which parallax is derived in the distance image 128 is less than the predetermined number of blocks, it is possible to switch from stereoscopic control to monocular control. When switched to monocular control, the object velocity derivation unit 168 can use the monocular velocity as the object velocity.

[0138] As described above, in the vehicle exterior environment recognition device 120 of this embodiment, the object velocity is derived by mixing stereo velocity, monocular velocity, and predicted velocity at an appropriate mixing ratio during stereo control. Therefore, in the vehicle exterior environment recognition device 120 of this embodiment, when an error occurs in stereo velocity, the weighting of the stereo velocity with error can be reduced compared to deriving the object velocity solely from stereo velocity. Thus, in the vehicle exterior environment recognition device 120 of this embodiment, even if raindrops or the like are present in the detection area 124 of an imaging device 110, errors in the derived object velocity can be suppressed.

[0139] Therefore, in the vehicle external environment recognition device 120 of this embodiment, the speed of three-dimensional objects can be stably derived. As a result, in the vehicle external environment recognition device 120 of this embodiment, cruise control utilizing the speed of objects can be stably performed.

[0140] The embodiments of the present invention have been described above with reference to the accompanying drawings, but the present invention is by no means limited to these embodiments. Obviously, those skilled in the art will be able to conceive of various modifications or alterations within the scope of the claims, and it should be understood that these modifications or alterations also fall within the technical scope of the present invention.

Claims

1. A vehicle external environment recognition device, characterized in that, have: A stereo velocity extraction unit that extracts stereo velocity, which represents the velocity of a predetermined stereoscopic object extracted based on a distance image derived from brightness images of two imaging devices; A monocular speed output unit that outputs monocular speed, which represents the speed of the stereoscopic object extracted based on a brightness image of the imaging device; A velocity prediction derivation unit derives a predicted velocity, which represents the current velocity of the solid object predicted based on its past velocities. A mixing ratio derivation unit derives the mixing ratio of the stereo velocity, the monocular velocity, and the predicted velocity; as well as The object velocity derivation unit, based on the mixing ratio, mixes the stereo velocity, the monocular velocity, and the predicted velocity to derive an object velocity representing the velocity of the stereoscopic object. The mixing ratio derivation unit has: The solid weight derivation unit derives solid weights that represent the proportion of the solid velocity; The monocular weight derivation unit derives a monocular weight representing the proportion of the monocular velocity; as well as The prediction weight derivation unit derives prediction weights that represent the proportion of the predicted speed. When the stereo velocity is set to Vs, the monocular velocity to Vm, the prediction velocity to Vp, the stereo weight to Ws, the monocular weight to Wm, the prediction weight to Wp, and the object velocity to V, the object velocity derivation unit derives the object velocity using the following equation (1): V=(Ws×Vs+Wm×Vm+Wp×Vp) / (Ws+Wm+Wp)…(1) The stereo weight derivation unit derives the stereo weights based on the confidence value of the distance image.

2. The vehicle exterior environment recognition device according to claim 1, characterized in that, The monocular weight derivation unit derives the monocular weight based on the confidence value of a brightness image of the imaging device.

3. The vehicle exterior environment recognition device according to claim 1 or 2, characterized in that, The prediction weight derivation unit derives the prediction weights based on the unreliability values ​​of the distance image and the brightness image of the shooting device.