Image processing device and vehicle

The image processing device achieves left-right symmetric object recognition by repeatedly performing convolution operations on stereo images, ensuring symmetric performance and standardized evaluation across different driving environments, thereby improving vehicle operations.

JP7742299B2Active Publication Date: 2025-09-19SUBARU CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021207407
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-09-19
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

Existing image processing devices face challenges in ensuring left-right symmetry of performance during object recognition based on stereo images, particularly in different driving environments, leading to asymmetric object recognition results and increased evaluation man-hours.

Method used

The image processing device employs a feature extraction unit that performs convolution operations multiple times on left and right images while ensuring bilateral symmetry, using a process that combines pixel values using an arithmetic process that guarantees left-right symmetry, and includes a vehicle control unit to utilize these results for vehicle operations.

Benefits of technology

This approach ensures left-right symmetric performance in object recognition, reduces the weight of the processing model, and standardizes evaluation work across different driving environments, enhancing convenience and efficiency in vehicle operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742299000001
    Figure 0007742299000001
  • Figure 0007742299000002
    Figure 0007742299000002
  • Figure 0007742299000003
    Figure 0007742299000003
Patent Text Reader

Abstract

To provide an image processing apparatus and the like capable of ensuring bilateral symmetry of performance in object recognition based on stereo images.SOLUTION: An image processing apparatus according to one embodiment of the present disclosure is provided with an extraction unit that extracts feature amounts included in right and left images as stereo images, and an object identification unit that identifies an object based on the feature amounts. The extraction unit repeats a convolution operation a plurality of times based on the right and left images so that the right feature amount and the left feature amount as the feature amounts ensure bilateral symmetric with each other, and performs the combination processing of pixel values themselves using the right feature amount and the left feature amount by using an operation process that ensures bilateral symmetry at the time of the above-described convolution operation.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image processing device that performs object recognition based on captured images, and a vehicle equipped with such an image processing device. [Background technology]

[0002] Images captured by an imaging device include images of various objects. For example, Patent Document 1 discloses an image processing device that performs object recognition based on such captured images. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-128350 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in such image processing devices, it is required to ensure left-right symmetry of performance (model performance) during image processing (object recognition) based on stereo images, for example. It is desirable to provide an image processing device that can ensure left-right symmetry of performance during object recognition based on stereo images, and a vehicle equipped with such an image processing device. [Means for solving the problem]

[0005] An image processing device according to an embodiment of the present disclosure includes an extraction unit that extracts feature amounts included in left and right images of stereo images, and an object identification unit that identifies an object based on the feature amounts. The extraction unit performs a convolution operation multiple times based on the left and right images to extract left and right feature amounts as the feature amounts while ensuring bilateral symmetry with respect to each other, and performs a process of combining pixel values ​​using the left and right feature amounts during the convolution operation by using a process that ensures bilateral symmetry.

[0006] A vehicle according to one embodiment of the present disclosure includes an image processing device according to the embodiment of the present disclosure, and a vehicle control unit that controls the vehicle using the object identification results obtained by the object identification unit. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a block diagram illustrating an example of a schematic configuration of a vehicle according to an embodiment of the present disclosure. [Figure 2] 2 is a top view schematically illustrating an example of the external configuration of the vehicle shown in FIG. 1. FIG. [Figure 3] 2 is a schematic diagram illustrating an example of a left image and a right image generated by the stereo camera illustrated in FIG. 1. [Figure 4] FIG. 2 is a schematic diagram illustrating an example of an image area set in a captured image. [Figure 5] FIG. 10 is a schematic diagram for explaining an overview of a filter update process used in a convolution operation. [Figure 6] 2 is a schematic diagram illustrating an example of application of a convolution operation and an activation function in the feature extraction unit shown in FIG. 1. FIG. [Figure 7] 7 is a schematic diagram illustrating a specific processing example of the convolution operation shown in FIG. 6. [Figure 8] FIG. 7 is a schematic diagram illustrating a specific configuration example of the activation function shown in FIG. 6. [Figure 9] FIG. 10 is a schematic diagram illustrating an example of an object recognition result according to Comparative Example 1. [Figure 10] 5A and 5B are schematic diagrams illustrating an example of an object recognition result in a stereo image according to an embodiment. [Figure 11] FIG. 10 is a schematic diagram illustrating an example of the configuration of a feature amount extraction unit according to Comparative Example 2. [Figure 12] FIG. 10 is a schematic diagram for explaining image transformation during convolution calculation. [Figure 13] FIG. 10 is a schematic diagram for explaining a general joining process. [Figure 14] 10 is a schematic diagram illustrating an example of a joining process according to Comparative Example 2. FIG. [Figure 15] FIG. 2 is a schematic diagram illustrating an example of the configuration of a feature extraction unit according to the embodiment. [Figure 16] FIG. 10 is a schematic diagram illustrating another example configuration of the feature extraction unit according to the embodiment. [Figure 17] 10A and 10B are schematic diagrams for explaining calculation processing during combining processing according to the embodiment. [Figure 18] FIG. 18 is a schematic diagram for explaining a combining process using the arithmetic process shown in FIG. 17. [Figure 19] 10A to 10C are schematic diagrams illustrating an example of a combining process according to an embodiment. [Figure 20] 10 is a schematic diagram for explaining object recognition according to an example and a comparative example 2. FIG. [Figure 21] FIG. 10 is a schematic diagram illustrating an example of an object recognition result according to Comparative Example 2. [Figure 22] FIG. 10 is a schematic diagram illustrating an example of an object recognition result according to the embodiment. [Figure 23] FIG. 10 is a schematic diagram illustrating an example of the configuration of a feature extraction unit according to Modification 1. [Figure 24] FIG. 10 is a schematic diagram illustrating an example of the configuration of a feature amount extraction unit according to Modification 2. [Figure 25] 10 is a schematic diagram illustrating an example of a process for updating a filter value in a filter according to Modification 2. FIG. [Figure 26] FIG. 10 is a schematic diagram illustrating a configuration example of a filter according to Modification 2. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. The description will be made in the following order. 1. Embodiment (Example of a case where left-right reversal processing units are arranged before and after the encoder) 2. Working Example (Example of Specific Object Recognition Results) 3. Variations Modification 1 (Example of a case where separate encoders are provided for left and right) Modification 2 (Example of setting the filter values ​​in the encoder symmetrically) Modification 3 (Example of a superordinate join process using commutative operations) 4. Other Modifications

[0009] <1. Embodiment> [composition] Fig. 1 is a block diagram illustrating an example of the general configuration of a vehicle (vehicle 10) according to an embodiment of the present disclosure. Fig. 2 is a schematic top view illustrating an example of the external configuration of vehicle 10 illustrated in Fig. 1.

[0010] As shown in Fig. 1, vehicle 10 includes a stereo camera 11, an image processing device 12, and a vehicle control unit 13. Note that Fig. 1 does not illustrate the driving power sources (engine, motor, etc.) of vehicle 10. Vehicle 10 may be, for example, an electrically powered vehicle such as a hybrid electric vehicle (HEV) or an electric vehicle (EV), or a gasoline-powered vehicle.

[0011] (A. Stereo Camera 11) 2, the stereo camera 11 is a camera that captures an image of the area ahead of the vehicle 10 to generate a pair of images (a left image PL and a right image PR) having a parallax therebetween. As shown in FIGS. 1 and 2, the stereo camera 11 has a left camera 11L and a right camera 11R.

[0012] The left camera 11L and the right camera 11R each include, for example, a lens and an image sensor. The left camera 11L and the right camera 11R are arranged near the top of the windshield 19 of the vehicle 10, spaced a predetermined distance apart along the width direction of the vehicle 10, as shown in FIG. 2, for example. The left camera 11L and the right camera 11R are configured to perform imaging operations in synchronization with each other. Specifically, as shown in FIG. 1, the left camera 11L generates a left image PL, and the right camera 11R generates a right image PR. The left image PL includes a plurality of pixel values, and the right image PR includes a plurality of pixel values. The left image PL and the right image PR constitute a stereo image PIC, as shown in FIG. 1.

[0013] FIG. 3 shows an example of such a stereo image PIC. Specifically, FIG. 3(A) shows an example of a left image PL, and FIG. 3(B) shows an example of a right image PR. Note that x and y shown in FIG. 3 represent the x-axis and y-axis, respectively. In this example, another vehicle (a preceding vehicle 90) is traveling ahead of the vehicle 10 on the road on which the vehicle 10 is traveling. The left camera 11L captures an image of the preceding vehicle 90 to generate the left image PL, and the right camera 11R captures an image of the preceding vehicle 90 to generate the right image PR.

[0014] The stereo camera 11 generates a stereo image PIC including such a left image PL and a right image PR. The stereo camera 11 also performs an imaging operation at a predetermined frame rate (e.g., 60 fps) to generate a series of stereo images PIC.

[0015] (B. Image processing device 12) The image processing device 12 is a device that performs various image processing (such as recognition processing of objects ahead of the vehicle 10) based on the stereo images PIC supplied from the stereo camera 11. As shown in FIG. 1 , the image processing device 12 has an image memory 121, a feature extraction unit 122, and an object identification unit 123.

[0016] Such an image processing device 12 includes, for example, one or more processors (CPUs: Central Processing Units) that execute programs, and one or more memories communicably connected to these processors. Such memories include, for example, RAMs (Random Access Memory) that temporarily store processing data, and ROMs (Read Only Memory) that store programs.

[0017] The feature extraction unit 122 described above corresponds to a specific example of an "extraction unit" in the present disclosure.

[0018] (Image memory 121) 1, the image memory 121 is a memory that temporarily stores the left image PL and the right image PR included in the stereo image PIC. The image memory 121 sequentially supplies the left image PL and the right image PR thus stored to the feature extraction unit 122 as captured images P (see FIG. 1).

[0019] (Feature extraction unit 122) The feature extraction unit 122 extracts feature values ​​F contained in one or more image regions R in the captured image P (left image PL and right image PR) read out from the image memory 121 (see FIG. 1). Specifically, the feature extraction unit 122 extracts feature values ​​F (left feature values ​​FL) contained in the left image PL and feature values ​​F (right feature values ​​FR) contained in the right image PR, as will be described in detail later. These feature values ​​F (left feature values ​​FL and right feature values ​​FR), as will be described in detail later (FIG. 7), are composed of pixel values ​​of multiple pixels arranged in a matrix (two-dimensional arrangement). Examples of such feature values ​​F include RGB (Red, Green, Blue) feature values ​​and HOG (Histograms of Oriented Gradients) feature values.

[0020] The feature extraction unit 122, the details of which will be described later, uses a trained model such as a DNN (Deep Neural Network) (machine learning) to set the image region R in the captured image P and extract the feature F. When setting the image region R, the feature extraction unit 122, for example, identifies an object in the captured image P and outputs the coordinates of the identified object, thereby setting the image region R, which is a rectangular region.

[0021] Fig. 4 is a schematic representation of an example of such an image region R. In the example shown in Fig. 4, an image region R is set for each of two vehicles in the captured image P. Note that in this example, an image region R is set for the vehicles, but the present invention is not limited to this example, and for example, an image region R may also be set for a person, a guardrail, a wall, etc.

[0022] Here, the process of extracting the feature amount F contained in the captured image P (in one or more image regions R) by the feature amount extracting section 122 will be described in detail with reference to FIGS.

[0023] Fig. 5 is a schematic diagram showing an outline of the update process of the filter FLT used in the convolution operation described later. Fig. 6 is a schematic diagram showing an example of application of the convolution operation and activation function described later in the feature extraction unit 122. Fig. 7 is a schematic diagram showing a specific processing example of the convolution operation shown in Fig. 6. Fig. 8 is a schematic diagram showing a specific configuration example of the activation function shown in Fig. 6.

[0024] First, as shown in Fig. 5, for example, the feature extraction unit 122 performs a convolution operation or the like using a filter FLT (described later) on the input captured image P to obtain an inference result of object recognition by machine learning (such as the extraction result of the feature F within the image region R described above). This inference result is compared with the correct answer data for object recognition as needed (see the dashed arrow CF in Fig. 5), and an update process for the parameters of the filter FLT (each filter value described later) is performed as needed to reduce the difference between these inference results and the correct answer data. In other words, each time the filter FLT is updated by such machine learning, an update process is performed on each filter value in the filter FLT as needed, and a trained model of machine learning is generated.

[0025] In this way, instead of specifying specific processing formulas as in conventional rule-based development, a large amount of training data for machine learning and corresponding correct answer data is prepared, and by repeating the update process described above, the same inference results as the correct answer data can ultimately be obtained.

[0026] 6, for example, the feature extraction unit 122 uses the trained model obtained in this way to repeatedly perform various arithmetic processes based on the input captured image P (left image PL and right image PR) multiple times, thereby performing object recognition (extraction of feature F, etc.) within each image region R in the captured image P. Specifically, as such various arithmetic processes, the feature extraction unit 122 alternately repeats a convolution operation CN using the filter FLT and an operation using the activation function CA multiple times (see FIG. 6).

[0027] Here, as shown in FIG. 7, for example, the above-mentioned convolution operation CN is performed as follows. That is, the feature extraction unit 122 first sets an area of ​​a predetermined size (3 pixels × 3 pixels in this example) in the captured image P having a plurality of pixels PX arranged two-dimensionally in a matrix. The feature extraction unit 122 then weights and adds the nine filter values ​​of the filter FLT to the nine pixel values ​​(values ​​of “0” or “1” in this example) in this set area as weighting coefficients. This results in a value of the feature F in that area (value of “4” in this example). In the example shown in FIG. 7, the filter values ​​of the filter FLT (written as “×0” or “×1”) are arranged two-dimensionally in a matrix, with nine values ​​(3 along the row direction (x-axis direction) × 3 along the column direction (y-axis direction)). The feature extraction unit 122 then sequentially sets the above-mentioned regions in the captured image P while shifting them by one pixel at a time, and sequentially calculates the value of the feature F in each set region by individually performing weighted addition using the above-mentioned filter FLT in each set region. As a result, a feature F having a plurality of pixels PX arranged two-dimensionally in a matrix form is extracted, as shown in Fig. 7, for example. Note that the above-mentioned filter FLT is individually set for each of the multiple convolution operations CN shown in Fig. 6, for example.

[0028] 8, the calculation using the activation function CA is performed as follows. That is, by applying the activation function CA shown in FIG. 8 to the input value (the value of each pixel PX in the feature F obtained by each convolution calculation CN), an output value after application of the activation function CA is obtained. In the example of FIG. 8, the output value is set so that if the input value is less than a predetermined value, the output value is set to a fixed value (for example, "0"), and if the input value is equal to or greater than the predetermined value, the output value increases linearly according to the magnitude of the input value.

[0029] The final feature amount F obtained by repeating such various calculation processes multiple times is supplied from the feature amount extraction unit 122 to the object identification unit 123 (see FIG. 1).

[0030] Note that details of an example configuration (processing example) of the feature amount extraction unit 122 will be described later (FIGS. 15 to 19).

[0031] (Object identification unit 123) The object identification unit 123 identifies an object in the captured image P (each of the one or more image regions R described above) based on the feature amount F extracted by the feature amount extraction unit 122. That is, for example, if the image in the image region R shows a vehicle, the feature amount F includes the feature of the vehicle, and if the image in the image region R shows a person, the feature amount F includes the feature of the person, and therefore the object identification unit 123 identifies an object in each image region R based on such feature amount F.

[0032] Then, the object identification unit 123 assigns a category indicating what the object is to each image region R. Specifically, if the object in the image of image region R is a vehicle, the object identification unit 123 assigns a category indicating a vehicle to that image region R, and if the object in the image of image region R is a person, the object identification unit 123 assigns a category indicating a person to that image region R.

[0033] (C. Vehicle control unit 13) The vehicle control unit 13 uses the object identification result (object recognition result in the image processing device 12) by the object identification unit 123 to perform various vehicle controls on the vehicle 10 (see FIG. 1). Specifically, the vehicle control unit 13 performs, for example, driving control of the vehicle 10 and operation control of various members in the vehicle 10 based on information on such object identification result (object recognition result).

[0034] The vehicle control unit 13 is configured to include, for example, one or more processors (CPUs) that execute programs and one or more memories communicably connected to these processors, similar to the image processing device 12. Also, similar to the image processing device 12, such memories are configured, for example, with a RAM that temporarily stores processed data and a ROM that stores programs.

[0035] [Operation, Actions and Effects] Next, the operation (detailed configuration and processing of the feature extraction unit 122, etc.) and effects of this embodiment will be described in detail while comparing with comparative examples (comparative examples 1 and 2).

[0036] (A. Comparative examples 1 and 2) The convolution operations in the DNNs mentioned above generally have the following issues:

[0037] That is, first, as mentioned above, since filters for convolution operations are generally provided separately for each of the multiple convolution operations, the number of parameters set for each filter (the number of values ​​indicated by the filter value) becomes enormous (for example, on the order of millions) for the entire trained model. This makes it difficult to reduce the weight of the processing model (trained model) used for image processing (object recognition), making it difficult to implement the model in small-scale hardware, such as embedded systems. Note that while methods such as reducing the model size itself or lowering the accuracy of the convolution operation are conceivable, there is a trade-off with model performance (recognition performance).

[0038] Furthermore, since vehicle driving environments (left-hand driving environments or right-hand driving environments) generally differ from country to country, it is desirable for object recognition performance to be symmetrical, but the convolutional operations in a typical DNN result in asymmetric object recognition performance. Therefore, separate evaluation work is required during machine learning for both left-hand driving environments and right-hand driving environments, which increases the evaluation man-hours.

[0039] Here, for example, a method of performing machine learning on an artificial image that is flipped left and right (a flipped image) can be considered. However, even when this method is used, there are cases where strict left-right symmetry cannot be obtained, as in the following Comparative Example 1, and in such cases, the number of evaluation steps ends up increasing.

[0040] Fig. 9 is a schematic representation of an example of an object recognition result (object identification result) according to Comparative Example 1. In the object recognition result according to Comparative Example 1 shown in Fig. 9, when the vehicle driving environment in the original captured image P is a left-hand driving environment (see Fig. 9(A)), the object recognition result is as follows in the above-mentioned artificially reversed image PLR ​​(see Fig. 9(B)).

[0041] In Figures 9(A) and 9(B), the front part of the recognized vehicle is shown by a solid line and the rear part of the recognized vehicle is shown by a dashed line in each image area R set during object recognition.

[0042] Here, in the object recognition result for the original captured image P shown in FIG. 9(A), the front and rear portions of the recognized vehicle are accurately recognized, for example, as in image region R within the region indicated by the dashed circle. On the other hand, in the object recognition result for the left-right reversed image PLR ​​shown in FIG. 9(B), unlike the case of the original captured image P, partially inaccurate recognition results are obtained. Specifically, for example, as in image region R within the region indicated by the dashed circle in FIG. 9(B), the front and rear portions of the recognized vehicle are reversed. In other words, it can be seen that the object recognition performance is not symmetrical in the case of Comparative Example 1 shown in FIG. 9.

[0043] In particular, when performing image processing (object recognition) based on the stereo images PIC (left image PL and right image PR), the feature extraction unit 122 is required to ensure the following left-right symmetry.

[0044] That is, as in the example of the object recognition result in the stereo image PIC according to this embodiment shown in Fig. 10, it is desirable that the left image PL and the right image PR included in the stereo image PIC be as follows. Specifically, for example, as indicated by the dashed arrows in Fig. 10, even when FLIP (left-right inversion processing) and SWAP (exchanging the left and right inputs) are performed on the left image PL and the right image PR input to the feature amount extraction unit 122, it is desirable that left-right symmetry is completely ensured in the inference result output from the feature amount extraction unit 122. In other words, when the left and right inputs are exchanged between an inverted image PL' obtained by performing left-right inversion processing on the left image PL and an inverted image PR' obtained by performing left-right inversion processing on the right image PR, it is desirable that the inference results for these inverted images PL', PR' also undergo the same FLIP (left-right inversion processing) and SWAP (exchanging the left and right inputs) as those on the input side.

[0045] Here, Fig. 11 is a schematic diagram showing an example of the configuration (processing example) of the feature extraction unit 202 according to Comparative Example 2. Fig. 12 is a schematic diagram for explaining image conversion during the above-mentioned convolution operation, and Fig. 13 is a schematic diagram for explaining a conventional general combining process. Fig. 14 is a schematic diagram showing an example of the combining process according to Comparative Example 2 shown in Fig. 11.

[0046] As shown in FIG. 11, the feature extraction unit 202 according to this comparative example 2 has two encoders EnC common to the left image PL and the right image PR, a concatenation processing unit Con202 that performs the concatenation processing (Concat processing) described later, and two decoders DeL and DeR.

[0047] The encoder EnC extracts a left feature value FL or a right feature value FR based on the left image PL or the right image PR by using the filter FLT described above. The combining processor Con202 performs combining processing (a conventional general combining processing) between pixel values ​​using the left feature value FL and the right feature value FR. The decoders DeL and DeR output an inference result for the left image PL or the right image PR, respectively, based on the output result from the combining processor Con202 (the combining processing result between the left feature value FL and the right feature value FR).

[0048] Here, a conventional general joining process will be described with reference to FIGS.

[0049] First, in the above-mentioned convolution operation CN, image conversion is performed in each convolution operation CN, for example, as shown in Fig. 12. Specifically, in the example of Fig. 12, image conversion is performed from captured image P (width W, height H, channel C) to captured image P' (width W', height H', channel C') in the convolution operation CN. In this way, the image size before and after this image conversion generally differs.

[0050] Furthermore, as shown in Fig. 13, the process of combining multiple images of the same shape (size) along the channel direction is generally called a concatenation process (concat process). Specifically, in the example of Fig. 13, the above-mentioned concatenation processor Con202 performs a conventional general concatenation process, so that pixel values ​​in the two captured images Pa and Pb are combined along the channel C.

[0051] For these reasons, when the general combining process described above is applied to the left feature amount FL and the right feature amount FR in the combining processor Con202 of Comparative Example 2 shown in Fig. 11, the result is as shown in Fig. 14, for example. Note that the gradation of light and shade along the left-right (width W) direction in the left feature amount FL and right feature amount FR shown in Fig. 14 and the left feature amount FL' and right feature amount FR' after the left-right reversal process are shown for convenience in terms of the pixel values ​​(reversal state) before and after the left-right reversal process. The meaning of such gradation of light and shade is the same in other drawings (Figs. 17 to 19) described later.

[0052] 14, when the above-described FLIP (left-right reversal processing) and SWAP (exchanging the left and right inputs) are performed on the left feature amount FL and the right feature amount FR, the left feature amount FL′ and the right feature amount FR′ after the left-right reversal processing are compared as follows: That is, the result of the general recombination processing by the recombination processing unit Con202 based on the original left feature amount FL and right feature amount FR and the result of the general recombination processing by the recombination processing unit Con202 based on the left feature amount FL′ and right feature amount FR′ after the left-right reversal processing are not left-right reversed with respect to each other along the width W direction (see the dashed arrows in FIG. 14).

[0053] For these reasons, it can be said that in this Comparative Example 2, it is difficult to ensure left-right symmetry of performance (model performance) when performing image processing (object recognition) based on stereo images PIC (left image PL and right image PR).

[0054] (B. This embodiment) Therefore, in this embodiment, a feature extraction unit 122 having the following configuration is provided instead of the feature extraction unit 202 in Comparative Example 2. In this embodiment, instead of the general combining process in the combining processor Con202, the combining process (combining process that ensures left-right symmetry) described below is performed in the combining processor Con.

[0055] 15 and 16 each schematically show an example of the configuration (processing example) of the feature extraction unit 122 according to this embodiment. FIG. 17 is a schematic diagram for explaining the calculation process during the combining process according to this embodiment (the combining process in the combining processor Con, which will be described later). FIG. 18 is a schematic diagram for explaining the combining process using the calculation process shown in FIG. 17. FIG. 19 is a schematic diagram for explaining an example of the combining process according to this embodiment.

[0056] The configuration example of the feature extraction unit 122 shown in Fig. 15 is similar to the configuration example of the feature extraction unit 202 of Comparative Example 2 shown in Fig. 11, except that left-right reversal processing units Fp that perform left-right reversal processing of pixel values ​​are provided before and after the encoder EnC for the left image PL, and a joining processing unit Con is provided instead of the joining processing unit Con202. On the other hand, the configuration example of the feature extraction unit 122 shown in Fig. 16 is similar to the configuration example of the feature extraction unit 122 shown in Fig. 15, except that the decoder DeL for the left image PL is omitted (not provided).

[0057] In the configuration examples shown in Figures 15 and 16, left-right reversal processing units Fp may be provided before and after the encoder EnC for the right image PR, instead of before and after the encoder EnC for the left image PL.

[0058] In this way, in the feature extraction unit 122 of this embodiment, for one of the left image PL and the right image PR, left-right inversion processing of pixel values ​​is performed before and after the encoder EnC (process for extracting feature F), thereby ensuring left-right symmetry between the left feature FL and the right feature FR.

[0059] Furthermore, unlike the combination processing unit Con202 of Comparative Example 2 described above, the combination processing unit Con of this embodiment performs a combination process (Concat process) of pixel values ​​using the left feature amount FL and the right feature amount FR using, for example, an arithmetic process that ensures left-right symmetry, as described below.

[0060] In this calculation process, first, an added image Fs and a difference image Fd are respectively obtained as shown in, for example, Fig. 17(A) and Fig. 17(B). Specifically, as shown in, for example, Fig. 17(A), the added image Fs is an image corresponding to the added value of pixel values ​​between the left feature amount FL and the right feature amount FR. On the other hand, as shown in, for example, Fig. 17(B), the difference image Fd is an image corresponding to the absolute value of the difference between pixel values ​​between the left feature amount FL and the right feature amount FR. However, in order to avoid a discrepancy in size (absolute value) between the added image Fs and the difference image Fd, in practice, √2 (= 2 1 / 2 ) is the value after division by Fs=(FL+FR) / √2 ……(1) Fd=|FL-FR| / √2 ……(2)

[0061] Next, in this calculation process, a concatenation process (concat process) is performed on pixel values ​​between the obtained added image Fs (the added value of pixel values ​​between the left feature amount FL and the right feature amount FR) and the difference image Fd (the absolute value of the difference between pixel values ​​between the left feature amount FL and the right feature amount FR), as shown in Fig. 18. In other words, in the concatenation process in the concatenation process unit Con of this embodiment, the above-mentioned general concatenation process (concatenation process in the concatenation process unit Con202) is performed on such added image Fs and difference image Fd.

[0062] The feature extraction unit 122 of this embodiment differs from the feature extraction unit 202 of the comparative example 2 in the following manner.

[0063] 19, for example, when the above-described FLIP (left-right reversal processing) and SWAP (exchanging the left and right inputs) are performed on the left feature amount FL and the right feature amount FR, the left feature amount FL' and the right feature amount FR' after the left-right reversal processing are compared as follows: That is, the result of the combining processing by the combining processor Con based on the original left feature amount FL and right feature amount FR and the result of the combining processing by the combining processor Con based on the left feature amount FL' and right feature amount FR' after the left-right reversal processing are left-right reversed with respect to each other along the width W direction (see the dashed arrow in FIG. 19).

[0064] Therefore, in this embodiment, unlike Comparative Example 2, it can be said that left-right symmetry of performance (model performance) is guaranteed when performing image processing (object recognition) based on stereo images PIC (left image PL and right image PR).

[0065] Furthermore, in this embodiment, left-right symmetry is ensured for the object identification result (object recognition result) by object identification unit 123. Specifically, for example, left-right symmetry is ensured for the object identification result by object identification unit 123 when the driving environment of vehicle 10 is a left-side driving environment and the object identification result by object identification unit 123 when the driving environment of vehicle 10 is a right-side driving environment.

[0066] (C. Actions and Effects) In this manner, in this embodiment, the left feature amount FL and the right feature amount FR are extracted while ensuring bilateral symmetry by repeatedly performing the convolution operation multiple times based on the stereo images PIC (left image PL and right image PR). During such convolution operation, the concatenation process (concat process) of pixel values ​​using the left feature amount FL and the right feature amount FR is performed using the above-described arithmetic process that ensures bilateral symmetry.

[0067] As a result, in this embodiment, as described above, left-right symmetric performance is ensured when object identification (object recognition) is performed based on the extracted feature amounts F (left feature amount FL and right feature amount FR). As a result, in this embodiment, it is possible to ensure symmetry of performance when performing image processing (object recognition) based on the stereo images PIC.

[0068] Furthermore, in this embodiment, the calculation process for ensuring the left-right symmetry is a process for combining pixel values ​​between the aforementioned added image Fs (the added value of pixel values ​​between the left feature amount FL and the right feature amount FR) and the difference image Fd (the absolute value of the difference between pixel values ​​between the left feature amount FL and the right feature amount FR), and therefore the following occurs: By utilizing such a combining process using the added image Fs and the difference image Fd, it becomes possible to easily ensure the left-right symmetry.

[0069] Furthermore, in this embodiment, pixel values ​​of one of the left image PL and the right image PR are subjected to left-right flipping processing before and after the process of extracting the feature amount F (a left-right flipping processing unit Fp is provided). The left-right flipping processing before and after the process ensures that the left feature amount FL and the right feature amount FR are symmetrical with each other, resulting in the following: In other words, compared to, for example, Modification 1 described below, it is possible to reduce the weight of the processing model (trained model) used in image processing (object recognition).

[0070] Additionally, in this embodiment, the image processing device 12 is mounted on the vehicle 10. As described above, bilateral symmetry is ensured for the object identification results by the object identification unit 123 in the case of a left-side driving environment and a right-side driving environment in the vehicle 10, resulting in the following: In other words, since bilateral symmetry of the object identification performance is ensured in both the case of a left-side driving environment and a right-side driving environment, it is possible to improve convenience and also to standardize the evaluation work during machine learning, thereby enabling a reduction in the number of evaluation steps.

[0071] <2. Example> Next, a specific example according to the above embodiment will be described in detail while comparing it with the above-mentioned comparative example 2.

[0072] Fig. 20 is a schematic diagram for explaining object recognition according to the example and comparative example 2. Fig. 21 is a schematic representation of an example of an object recognition result according to comparative example 2, and Fig. 22 is a schematic representation of an example of an object recognition result according to the example.

[0073] First, the feature extraction unit 122 according to the embodiment and the feature extraction unit 202 according to the comparative example 2 shown in Fig. 20 used the configurations shown in Fig. 16 and Fig. 11, respectively. Then, in the embodiment and the comparative example 2, feature extraction and object recognition were performed based on the stereo images PIC (left image PL and right image PR) described above, with another vehicle traveling ahead as the detection target.

[0074] The inference results by object recognition were carried out only for the vehicle position on the right image PR, such as the rear image area Rb and the side image area Rs as the image area R, respectively shown in Figures 20 to 22. Furthermore, in a good environment that can be handled by a monocular DNN, it is difficult to see the difference between the working example and comparative example 2, so each of these working example and comparative example 2 was carried out in a rainy situation where strong noise was included in the monocular image (right image PR).

[0075] First, in Comparative Example 2 shown in FIG. 21, the results of object recognition (side recognition) in the side image area Rs (disturbance during recognition) are similar before and after the aforementioned FLIP (left-right reversal process) and SWAP (exchanging left and right inputs). Specifically, in Comparative Example 2, before and after each of these processes, the disturbance (mistake) in the side image area Rs is recognized on the right side of the vehicle (other vehicle). Incidentally, if left-right symmetry performance is guaranteed during object recognition, it can be said that the disturbance in the side image area Rs described above should also be reversed left-right in response to the left-right reversal of the input images (left image PL and right image PR) (reversed images PL', PR').

[0076] 22, unlike Comparative Example 2, the object recognition results (disturbance during recognition) in the side image region Rs are reversed left and right before and after each of the FLIP (left-right reversal process) and SWAP (exchange of left and right inputs) processes. In other words, as described above, in this example, it can be seen that the disturbance in the side image region Rs is also reversed left and right in response to the left-right reversal (reversed images PL', PR') of the input images (left image PL and right image PR).

[0077] From the above, it was actually confirmed that in this example, unlike Comparative Example 2, as described above, left-right symmetric performance is guaranteed when recognizing an object based on stereo images PIC (when a vehicle is the detection target). Note that the implementation examples in the above-described Example and Comparative Example 2 are merely examples, and similar evaluation results (object recognition results) were obtained in the Example and Comparative Example 2 even in the case of other implementation examples.

[0078] <3. Modifications> Next, modifications of the above embodiment (Modifications 1 to 3) will be described. Note that, in the following, the same components as those in the embodiment will be given the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0079] [Variation 1] (composition) In the above embodiment, an example has been described in which the left-right reversal processor Fp is arranged before and after the encoder EnC common to the left image PL and the right image PR. In contrast, in Modification 1, an example will be described in which separate (dedicated) encoders EnL and EnR are provided for the left image PL and the right image PR, respectively.

[0080] Fig. 23 is a schematic diagram showing an example configuration (processing example) of a feature extraction unit 122A according to Modification 1. This feature extraction unit 122A is similar to the feature extraction unit 122 of the embodiment shown in Figs. 15 and 16 in that it does not include (omits) the two left-right reversal processing units Fp (pre-stage and post-stage), and it is provided with the dedicated left-right encoders EnL and EnR instead of the common left-right encoder EnC.

[0081] 23, the encoder EnL is an encoder dedicated to the left image PL, and has a filter FLT(L) dedicated to the left image PL as the filter FLT described above. The encoder EnR is an encoder dedicated to the right image PR, and has a filter FLT(R) dedicated to the right image PR as the filter FLT described above.

[0082] With this configuration, the feature extraction unit 122A also ensures bilateral symmetry between the left feature FL and the right feature FR, similar to the feature extraction unit 122 of the embodiment.

[0083] (Actions and Effects) In such a first modification, it is possible to obtain the same effects as those of the above-described embodiment through the same actions as those of the above-described embodiment. That is, it is possible to ensure symmetry of performance when performing image processing (object recognition) based on stereo images PIC.

[0084] [Variation 2] (composition) In the above embodiment, an example has been described in which the left-right reversal processors Fp are arranged before and after the encoder EnC common to the left image PL and the right image PR. In contrast, in Modification 2, an example will be described in which the filter values ​​Vf in a filter (filter FLT2 described later) in the encoder EnC common to the left image PL and the right image PR are set symmetrically.

[0085] Fig. 24 schematically shows an example of the configuration (processing example) of the feature extraction unit 122B according to Modification 2. Fig. 25 also schematically shows an example of the update processing of the filter value Vf in the filter FLT2 according to Modification 2. Fig. 26 schematically shows an example of the configuration of the filter FLT2 according to Modification 2.

[0086] The feature extraction unit 122B shown in Figure 24 is similar to the feature extraction unit 122 of the embodiment shown in Figures 15 and 16 in that it does not have (omits) the two left-right reversal processing units Fp (pre-stage and post-stage) described above, and also has a filter FLT2 instead of the filter FLT described above in the above-mentioned common left-right encoder EnC.

[0087] Specifically, as shown in FIG. 24, the encoder EnC is an encoder common to the left image PL and the right image PR, and has a filter FLT2 common to the left image PL and the right image PR instead of the filter FLT described above.

[0088] Here, with reference to FIGS. 25 and 26, an example of the configuration of the filter FLT2 of the second modification will be described in comparison with the filter FLT described above.

[0089] First, as shown in Fig. 25 , in the above-described filter FLT, a plurality of filter values ​​Vf are each set arbitrarily, unlike the filter FLT2 of Modification Example 2. Specifically, the filter values ​​Vf in this filter FLT are not line-symmetric (bilaterally symmetric) values ​​about a symmetry axis As that runs along a predetermined direction (the y-axis direction in this example) (see the dashed arrows in Fig. 25 ).

[0090] In contrast to this, in the filter FLT2 of the second modification, unlike the above-described filter FLT, a plurality of filter values ​​Vf are set as follows, as shown in, for example, FIGS.

[0091] 26, in the filter FLT2 of Modification 2, the plurality of filter values ​​Vf are set to values ​​that are line-symmetrical about the above-described axis of symmetry As. Specifically, in this example, such line symmetry is bilateral symmetry (symmetry along the x-axis direction) about the axis of symmetry As, and the plurality of filter values ​​Vf are set to values ​​that are bilaterally symmetrical (see the dashed arrows in FIG. 26).

[0092] Furthermore, such left-right symmetry setting for each filter value Vf is performed, for example, as shown in Fig. 25. That is, every time the filter (filter FLT2) based on the above-described machine learning is updated (see Fig. 5), an update process is performed on the multiple filter values ​​Vf as needed, so that the multiple filter values ​​Vf in the filter FLT2 are each set to the above-described line-symmetric (left-right symmetry) values.

[0093] Specifically, as shown by the dashed arrows and the calculation formula (division formula) in Fig. 25, the update process of the filter value Vf in this case is as follows: That is, the filter values ​​Vf at two line-symmetric positions (left-right symmetric positions in this example) about the symmetry axis As are updated to the average value of the filter values ​​Vf at the two line-symmetric positions. By this update process, as shown in Fig. 25, for example, a configuration in which the plurality of filter values ​​Vf are not line-symmetric (left-right symmetric) like the above-mentioned filter FLT (each filter value Vf is set arbitrarily) is updated to the above-mentioned filter FLT2 exhibiting line symmetry (left-right symmetry).

[0094] With this configuration, the feature extraction unit 122B also ensures bilateral symmetry between the left feature amount FL and the right feature amount FR, similarly to the feature extraction unit 122 of the embodiment. That is, by setting the plurality of filter values ​​Vf in the filter FLT2 to bilaterally symmetrical values, the left feature amount FL and the right feature amount FR are ensured to be bilaterally symmetric.

[0095] Furthermore, in the filter FLT2 of this modified example 2, as described above, the multiple filter values ​​Vf are set to values ​​that are symmetrical on the left and right, thereby ensuring symmetry in the object identification results (object recognition results) by the object identification unit 123.

[0096] (Actions and Effects) In such a modified example 2, it is possible to obtain the same effects as those of the above-described embodiment through the same actions as those of the above-described embodiment. That is, it is possible to ensure symmetry of performance when performing image processing (object recognition) based on stereo images PIC.

[0097] In particular, in this modification 2, a convolution operation is performed using a filter FLT2 having a plurality of filter values ​​Vf arranged two-dimensionally, thereby extracting feature values ​​F (left feature values ​​FL and right feature values ​​FR) contained in a stereo image PIC (left image PL and right image PR). The plurality of filter values ​​Vf in this filter FLT2 are set to values ​​symmetrical about a symmetry axis As extending in a predetermined direction, thereby ensuring symmetry between the left feature values ​​FL and the right feature values ​​FR.

[0098] As a result, in Modification 2, the number of parameters included in filter FLT2 (the number of values ​​indicated by filter values ​​Vf) is reduced compared to when, for example, multiple filter values ​​Vf are not symmetrical (each filter value Vf is set arbitrarily), as in the above-mentioned filter FLT. Specifically, in the examples of FIGS. 25 and 26 described above, the number of such parameters is reduced to about half in filter FLT2 of Modification 2 compared to filter FLT. Therefore, in Modification 2, it is possible to further reduce the weight of the processing model (trained model) used in image processing (object recognition).

[0099] Furthermore, in this Modification 2, an update process for the multiple filter values ​​Vf is executed as needed each time the filter FLT2 is updated by the machine learning described above, so that the multiple filter values ​​Vf in the filter FLT2 are set to symmetrical values. In this Modification 2, the update process for each filter value Vf is a process of updating each of the filter values ​​Vf at two symmetrical positions about the symmetry axis As to the average value of the filter values ​​Vf at the two symmetrical positions. For these reasons, in this Modification 2, the process of setting each filter value Vf to a symmetrical value can be easily performed.

[0100] [Variation 3] In each of the above-described embodiment and variants 1 and 2, the calculation process to ensure the left-right symmetry described above involves combining pixel values ​​between the additive image Fs (the additive value of pixel values ​​between the left feature amount FL and the right feature amount FR) and the differential image Fd (the absolute value of the difference between pixel values ​​between the left feature amount FL and the right feature amount FR).

[0101] In contrast to this, in variant example 3, a higher-level (generalized) combining process is used, including combining process using such an additive image Fs and a differential image Fd, to realize the calculation process that ensures the aforementioned left-right symmetry.

[0102] Specifically, in this modification 3, the calculation process for ensuring the left-right symmetry described above is a combining process in which pixel values ​​between the left feature amount FL and the right feature amount FR are combined using one or more commutative calculations. Here, the commutative calculation (calculation process) is the calculation f(X,Y) that calculates a new value based on the variable X and the variable Y, and is a combination process in which f(X,Y )=f(Y,X) (operation that satisfies the commutative law).

[0103] The two types of operations described in the above embodiment, the addition operation (f(X,Y)=X+Y) to obtain the added image Fs and the absolute difference operation (f(X,Y)=|XY|) to obtain the difference image Fd, are both commutative operations. In addition, for example, the multiplication shown in the following equation (3), the operation shown in the following equation (4) (operation to obtain the maximum value of X and Y), and the operation shown in the following equation (5) (operation to obtain the minimum value of X and Y) are also commutative operations. f(X,Y)=X×Y ……(3) f(X,Y)=max(X,Y) ……(4) f(X,Y)=min(X,Y) ……(5)

[0104] Furthermore, for example, the operation shown in the following formula (6), which is an operation combining a plurality of operations having commutativity (commutative operation) f, g, and h, is also a commutative operation. f(g(X,Y)),h(X,Y)) ……(6)

[0105] In this way, even when a process (combining process) is used in which pixel values ​​between the left feature amount FL and the right feature amount FR are combined using one or more of the above-mentioned commutative operations, it is possible to realize the above-mentioned arithmetic process that ensures left-right symmetry. Note that examples of commutative operations are not limited to the above-mentioned examples, and other commutative operations may also be used.

[0106] In such a third modification, it is possible to obtain the same effects as those of the above-described embodiment through the same actions as those of the above-described embodiment. That is, it is possible to ensure symmetry of performance when performing image processing (object recognition) based on stereo images PIC.

[0107] <4. Other Modifications> The present disclosure has been described above by giving embodiments, modifications, and examples, but the present disclosure is not limited to these embodiments, etc., and various modifications are possible.

[0108] For example, the configuration (type, shape, arrangement, number, etc.) of each component in the vehicle 10 and the image processing device 12 is not limited to that described in the above embodiment, etc. In other words, the configuration of each component may be of a different type, shape, arrangement, number, etc. Furthermore, the values, ranges, magnitude relationships, etc. of the various parameters described in the above embodiment, etc. are not limited to those described in the above embodiment, etc., and may be other values, ranges, magnitude relationships, etc.

[0109] Specifically, for example, in the above-described embodiment, the stereo camera 11 is configured to capture an image in front of the vehicle 10, but this is not limited to such a configuration, and for example, the stereo camera 11 may be configured to capture an image of the side or rear of the vehicle 10.

[0110] Furthermore, for example, in the above-described embodiments, the various processes performed in the vehicle 10 and the image processing device 12 have been described using specific examples, but the present invention is not limited to these specific examples. That is, other methods may be used to perform these various processes. Specifically, for example, the above-described combining process method and filter update process method are not limited to the methods described in the above-described embodiments, but other methods may be used. Furthermore, filter values ​​may be set symmetrically using methods other than those described in the above-described embodiments. Furthermore, in the above-described embodiments, the case where a convolution operation is performed multiple times has been described as an example, but the present invention is not limited to this example. That is, for example, feature values ​​may be extracted by performing a single convolution operation in combination with another calculation method.

[0111] Furthermore, the series of processes described in the above embodiments may be performed by hardware (circuits) or software (programs). When performed by software, the software is composed of a group of programs for causing a computer to execute each function. Each program may be pre-installed in the computer, or may be installed on the computer from a network or a recording medium.

[0112] Furthermore, in the above embodiments, an example has been described in which the image processing device 12 is mounted on a vehicle, but this is not limited to this example, and such an image processing device 12 may be provided, for example, in a moving body other than a vehicle or in a device other than a moving body.

[0113] Furthermore, the various examples described above may be applied in any combination.

[0114] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0115] The present disclosure can also be configured as follows. (1) an extraction unit that extracts feature amounts included in the left image and the right image as stereo images; an object identification unit that identifies an object based on the feature amount; Equipped with The extraction unit By repeating a convolution operation multiple times based on the left image and the right image, the left feature amount and the right feature amount are extracted so as to ensure bilateral symmetry with each other, and During the convolution operation, a process of combining pixel values ​​using the left feature amount and the right feature amount is performed using a calculation process that ensures left-right symmetry. Image processing device. (2) The arithmetic processing The pixel values ​​between the left feature amount and the right feature amount are A join operation that combines data using one or more commutative operations. The image processing device according to (1) above. (3) The arithmetic processing an addition value of pixel values ​​between the left feature amount and the right feature amount; the absolute value of the difference between pixel values ​​between the left feature amount and the right feature amount; This is a process of combining pixel values ​​between The image processing device according to (2) above. (4) For one of the left image and the right image, By performing left-right inversion processing of pixel values ​​at the front stage and the back stage of the processing for extracting the feature amount, The left and right features are symmetrical with each other. The image processing device according to any one of (1) to (3) above. (5) the extraction unit extracts the feature by performing the convolution operation using a filter having a plurality of filter values ​​arranged two-dimensionally; The plurality of filter values ​​in the filter are set to symmetrical values, The left and right features are symmetrical with each other. The image processing device according to any one of (1) to (3) above. (6) Each time the filter is updated by machine learning, an update process is performed on the plurality of filter values ​​as needed, the plurality of filter values ​​in the filter are set to the symmetrical values, The update process is a process of updating the filter values ​​at two symmetrical positions with respect to a predetermined symmetry axis to the average value of the filter values ​​at the two symmetrical positions. The image processing device according to (5) above. (7) The image processing device is mounted on a vehicle, a result of identifying the object by the object identifying unit when the vehicle is in a left-side driving environment; and The object identification result by the object identification unit when the vehicle driving environment is a right-hand driving environment. Regarding the The image processing device according to any one of (1) to (6) above. (8) An image processing device according to any one of (1) to (7) above; a vehicle control unit that controls the vehicle by using the object identification result by the object identification unit; A vehicle equipped with. (9) one or more processors; one or more memories communicatively coupled to the one or more processors; Equipped with the one or more processors: Extracting feature amounts contained in the left and right images as stereo images; identifying an object based on the feature amount; In addition to carrying out the above, extracting left and right feature amounts as the feature amounts by repeating a convolution operation multiple times based on the left image and the right image, while ensuring bilateral symmetry between the left feature amount and the right feature amount; During the convolution operation, a process of combining pixel values ​​using the left feature amount and the right feature amount is performed using a calculation process that ensures left-right symmetry. Image processing device. [Explanation of symbols]

[0116] 10...vehicle, 11...stereo camera, 11L...left camera, 11R...right camera, 12...image processing device, 121...image memory, 122, 122A, 122B...feature extraction unit, 123...object identification unit, 13...vehicle control unit, 19...windshield, 90...leading vehicle, PL...left image, PR...right image, PL', PR'...inverted image, PIC...stereo image, P, P', Pa, Pb...captured image, R...image area, Rb...rear image area, Rs...side image area, F...feature FL, FL'...left feature, FR, FR'...right feature, Fs, Fs'...addition image, Fd, Fd'...difference image, FLT, FLT(L), FLT(R), FLT2...filter, Vf...filter value, As...axis of symmetry, CN...convolution operation, CA...activation function, PX...pixel, EnC, EnL, EnR...encoder, DeL, DeR...decoder, Con, Con202...combination processing unit, W, W'...width, H, H'...height, C, C'...channel, Fp...left-right flip processing unit.

Claims

1. an extraction unit that extracts feature amounts included in the left image and the right image as stereo images; an object identification unit that identifies an object based on the feature amount; Equipped with The extraction unit By repeating a convolution operation multiple times based on the left image and the right image, the left feature amount and the right feature amount are extracted so as to ensure bilateral symmetry with each other, and During the convolution operation, a process of combining pixel values ​​using the left feature amount and the right feature amount is performed using a calculation process that ensures left-right symmetry. Image processing device.

2. The arithmetic processing The pixel values ​​between the left feature amount and the right feature amount are A join operation that combines data using one or more commutative operations. The image processing device according to claim 1 .

3. The arithmetic processing an addition value of pixel values ​​between the left feature amount and the right feature amount; the absolute value of the difference between pixel values ​​between the left feature amount and the right feature amount; This is a process of combining pixel values ​​between The image processing device according to claim 2 .

4. For one of the left image and the right image, By performing left-right inversion processing of pixel values ​​at the front stage and the back stage of the processing for extracting the feature amount, The left and right features are symmetrical with each other.

4. The image processing device according to claim 1.

5. the extraction unit extracts the feature by performing the convolution operation using a filter having a plurality of filter values ​​arranged two-dimensionally; The plurality of filter values ​​in the filter are set to symmetrical values, The left and right features are symmetrical with each other.

4. The image processing device according to claim 1.

6. Each time the filter is updated by machine learning, an update process is performed on the plurality of filter values ​​as needed, the plurality of filter values ​​in the filter are set to the symmetrical values, The update process is a process of updating the filter values ​​at two symmetrical positions with respect to a predetermined symmetry axis to an average value of the filter values ​​at the two symmetrical positions. The image processing device according to claim 5 .

7. The image processing device is mounted on a vehicle, a result of identifying the object by the object identifying unit when the vehicle is in a left-side driving environment; and The object identification result by the object identification unit when the vehicle driving environment is a right-hand driving environment. Regarding the 7. The image processing device according to claim 1.

8. The image processing device according to any one of claims 1 to 7, a vehicle control unit that controls the vehicle by using the object identification result by the object identification unit; A vehicle equipped with.

Citation Information

Patent Citations

  • Moving object detector and program

    JP2012064153A

  • Image processing method, image processing device, on-vehicle device, moving body and system

    JP2019128350A

  • Program, recognition apparatus, and recognition method

    JP2019219904A

  • Method and device for 3D object detection

    WO2021212420A1