Image recognition device, in-vehicle system, angle of view adjustment method, and angle of view adjustment program
The image recognition device adjusts the camera angle based on the distance to the most dangerous object, optimizing object size and enhancing recognition accuracy for safety-critical objects.
Patent Information
- Application Number
- JP2024024852
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-21
- Publication Date
- 2025-09-02
AI Technical Summary
Existing image recognition systems in vehicles struggle to accurately adjust the camera angle of view to optimize object size in target images, particularly for objects requiring attention for safety, and fail to effectively switch cameras based on the level of danger posed by objects around the vehicle.
An image recognition device that adjusts the camera angle of view based on the distance between the vehicle and the object posing the greatest risk, optimizing the size of the target object in the image to improve accuracy and safety.
Enhances the accuracy of image recognition for high-risk objects by optimizing their size in the target image, thereby improving vehicle safety and pedestrian protection.
Smart Images

Figure 2025127871000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image recognition device, an in-vehicle system, a method for adjusting an angle of view, and a program for adjusting an angle of view. [Background technology]
[0002] A technology has been put into practical use that detects objects around a vehicle by performing image recognition processing based on images captured by a camera installed in the vehicle, and notifies the driver based on the detection results or controls driving based on the detection results. Note that Patent Document 1 listed below discloses a method in which multiple cameras with different angles of view are installed and the camera to be used is switched depending on the vehicle speed. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-24120 Summary of the Invention [Problem to be solved by the invention]
[0004] In order to properly perform image recognition processing (to improve the accuracy of image recognition processing), it is desirable to properly size objects (image size) in the target image for image recognition. Meanwhile, to ensure the safety of vehicles, pedestrians, etc., it is essential to improve the accuracy of image recognition processing for objects that, among the objects around the vehicle, require particular attention for ensuring safety. While the image size of an object can be adjusted by adjusting the camera's angle of view, ingenuity is required in determining how to adjust the angle of view. Furthermore, dangerous situations that can be detected or predicted using image recognition (such as a pedestrian running into the roadway) can occur regardless of vehicle speed, and therefore there is room for improvement in the method of switching cameras based on vehicle speed (see Patent Document 1).
[0005] The present invention aims to improve the accuracy of image recognition processing for objects that require particular attention for ensuring safety. [Means for solving the problem]
[0006] The image recognition device of the present invention is an image recognition device that performs image recognition processing on target images belonging to images captured by a camera installed in a vehicle, and is equipped with a controller that identifies the level of danger of each object around the vehicle in relation to the vehicle based on the captured images, and adjusts the angle of view of the camera to obtain the target image depending on the distance between the vehicle and the object that corresponds to the greatest level of danger. [Effects of the Invention]
[0007] In order to perform image recognition processing appropriately (to improve the accuracy of image recognition), it is desirable to optimize the size of objects in the target image for image recognition. On the other hand, to ensure the safety of vehicles, pedestrians, etc., it is preferable to perform appropriate image recognition processing on objects that pose a relatively high risk rather than objects that pose a relatively low risk, and therefore it is preferable to optimize the size (image size) of the latter objects. According to the image recognition device of the present invention, the angle of view is adjusted according to the distance between the vehicle and the object (target object) that poses the greatest risk among the objects around the vehicle. Therefore, the size of the target object in the target image can be optimized according to the distance, which is expected to improve the accuracy of image recognition processing for high-risk objects. As a result, the safety of vehicles, pedestrians, etc. is promoted. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 2 is a diagram illustrating the relationship between a user and other components according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an internal configuration of an in-vehicle system according to an embodiment of the present invention. [Figure 3] 1A and 1B are diagrams illustrating the imaging areas of a wide-angle camera, a standard camera, and a telephoto camera according to an embodiment of the present invention. [Figure 4] 1A and 1B are diagrams illustrating the relationship between a scene involved in a photograph and a wide-angle camera image, a standard camera image, and a telephoto camera image obtained by photographing the scene, according to an embodiment of the present invention. [Figure 5] 1A and 1B are diagrams showing examples of two-dimensional images obtained by imaging according to an embodiment of the present invention. [Figure 6] 1 is an internal block diagram of an image recognition device according to an embodiment of the present invention. [Figure 7] 10 is a flowchart showing the operation of a controller according to Example EX1_A of the embodiment of the present invention. [Figure 8] FIG. 10 is a diagram showing an example of an input image according to Example EX1_A of the embodiment of the present invention. [Figure 9] FIG. 10 is a diagram showing recognition result information obtained by image recognition processing according to Example EX1_A of the embodiment of the present invention. [Figure 10] FIG. 10 is a diagram showing recognition result information obtained by image recognition processing and risk assessment information obtained by risk assessment processing according to Example EX1_A of the embodiment of the present invention. [Figure 11] FIG. 10 is a diagram showing the configuration of a risk determination table according to Example EX1_A of the embodiment of the present invention. [Figure 12] 10 is an explanatory diagram of a method for setting a target camera according to the distance between the target object and the vehicle, according to Example EX1_A belonging to an embodiment of the present invention. FIG. [Figure 13] FIG. 10 is a diagram showing a time series change in an image (target image) captured by a target camera according to Example EX1_A belonging to an embodiment of the present invention. [Figure 14] FIG. 10 is a diagram showing a partial image area of a target image according to Example EX1_C of the embodiment of the present invention. [Figure 15] FIG. 10 is a diagram showing an example EX2_A according to an embodiment of the present invention, in which a camera with a variable angle of view is provided in a camera block. [Figure 16] 10 is a flowchart showing the operation of a controller according to Example EX2_A of the embodiment of the present invention. [Figure 17] FIG. 10 is a diagram showing a sequence of captured images obtained from a target camera according to Example EX3 belonging to an embodiment of the present invention. [Figure 18]FIG. 10 is a diagram showing a view in front of a vehicle according to Example EX3 of the present invention. [Figure 19] FIG. 10 is a functional block diagram of a controller according to Example EX4 of the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, examples of embodiments of the present invention will be described in detail with reference to the drawings. In each of the drawings, identical parts are designated by the same reference numerals, and redundant descriptions of identical parts will be omitted as a general rule. For the sake of simplicity, this specification may use symbols or signs referring to information, signals, physical quantities, functional units, circuits, elements, or components, and may omit or abbreviate the names of the information, signals, physical quantities, functional units, circuits, elements, or components corresponding to the symbols or signs. For example, the wide-angle camera referred to by "CM1" (see FIG. 2) described below may be written as wide-angle camera CM1 or abbreviated as camera CM1, but these all refer to the same thing.
[0010] FIG. 1 shows the relationship between user U1 and other components assumed in an embodiment of the present invention. User U1 is an occupant of vehicle V1. User U1 is the driver of vehicle V1. Hereinafter, when simply referring to a driver, this refers to the driver of vehicle V1 (hence user U1). However, user U1 may also be an occupant other than the driver (i.e., a passenger in vehicle V1). Vehicle V1 is any type of vehicle. Here, vehicle V1 is assumed to be an automobile or the like that travels on a road. An in-vehicle system SYS is installed in vehicle V1.
[0011] A seat ST1 is installed in the cabin of the vehicle V1. A user U1 sits in the seat ST1. Since it is assumed that the user U1 is the driver, the seat ST1 is the driver's seat. Hereinafter, when simply referring to the cabin, unless otherwise specified, this refers to the cabin of the vehicle V1. Furthermore, below, unless otherwise specified, the inside of the vehicle refers to the internal area of the vehicle V1, and the outside of the vehicle refers to the external area of the vehicle V1.
[0012] The direction from the driver's seat of vehicle V1 toward the steering wheel is defined as "forward," and the direction from the steering wheel of vehicle V1 toward the driver's seat is defined as "rearward." The direction perpendicular to the front-to-rear direction and parallel to the road surface on which vehicle V1 is traveling is defined as the left-to-right direction. The direction perpendicular to the front-to-rear direction and perpendicular to the left-to-right direction is defined as the up-to-down direction. User U1 sits in seat ST1 facing forward. The front-to-rear direction, left-to-right direction, and up-to-down direction correspond to the front-to-rear direction, left-to-right direction, and up-to-down direction as seen from user U1. Unless otherwise specified below, vehicle V1 is assumed to be located on a horizontal road surface, and the traveling direction of vehicle V1 is assumed to be forward.
[0013] The world coordinate system is defined as the coordinate system of the real space (actual three-dimensional space) in which the vehicle V1 exists. The world coordinate system is a three-dimensional coordinate system with three mutually perpendicular axes: the WX-axis, the WY-axis, and the WZ-axis. The relationship between the WX-axis, the WY-axis, and the WZ-axis and the front-to-rear, left-to-right, and up-to-down directions is defined as follows: The WX-axis is parallel to the left-to-right direction. The WY-axis is parallel to the front-to-rear direction. The WZ-axis is parallel to the up-to-down direction. The direction from rear to front coincides with the direction from the negative side to the positive side of the WY-axis. The direction from bottom to top coincides with the direction from the negative side to the positive side of the WZ-axis. The direction from left to right coincides with the direction from the negative side to the positive side of the WX-axis. Here, it is assumed that the vehicle V1 is located on a horizontal road surface, so in real space the WX-axis and WY-axis are parallel to the horizontal plane (hence the road surface), and the WZ-axis is parallel to the vertical line.
[0014] The configuration of the in-vehicle system SYS is shown in Figure 2. The in-vehicle system SYS includes a camera block CB, an image recognition device 10, a vehicle control device 20, an actuator unit 30, a vehicle sensor unit 40, and an HMI 50. The components of the in-vehicle system SYS can exchange signals and information with each other through an in-vehicle network formed in the vehicle V1. The in-vehicle network includes, for example, a CAN (Controller Area Network) and an AVCLAN (Audio Visual Communication Local Area Network).
[0015] The camera block CB includes one or more unit cameras. In this embodiment, unless otherwise specified, it is assumed that the camera block CB includes three unit cameras: a wide-angle camera CM1, a standard camera CM2, and a telephoto camera CM3. However, a modification is also possible in which only two cameras out of the wide-angle camera CM1, the standard camera CM2, and the telephoto camera CM3 are provided in the camera block CB. Alternatively, the total number of unit cameras provided in the camera block CB may be four or more.
[0016] Each unit camera includes an optical system and an imaging element configured with a CMOS (Complementary Metal Oxide Semiconductor) image sensor or a CCD (Charge-Coupled Device) image sensor. Each unit camera captures images within its own imaging area (i.e., field of view). The image captured by each unit camera is called a captured image, and image data representing the captured image is called captured image data. The captured image is a two-dimensional image. Each unit camera periodically captures images within its own imaging area at a predetermined frame rate. Each unit camera generates and outputs captured image data each time it captures an image. The captured image data obtained by each unit camera is supplied to the image recognition device 10. Note that when multiple cameras are provided in the camera block CB, the multiple cameras are assumed to have the same frame rate. Therefore, the cameras CM1 to CM3 in FIG. 2 are assumed to have the same frame rate (however, modifications are possible in which the cameras CM1 to CM3 have different frame rates). The image captured by the wide-angle camera CM1 is particularly referred to as a wide-angle camera image. Images captured by the standard camera CM2 are specifically referred to as standard camera images. Images captured by the telephoto camera CM3 are specifically referred to as telephoto camera images. In this specification, obtaining a wide-angle camera image by shooting is sometimes referred to as shooting a wide-angle camera image. The same applies to standard camera images, telephoto camera images, etc.
[0017] The shooting areas of cameras CM1 to CM3 will be compared with reference to Figures 3(a) to 3(c). The hatched area SR1 shown in Figure 3(a) represents the shooting area of wide-angle camera CM1. The hatched area SR2 shown in Figure 3(b) represents the shooting area of standard camera CM2. The hatched area SR3 shown in Figure 3(c) represents the shooting area of telephoto camera CM3. The shooting areas of each of cameras CM1 to CM3 extend from the installation position of cameras CM1 to CM3 toward the front of vehicle V1. Therefore, each of shooting areas SR1 to SR3 is an area outside the vehicle located in front of vehicle V1. Cameras CM1 to CM3 are installed and fixed in appropriate positions on the body of vehicle V1. In view of the size of shooting areas SR1 to SR3, the difference between the installation positions of cameras CM1 to CM3 is sufficiently small that the difference can be considered to be essentially zero. In the following, unless otherwise required, it is assumed that the cameras CM1 to CM3 are installed at the same positions (more specifically, it is assumed that the optical centers of the cameras CM1 to CM3 are the same and that the positions of the imaging elements of the cameras CM1 to CM3 are the same).
[0018] The angle of view of the standard camera CM2 is larger than that of the telephoto camera CM3, and the angle of view of the wide-angle camera CM1 is even larger than that of the standard camera CM2. The angle of view of any unit camera includes a horizontal angle of view (horizontal angle of view) and a vertical angle of view (vertical direction). The horizontal angle of view corresponds to the angle of view in the WX axis direction, and the vertical angle of view corresponds to the angle of view in the WZ axis direction. As shown schematically in Figures 3(a) to 3(c), the horizontal angle of view of the standard camera CM2 is larger than that of the telephoto camera CM3, and the horizontal angle of view of the wide-angle camera CM1 is even larger than that of the standard camera CM2. However, a similar relationship applies to the vertical angle of view. That is, the vertical angle of view of the standard camera CM2 is larger than that of the telephoto camera CM3, and the vertical angle of view of the wide-angle camera CM1 is even larger than that of the standard camera CM2.
[0019] Scenery 600 in FIG. 4 is an arbitrary scene that spreads out in front of vehicle V1. However, scene 600 shown in FIG. 4 is a scene in three-dimensional space projected onto a plane parallel to the WX axis and the WZ axis. Assume that scene 600 is photographed by each of cameras CM1 to CM3. In FIG. 4, the outer edges of shooting areas SR1 to SR3 are shown superimposed on scene 600. In FIG. 4, two-dimensional image 610 is a wide-angle camera image obtained by photographing scene 600 with wide-angle camera CM1. In FIG. 4, two-dimensional image 620 is a standard camera image obtained by photographing scene 600 with standard camera CM2. In FIG. 4, two-dimensional image 630 is a telephoto camera image obtained by photographing scene 600 with telephoto camera CM3. Assume that images 610, 620, and 630 are photographed simultaneously.
[0020] The entire imaging region SR3 is included in the imaging region SR2, and the imaging region SR2 is larger than the imaging region SR3 in both the WX-axis direction and the WZ-axis direction. The entire imaging region SR2 is included in the imaging region SR1, and the imaging region SR1 is larger than the imaging region SR2 in both the WX-axis direction and the WZ-axis direction. In FIG. 4, it is assumed that the positions of the centers of the three regions obtained by projecting the imaging regions SR1 to SR3 onto a plane parallel to the WX-axis and WZ-axis coincide with each other. However, the positions of the centers of these three regions may differ from each other.
[0021] Any two-dimensional image is a collection of multiple pixels arranged in a matrix in the width and height directions. The width and height directions are orthogonal to each other. The image size of any two-dimensional image is expressed by the number of pixels in the width and height directions of the two-dimensional image. The number of pixels in the width and height directions of any wide-angle camera image (wide-angle camera image 610 in FIG. 4) are represented by W1 and H1, respectively. The number of pixels in the width and height directions of any standard camera image (standard camera image 620 in FIG. 4) are represented by W2 and H2, respectively. The number of pixels in the width and height directions of any telephoto camera image (telephoto camera image 630 in FIG. 4) are represented by W3 and H3, respectively. Here, it is assumed that the size of the image sensor of wide-angle camera CM1, the size of the image sensor of standard camera CM2, and the size of the image sensor of telephoto camera CM3 are all the same. Therefore, the image size of the wide-angle camera image, the image size of the standard camera image, and the image size of the telephoto camera image are all equal in the width and height directions, i.e., "W1 = W2 = W3" and "H1 = H2 = H3" are true (however, variations in which these equations do not hold are possible).
[0022] For this reason, when an object OBJ that fits within the scenery 600 and is located within the shooting region SR3 is simultaneously photographed by cameras CM1 to CM3, the image size of the object OBJ will be as follows: That is, the image size of the object OBJ in the standard camera image 620 is larger than the image size of the object OBJ in the wide-angle camera image 610. In addition, the image size of the object OBJ in the telephoto camera image 630 is even larger than the image size of the object OBJ in the standard camera image 620. Note that in any two-dimensional image, the image size of an object can also be expressed as the size of the object.
[0023] In the following, the image coordinate system in which an arbitrary two-dimensional image is defined is defined as follows. The image 650 shown in FIG. 5 is an arbitrary two-dimensional image. The image coordinate system is a two-dimensional coordinate system having two axes, the X-axis and the Y-axis, which are orthogonal to each other. The X-axis direction corresponds to the width direction, and the Y-axis direction corresponds to the height direction. The X-axis is parallel to the horizontal direction of the two-dimensional image 650, and the Y-axis is parallel to the vertical direction of the two-dimensional image 650. In the two-dimensional image 650, the direction toward the positive side of the X-axis is the rightward direction, and the direction toward the negative side of the X-axis is the leftward direction. The rightward and leftward directions in the two-dimensional image 650 correspond to the rightward and leftward directions as seen by the driver of the vehicle V1. Therefore, with respect to an arbitrary target object located around the vehicle V1, the more the target object moves to the right as seen by the driver of the vehicle V1, the more the position of the target object in the two-dimensional image 650 moves rightward. Conversely, as the target object moves further left as seen by the driver of the vehicle V1, the position of the target object in the two-dimensional image 650 moves further left.
[0024] The components other than the camera block CB shown in FIG. 2 will be described below. The image recognition device 10 performs image recognition processing on images captured by any of the cameras in the camera block CB. Details of the image recognition processing will be described later. The vehicle control device 20 controls the driving of the vehicle V1 using an actuator unit 30. The actuator unit 30 has various driving components such as a motor that realizes the driving of the vehicle V1. The actuator unit 30 includes an engine and a motor that generate driving force for the vehicle V1, a steering actuator that drives the steering of the vehicle V1, and a brake actuator that drives the brakes of the vehicle V1.
[0025] The vehicle sensor unit 40 includes various sensors, including sensors that detect the details of the driving operation of the vehicle V1 by the driver of the vehicle V1 and sensors that detect various states of the vehicle V1, and outputs vehicle sensor information containing the detection results of each sensor. The vehicle control device 20 realizes driving control of the vehicle V1 by driving and controlling the actuator unit 30 in accordance with the vehicle sensor information. At this time, the vehicle control device 20 can also perform driving control in accordance with the results of image recognition processing by the image recognition device 10.
[0026] The HMI 50 is a human-machine interface. The HMI 50 is provided with a display device 51 and a speaker 52. The display device 51 has a display screen such as a liquid crystal display panel, and displays any image under the control of the image recognition device 10, the vehicle control device 20, or a display control device (not shown). The display device 51 is installed at an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can see the display content of the display device 51. Multiple display devices 51 may be installed in the cabin of the vehicle V1. The display device 51 may be a component of a car navigation system installed in the vehicle V1. The car navigation system may be included in the in-vehicle system SYS. The speaker 52 outputs any sound (message, warning sound, music, etc.) under the control of the image recognition device 10, the vehicle control device 20, or an audio device (not shown). The speaker 52 is installed at an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can hear the output sound of the speaker 52. A plurality of speakers 52 may be installed in the cabin of the vehicle V1. In addition, the HMI 50 may be provided with a vibration device that applies vibrations to the occupants (particularly the driver) of the vehicle V1, an operation input device that accepts any operation (including voice operation) from the occupants of the vehicle V1, and the like.
[0027] 6 shows the internal configuration of the image recognition device 10. The image recognition device 10 includes a controller 11, a memory 12, and a communication unit 13.
[0028] The controller 11 includes, as hardware resources, a processing unit including a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), etc. The controller 11 may implement any function, operation, or process that should be implemented by the controller 11 by executing a program recorded in the memory 12 or any other recording medium (not shown).
[0029] The memory 12 is configured to include a non-volatile memory such as a ROM (Read Only Memory) or a flash memory, and a volatile memory such as a RAM (Random Access Memory). The memory 12 stores various data referenced by the controller 11 as well as various programs to be executed by the controller 11.
[0030] The communication unit 13 transmits and receives any signal to and from a counterpart device different from the image recognition device 10. The counterpart device for the communication unit 13 includes components of the in-vehicle system SYS shown in FIG. 2 other than the image recognition device 10. The communication unit 13 can communicate with the counterpart device via an in-vehicle network formed in the vehicle V1. The counterpart device for the communication unit 13 may further include an external device (e.g., a server device connected to the Internet) provided outside the vehicle V1. Note that the controller 11 can transmit and receive any information to and from the counterpart device using the communication unit 13, but the description of the communication unit 13 may be omitted below.
[0031] The controller 11 sets one of the unit cameras in the camera block CB as a target camera and performs image recognition processing on the image captured by the target camera. The image captured by the target camera to which the image recognition processing is applied is hereinafter referred to as a target image I. TG Therefore, the target image I TG belongs to the image captured by camera CM1, CM2, or CM3. If wide-angle camera CM1 is the target camera, target image I TG is the wide-angle camera image. If the standard camera CM2 is the target camera, the target image I TG is the standard camera image. If the telephoto camera CM3 is the target camera, the target image I TG is a telephoto camera image. The controller 11 converts all of the images captured by the target camera, which are generated sequentially in time series, into target image I TG You can treat it as a target image I, or you can treat only a part of the images taken by the target camera as a target image I. TG It may be treated as such.
[0032] Image recognition processing is performed on the target image ITG When focusing only on the object detection process, the image recognition device 10 can be said to be an object detection device. Since the cameras CM1 to CM3 all capture the area surrounding the vehicle V1, the target image I TG The objects within are objects located around the vehicle V1 (hereinafter, may be referred to as objects around the vehicle).
[0033] In order to perform image recognition processing properly, the target image I TG It is desirable to optimize the size of the object in the target image I. TG The wide-angle camera image, standard camera image, or telephoto camera image is used as the target image I so that the size of the object in the image is appropriate. TG However, if there are multiple objects around the vehicle V1, it is necessary to select and set the target image I TG (Which of the cameras CM1 to CM3 is the target image I?) TG Some ingenuity will be required as to whether or not to obtain a license.
[0034] Below, several examples of operation, application techniques, modified techniques, etc. related to the device will be described in multiple embodiments. The matters described above in this embodiment are applied to each of the following embodiments unless otherwise specified and unless there is a contradiction. If there are any matters in each embodiment that contradict the matters described above, the description in each embodiment may take precedence. Furthermore, unless there is a contradiction, matters described in any of the multiple embodiments shown below can also be applied to any other embodiment (i.e., any two or more of the multiple embodiments can be combined).
[0035] <<Example EX1_A>> In the embodiment EX1_A and the embodiments EX1_B and EX1_C described below, it is assumed that a wide-angle camera CM1, a standard camera CM2, and a telephoto camera CM3 are provided in the camera block CB, as shown in FIG.
[0036] FIG. 7 is an operation flowchart of the controller 11 according to the embodiment EX1_A. Each process of steps S10 to S17 shown in FIG. 7 is executed by the controller 11. When the vehicle V1 starts, the controller 11 also starts, and the operation of the controller 11 starts from the process of step S10. In step S10, the controller 11 executes an initialization process. In the initialization process, the controller 11 sets a predetermined one of the cameras CM1 to CM3 as the target camera. In order to enable detection of objects over a wider range in the image recognition process, in step S10, the wide-angle camera CM1 is set as the target camera. However, in step S10, the camera CM2 or CM3 may also be set as the target camera. Also, in step S10, the controller 11 assigns 1 to a variable i that it manages. After step S10, the process proceeds to step S11.
[0037] As described above, the image captured by the target camera and to which the image recognition process is applied is the target image I TG The image captured by the target camera at time t[i] is called the target image I TG The time t[i] is referred to as [i]. For any variable i, the time t[i+1] is one control cycle later than the time t[i]. One control cycle has a fixed time length and may be an integer multiple (including 1) of the above-mentioned shooting frame rate. As will be clear from the explanation below, a series of loop processes consisting of steps S11 to S17 are repeatedly executed. In the embodiment EX1_A, the execution interval between two loop processes executed adjacently in time corresponds to one control cycle. The time t[i] is understood as a concept having a time length of one control cycle or less, and therefore the time t[i] may be read as the ith unit period. At the time t[i] (ith unit period), the target image I TG [i] Shooting and target image I TG Various processes based on [i] (processing of steps S11 to S17) are executed.
[0038] In step S11, the controller 11 calculates the image captured by the target camera at time t[i], i.e., the target image I TGSet [i] to the input image IN, and the input image IN (hence the target image I TG [i]) is subjected to image recognition processing. The image recognition processing for the input image IN is performed based on the image data of the input image IN. The image recognition processing includes object detection processing, area detection processing, and state determination processing.
[0039] The object detection process for the input image IN will be described below. An object to be recognized in the object detection process is referred to as a recognition target object. The recognition target object includes at least a person, and may also include objects other than a person (e.g., a dog, a cat, a bicycle, a car, a kick scooter, or a traffic light). In the object detection process, the type of the recognition target object in the input image IN is detected, and a rectangular area in the input image IN in which the recognition target object is determined to exist is set as a bounding box. Any object in the following description may be understood to belong to the recognition target object unless otherwise specified. The bounding box will be referred to as a BBOX below. Furthermore, a person located near a vehicle may be referred to as a pedestrian, etc. below. Pedestrians, etc. include not only pedestrians, but also people moving or stationary on bicycles, kick scooters, etc.
[0040] Objects to be recognized in the input image IN are detected by the object detection process, and a BBOX is set for each detected object to be recognized. After setting a BBOX for each object to be recognized in the object detection process, the controller 11 generates BBOX information that specifies the position and shape of the BBOX in the input image IN. Object detection AI that realizes object detection processing has been put to practical use, and the object detection AI may be included in the controller 11. Note that AI is an abbreviation for artificial intelligence. Any one BBOX is referred to as a BBOX of interest. The BBOX information for the BBOX of interest indicates the type of object to be recognized corresponding to the BBOX of interest, as well as the origin coordinates, width, and height of the BBOX of interest in the input image IN. The origin coordinates of the BBOX of interest refer to the coordinates of one of the four vertices of the rectangular outline of the BBOX of interest. The width of the BBOX of interest is the length of the BBOX of interest in the X-axis direction and is expressed by the number of pixels of the BBOX of interest in the X-axis direction. The height of the BBOX of interest is the length of the BBOX of interest in the Y-axis direction and is expressed by the number of pixels of the BBOX of interest in the Y-axis direction.
[0041] The region detection process for the input image IN will now be described. In the region detection process, for each object present in the input image IN, the controller 11 detects the region in which the image of the object is located (in other words, the region in which the image data of the object is located) in pixel units. The controller 11 generates region detection information indicating the results of the detection in the region detection process. Therefore, for example, if the input image IN includes images of a roadway, a sidewalk, and a person, the region detection process detects the roadway region in which the image of the roadway is located, the sidewalk region in which the image of the sidewalk is located, and the person region in which the image of the person is located in pixel units. In this case, if the input image IN includes images of multiple people, a person region is detected for each person. The same applies to sidewalks, etc. The region detection process may be semantic segmentation, and a known AI (semantic segmentation AI) that realizes semantic segmentation may be included in the controller 11.
[0042] The state determination process for the input image IN will now be described. In the state determination process, the controller 11 determines the state of each recognition target object detected by the object detection process. In this determination, the result of the area detection process is referenced. The controller 11 generates state determination information indicating the result of the determination in the state determination process.
[0043] In step S11, the controller 11 generates recognition result information indicating the result of the image recognition process. The recognition result information includes BBOX information for each object to be recognized and state determination information for each object to be recognized, as well as area detection information.
[0044] FIG. 8 shows input image IN1, which is an example of input image IN, and will explain the image recognition process for input image IN1. The shooting area of the camera that captures input image IN1 by shooting includes roadway RD located in front of vehicle V1, sidewalk SW_L located on the left side of roadway RD, and sidewalk SW_R located on the right side of roadway RD. Therefore, input image IN1 includes images of roadway RD, sidewalk SW_L, and sidewalk SW_R. Note that the inclusion of an image of roadway RD in input image IN1 means, in other words, that input image IN1 includes image data of roadway RD. The same applies to sidewalks SW_L and SW_R and any objects (including people) described below.
[0045] The input image IN1 includes images of three objects 711 to 713. The objects 711 to 713 are each a person. At the time of capturing the input image IN1, the object 711 was located on the sidewalk SW_L, while the objects 712 and 713 were located on the roadway RD. In detail, at the time of capturing the input image IN1, the object 712 was a person crossing the vehicle RD while looking straight ahead and facing right, and the object 713 was a person crossing the vehicle RD while looking at an information terminal (such as a smartphone) held in his / her hand and facing left.
[0046] In the object detection process for input image IN1, a BBOX is set for each object to be recognized. Therefore, for input image IN1, BBOXes 721, 722, and 723 are set for objects 711, 712, and 713, respectively. Furthermore, by performing area detection process for input image IN1, the roadway area where the image of roadway RD is located, the sidewalk area where the images of sidewalks SW_L and SW_R are located, and the person area where the images of objects 711 to 713 are located are detected in pixel units. The controller 11 determines the state of each of the objects 711 to 713 based on the results of the object detection process and area detection process for input image IN1.
[0047] Specifically, the controller 11 involved in the state determination process determines that the person as object 711 is in a sidewalk-staying state. The sidewalk-staying state is a state in which the person is walking or standing still on the sidewalk (see FIG. 9). The controller 11 involved in the state determination process determines that the person as object 712 is in a type 1 crossing state. The type 1 crossing state is a state in which the person is crossing the roadway while looking straight ahead (see FIG. 9). The controller 11 involved in the state determination process determines that the person as object 713 is in a type 2 crossing state. The type 2 crossing state is a state in which the person is crossing the roadway while looking at the information terminal they are holding in their hand (see FIG. 9).
[0048] In step S11, the controller 11 stores the generated recognition result information in the memory 12. FIG. 9 shows an example of recognition result information 700 generated by the image recognition process for the input image IN1. TG The recognition result information when [i] is the input image IN is stored in memory 12 in association with time t[i]. Therefore, the recognition result information 700 is stored in memory 12 in association with time t[i]. In the recognition result information 700, an object ID is assigned to each recognition target object. In the recognition result information 700, for each object ID, the type of object, the position and size of the object in the input image IN (here, IN1), and the state of the object (the state determined by the state determination process) are indicated. In the recognition result information 700, the object ID of "001" corresponds to object 711, the object ID of "002" corresponds to object 712, and the object ID of "003" corresponds to object 713.
[0049] In step S12 following step S11, the controller 11 performs a risk determination process to determine and derive the risk of each object around the vehicle based on the recognition result information generated in step S11. The object whose risk is to be determined is a recognition target object detected in the object detection process of step S11. Assume that first to mth objects have been detected as recognition target objects in the object detection process of step S11. The risk derived for the jth object is referred to as risk DG[j]. Here, m represents an arbitrary integer equal to or greater than 2, and j represents a natural number equal to or less than m. In step S12, the controller 11 generates risk determination information including the derived risk levels DG[1] to DG[m], and stores the generated risk determination information in memory 12 in association with the recognition result information.
[0050] A certain risk level DG[j] is the risk level (degree of risk) of the jth object in relation to the vehicle V1. In detail, the risk level DG[j] represents the risk level of the vehicle V1 colliding with the jth object in the near future (the possibility that the vehicle V1 will collide with the jth object). The higher the risk level DG[j], the greater the risk level of the vehicle V1 colliding with the jth object in the near future (the possibility that the vehicle V1 will collide with the jth object). A person around the vehicle is an object whose risk level is to be determined. The object whose risk level is to be determined may be something other than a person (for example, a dog or a vehicle other than the vehicle V1), but hereinafter it is assumed that the object whose risk level is to be determined is a person.
[0051] 10 shows the above-mentioned recognition result information 700 and risk assessment information 702 generated by the risk assessment process for input image IN1. The risk assessment information 702 is stored in memory 12 in association with time t[i] and recognition result information 700. The risk assessment information 702 indicates a risk DG[1] for an object 711 on input image IN1, a risk DG[2] for an object 712 on input image IN1, and a risk DG[3] for an object 713 on input image IN1.
[0052] The controller 11 may determine the level of danger of each object using a danger level determination table stored in advance in the memory 12. Table TBL1 in FIG. 11 is an example of a danger level determination table. In the danger level determination table, a danger level is associated with each state determined in the state determination process. In table TBL1 in FIG. 11, a danger level of "1" is associated with the sidewalk stay state, a danger level of "3" is associated with the first-class crossing state, and a danger level of "5" is associated with the second-class crossing state. Therefore, if table TBL1 is used when input image IN1 in FIG. 8 is the input image IN in step S11, the danger levels DG[1], DG[2], and DG[3] are determined to be 1, 3, and 5, respectively, based on the recognition result information 700 (see FIG. 10).
[0053] In practice, for example, in the state determination process, the controller 11 determines to which of the first to pth classes the state of each person in the input image IN belongs. Here, p represents an integer of 2 or greater. The risk determination table defines the risk level for each of the first to pth classes. The sidewalk stay state, type 1 crossing state, and type 2 crossing state shown in FIG. 9 correspond to three classes among the first to pth classes. When the controller 11 determines that the state of a certain person of interest belongs to the qth class, it determines that the risk level of the person of interest is the risk level defined in association with the qth class in the risk determination table. q represents a natural number equal to or less than p.
[0054] After step S12, the process proceeds to step S13. In step S13, the controller 11 determines the maximum risk DG among the risk levels DG[1] to DG[m]. MAX Identify the maximum risk DG in the input image IN. MAX Check whether there is one object corresponding to the maximum risk DG in the input image IN. MAX If there is only one object corresponding to the maximum risk DG in the input image IN (Y in step S13), proceed to step S14. MAX If there are multiple objects that correspond to the object (N in step S13), proceed to step S15.
[0055] In step S14, the controller 11 determines the maximum risk level DG MAX After step S14, the controller 11 sets the object corresponding to the maximum danger level DG as the target TG. MAX From among the plurality of objects corresponding to the target object TG, one object is selected and set as the target object TG. The selection in step S15 is performed in accordance with predetermined selection rules. The selection rules will be described in the following examples. After step S15, the process proceeds to step S16.
[0056] In step S16, the controller 11 calculates the distance d between the object TG and the vehicle V1. TG Based on the distance d, the target camera is selected and set from among the cameras CM1 to CM3. TG is the distance between the object TG and the vehicle V1 in real space (more specifically, for example, the shortest distance between the object TG and the body of the vehicle V1). In setting the target camera in step S16, the target camera may be changed from one camera to another, or the target camera may be maintained as a certain camera.
[0057] Referring to Figure 12, distance d TG The controller 11 in step S16 sets the target camera in accordance with the TG ≦d TH1 If the condition "d" is satisfied, the wide-angle camera CM1 is set as the target camera. TH1 <d TG ≦d TH2 When the condition "d" is satisfied, the standard camera CM2 is set as the target camera. TH2 <d TG When the condition "is satisfied," the telephoto camera CM3 is set as the target camera. TH1 and d TH2 is a predetermined threshold distance, and "0 <d TH1 <d TH2 " holds true.
[0058] The controller 11 calculates the distance d according to triangulation based on the camera installation information, camera parameters, and captured image data (i.e., image data of the input image IN) of the target camera in step S11. TG The method of deriving the distance between the object on the captured image and the vehicle V1 by triangulation based on the camera installation information, camera parameters, and captured image data is well known. The position where the object TG touches the ground on the input image IN is identified from the image data of the input image IN (if the object 713 in FIG. 8 is the object TG, the position where the bottom of the foot of the person as the object 713 touches the roadway RD). The distance d TG can be derived. The memory 12 stores in advance the camera installation information and camera parameters of each camera provided in the camera block CB. The camera installation information and camera parameters of camera CM1 will be described. The camera installation information of camera CM1 indicates the installation position of camera CM1 relative to vehicle V1 and the mounting angle of camera CM1. The installation position of camera CM1 in the camera installation information of camera CM1 indicates the height of camera CM1 from the road surface (the bottom end of vehicle V1) and the distance between the front end of the body of vehicle V1 and camera CM1. The mounting angle of camera CM1 indicates the depression angle or elevation angle of camera CM1 with respect to the horizontal plane. The mounting angle of camera CM1 may also represent the Euler angles (pitch angle, roll angle, and yaw angle) of camera CM1. The camera parameters of camera CM1 are information for specifying the camera coordinate system for camera CM1 together with the camera installation information of camera CM1, such as the focal length and size of the image sensor of camera CM1. The camera installation information and camera parameters of camera CM1 can be used to determine the real-space distance between any point on the ground (such as the road surface or sidewalk) in the wide-angle camera image and vehicle V1. The same applies to the camera installation information and camera parameters of cameras other than camera CM1.
[0059] Alternatively, the controller 11 may calculate the distance d based on the camera parameters of the target camera and the size of the target TG on the input image IN. TG This method of derivation can be roughly calculated by the distance dTG This method is based on the fact that the size (image size) of the object TG on the input image IN decreases or increases as the size increases or decreases, and can be used when the object TG is a person (i.e., when the size of the object TG is known approximately).
[0060] In the case where a distance measurement sensor (not shown) is provided in the vehicle sensor unit 40, the distance d TG In the distance measurement, the distance measurement sensor detects the distance between the vehicle V1 and a three-dimensional object (a three-dimensional object) located around the vehicle V1, and also detects the orientation of the three-dimensional object as viewed from the vehicle V1. The distance d can be calculated by combining the detection result of the distance measurement by the distance measurement sensor and the result of the object detection process in the image recognition process. TG The distance measurement sensor is configured by a LIDAR (Light Detection and Ranging) that measures distance using light, a radar that measures distance using radio waves, or a combination of these.
[0061] After step S16, the process proceeds to step S17. In step S17, the controller 11 adds 1 to the variable i. After step S17, the process returns to step S11, and the processes from step S11 onwards are repeated.
[0062] Consider case CS1 in which input image IN1 in Fig. 8 is the input image IN in step S11 and table TBL1 in Fig. 11 is used as the risk determination table. In case CS1, the risks DG[1], DG[2], and DG[3] for objects 711, 712, and 713 are determined to be 1, 3, and 5, respectively (step S12).
[0063] In case CS1, the maximum risk DG MAX is "5" in the risk level DG[3], the maximum risk level DG MAX Therefore, the process proceeds from step S13 to step S14, and the object 713 is set as the target object TG. As a result, the distance between the object 713 and the vehicle V1 is set as the distance dTG In step S16, the target camera is set according to the distance between the object 713 and the vehicle V1.
[0064] In case CS1, at time t[i A ], in the case where the target camera at the stage of step S11 is the wide-angle camera CM1 (i.e., the target image I TG [i A ] is a wide-angle camera image) is called case CS1a. A represents any natural number and has the value of variable i at a certain timing. A ] distance d TG is the time t[i A ] is the distance (distance in real space) between the object 713 and the vehicle V1.
[0065] Case CS1a is the time t[i A ] distance d TG In case CS1a, the process branches to the following cases CS1a_1, CS1a_2, or CS1a_3 depending on the time t[i A ] distance d TG "d TG ≦d TH1 In case CS1a_1 that satisfies ", the target camera is maintained as the wide-angle camera CM1 in step S16. In case CS1a_1, at time t[i A +1], the input image IN (i.e., the target image I TG [i A +1]) is time t[i A +1]. In case CS1a, the wide-angle camera image is taken by the wide-angle camera CM1 at time t[i A ] distance d TG "d TH1 <d TG ≦d TH2 In case CS1a_2 that satisfies ", the target camera is changed to the standard camera CM2 in step S16. In case CS1a_2, at time t[i A +1], the input image IN (i.e., the target image I TG [i A +1]) is time t[i A+1] is the standard camera image taken by the standard camera CM2. In case CS1a, at time t[i A ] distance d TG "d TH2 <d TG In case CS1a_3, which satisfies ", the target camera is changed to the telephoto camera CM3 in step S16. In case CS1a_3, at time t[i A +1], the input image IN (i.e., the target image I TG [i A +1]) is time t[i A +1]. A ], we have considered a case CS1a in which the target camera at step S11 is the wide-angle camera CM1, but the same applies when the target camera is the camera CM2 or CM3.
[0066] In case CS1a, at time t[i A ], the object 713 continues to be set as the target TG and the distance d TG "d TH1 <d TG ≦d TH2 Target image I when continuing to satisfy " TG The sequence of is shown in Figure 13. In the example of Figure 13, the target camera is A ] and t[i A +1], the wide-angle camera CM1 is switched to the standard camera CM2, and thereafter, at least at time t[i A +2] and t[i A +3] and the standard camera CM2 is maintained.
[0067] Based on the recognition result information obtained by the image recognition processing, the image recognition device 10 (controller 11) or the vehicle control device 20 can perform recognition response processing that contributes to ensuring the safety of the vehicle V1, pedestrians, etc. Recognition result information is generated sequentially during the operation shown in FIG. 7 , and the image recognition device 10 (controller 11) or the vehicle control device 20 can perform recognition response processing based on the latest recognition result information. In the recognition response processing, the image recognition device 10 (controller 11) or the vehicle control device 20 can issue a notification to notify the driver of the vehicle V1 of the presence of the object TG. The notification is realized using a video display by the display device 51 and an audio output from the speaker 52. In the recognition response processing, the vehicle control device 20 may perform driving control (for example, driving control to stop or decelerate the vehicle V1) to reliably avoid a collision between the object TG and the vehicle V1.
[0068] In the embodiment EX1_A, as described above, the risk of each object around the vehicle V1 is derived and identified based on the captured image of the camera (CM1, CM2, or CM3) in the camera block CB (S11 and S12). MAX The object corresponding to the distance d between the object TG and the vehicle V1 is set as the object TG (S13 to S15). TG Depending on the target image I TG The angle of view of the camera (target camera) is adjusted to obtain the target image I TG The camera (target camera) to obtain the TG This is realized by selecting from the cameras CM1 to CM3 according to the situation (S16).
[0069] In order to properly perform image recognition processing (to improve the accuracy of image recognition processing), target image I TGOn the other hand, in order to ensure the safety of the vehicle V1 or pedestrians, it is preferable to perform appropriate image recognition processing on objects with a relatively high risk level rather than objects with a relatively low risk level, and therefore it is preferable to optimize the size (image size) of the latter objects. According to Example EX1_A, among the objects around the vehicle, the object with the highest risk level DG MAX The object corresponding to the target TG is set as the target TG, and the distance d between the target TG and the vehicle V1 is TG The angle of view is adjusted according to the distance d TG Depending on the target image I TG This is expected to improve the accuracy of image recognition processing for high-risk objects TG, thereby promoting the safety of the vehicle V1 and pedestrians.
[0070] Specifically, the controller 11 receives an image (I TG ) and determine the risk level (DG[1] to DG[m]) of each object around the vehicle based on the risk level (DG[1] to DG[m]) to set the object TG (S12 to S15). Then, the distance d between the object TG and the vehicle V1 is calculated. TG In response to the request, a target camera at a second time point that is later than the first time point is selected from the plurality of cameras (CM1 to CM3) (S16).
[0071] This allows the angle of view of the target camera at the second time to be appropriate for recognizing the target object TG. Therefore, it is expected that the accuracy of the image recognition process for the highly dangerous target object TG will be improved (the accuracy improvement at the second time), thereby promoting the safety of the vehicle V1 or pedestrians, etc. Note that the first time and the second time are, for example, the time t[i A ] and t[i A +1] (see Figure 13).
[0072] More specifically, the controller 11 calculates the distance d TG is the threshold distance (d TH1 or d TH2), the controller 11 sets the target camera at a second time point after the first time point to the first camera having the first angle of view. TG is the threshold distance (d TH1 or d TH2 ), the target camera at the second time point is set to a second camera having a second angle of view narrower than the first angle of view.
[0073] This allows the distance d TG If the distance d is relatively small, the object TG is photographed by the first camera with a relatively large angle of view, and the size of the object TG in the image of the object can be optimized for image recognition (the size of the object TG may be too large in the image photographed by the second camera). TG is relatively large, the object TG is photographed by the second camera with a relatively small angle of view, and the size of the object TG on the object image can be optimized for image recognition (the size of the object TG may be too small on the image photographed by the first camera). Note that in Example EX1_A, the first and second cameras are two of the cameras CM1 to CM3, for example, the wide-angle camera CM1 and the standard camera CM2.
[0074] <<Example EX1_B>> In the example EX1_B, the maximum risk DG MAX In step S13, it is assumed that there are two or more objects corresponding to the maximum danger level DG MAX Each of the two or more objects corresponding to each of the two or more objects is referred to as a candidate object. In step S15, one candidate object is selected as the target object TG from among the plurality of candidate objects in accordance with a selection rule. A method for selecting the target object TG in accordance with the selection rule will be described. For the sake of concreteness of explanation, the plurality of candidate objects will be referred to as first to Kth candidate objects (K is an integer of 2 or more).
[0075] In step S15, the controller 11 derives the distance (distance in real space) between each candidate object and the vehicle V1. The method for deriving the distance between the candidate object and the vehicle V1 is the same as the method for deriving the distance between the object TG and the vehicle V1 described in embodiment EX1_A. When applying the method for deriving the distance between the object TG and the vehicle V1 described in embodiment EX1_A to the method for deriving the distance between the candidate object and the vehicle V1, the object TG in the description of embodiment EX1_A can be read as the candidate object. For any integer i, the distance derived for the ith candidate object is referred to as the distance d[i]. Then, the controller 11 selects the object TG from the first to Kth candidate objects based on at least the distances d[1] to d[K].
[0076] Based on the distance between each candidate object and the vehicle V1, it is possible to estimate the remaining time until the vehicle V1 approaches each candidate object, and to set the candidate object with a higher urgency of response as the target object TG. As a result, it is possible to improve the accuracy of image recognition processing for candidate objects with a higher urgency of response.
[0077] Specifically, the controller 11 may identify the shortest distance among the distances d[1] to d[K], and set the candidate object corresponding to the shortest distance as the target TG. This is because the shorter the distance to the candidate object, the shorter the time remaining until the approach of the vehicle V1 is considered to be, and the shorter the time remaining until the approach of the vehicle V1, the more urgent the response. In reality, contact between the vehicle V1 and the candidate object is avoided almost 100%, but the approach of the vehicle V1 to the candidate object may be interpreted as referring to contact (collision) of the vehicle V1 with the candidate object.
[0078] The controller 11 may select the target object TG from among the first to Kth candidate objects based on the distances d[1] to d[K] and the relative velocities V[1] to V[K]. For any integer i, the relative velocity V[i] represents the relative velocity between the vehicle V1 and the i-th candidate object in the WY axis direction. The controller 11 selects a plurality of target images I arranged in time series. TGThe controller 11 can detect the relative velocity V[i] based on the movement vector of the ith candidate object in the WY direction. The controller 11 can estimate the remaining time until the distance between the vehicle V1 and the ith candidate object in the WY direction becomes zero based on the distance d[i] and the relative velocity V[i]. The controller 11 can perform this estimation for each candidate object and select the candidate object corresponding to the shortest remaining time among the remaining times estimated for the first to Kth candidate objects as the target object TG. In this selection, the predicted travel path of the vehicle V1 based on the steering angle of the vehicle V1 can be taken into consideration.
[0079] <<Example EX1_C>> Example EX1_C will be described. In Example EX1_C, several modified techniques applicable to Examples EX1_A and EX1_B will be described.
[0080] 14(a) and (b). The case where the target camera at time t[i] is the wide-angle camera CM1 is referred to as case CS2a. The target image I TG [i] is a wide-angle camera image.
[0081] In case CS2a, the controller 11 TG Based on the BBOX information for [i], target image I TG It is determined whether the entire object TG (the entire image of the object TG) on [i] is located within the image area 761 (see FIG. 14(a)). TG Only when the entire object TG on [i] is located within the image area 761, the controller 11 in case CS2a permits the target camera to be switched to the telephoto camera CM3. TG When all or part of the object TG on [i] is located outside the image area 761, the controller 11 according to the case CS2a determines the distance d TG When this prohibition is performed, the controller 11 in case CS2a prohibits switching of the target camera to the telephoto camera CM3 regardless of the TG ≦d TH1 ", set the wide-angle camera CM1 as the target camera, and "dTH1 <d TG ", then the standard camera CM2 is set as the target camera.
[0082] In case CS2a, the controller 11 TG Based on the BBOX information for [i], target image I TG It is determined whether the entire object TG (the entire image of the object TG) on [i] is located within the image area 762 (see FIG. 14(b)). TG Only when the entire object TG on [i] is located within the image area 762, the controller 11 in case CS2a permits the target camera to be switched to the standard camera CM2. TG When all or part of the object TG on [i] is located outside the image area 762, the controller 11 according to the case CS2a determines the distance d TG When this prohibition is performed, the controller 11 in case CS2a prohibits the target camera from being switched to the standard camera CM2 or the telephoto camera CM3 regardless of the distance d TG Regardless of the above, the target camera remains the wide-angle camera CM1.
[0083] The image areas 761 and 762 in case CS2a are the target image I captured by the wide-angle camera CM1. TG In case CS2a, the image area 762 is larger than the image area 761 in both the X-axis and Y-axis directions, and the entire image area 761 is included in the image area 762. TGThe center position of the entire image area [i], the center position of the image area 761, and the center position of the image area 762 may be the same or different from one another. The image area 761 is an image area corresponding to the shooting area SR3 of the telephoto camera CM3 (see FIG. 4), and objects outside the image area 761 are likely to protrude from or be outside the shooting area SR3. The image area 762 is an image area corresponding to the shooting area SR2 of the standard camera CM2 (see FIG. 4), and objects outside the image area 762 are likely to protrude from or be outside the shooting area SR2. For this reason, the above prohibition is performed to prevent the object TG from going out of the frame. The positions, shapes, and sizes of the image areas 761 and 762 are set in advance based on the positional and dimensional relationships between the shooting areas SR1 to SR3, and these relationships are specified by the camera installation information, camera parameters, etc. of the cameras CM1 to CM3.
[0084] See Figure 14(c). The case where the target camera at time t[i] is the standard camera CM2 is called case CS2b. The target image I TG [i] is a standard camera image.
[0085] In case CS2b, the controller 11 TG Based on the BBOX information for [i], target image I TG It is determined whether the entire object TG (the entire image of the object TG) on [i] is located within the image area 763 (see FIG. 14(c)). TG Only when the entire object TG on [i] is located within the image area 763, the controller 11 in case CS2b permits the target camera to be switched to the telephoto camera CM3. TG When all or part of the object TG on [i] is located outside the image area 763, the controller 11 according to the case CS2b determines the distance d TG When this prohibition is performed, the controller 11 in case CS2b prohibits switching of the target camera to the telephoto camera CM3 regardless of the TG ≦d TH1", set the wide-angle camera CM1 as the target camera, and "d TH1 <d TG ", then the standard camera CM2 is set as the target camera.
[0086] The image area 763 in case CS2b is the target image I captured by the standard camera CM2. TG In case CS2b, the target image I TG The center position of the entire image area [i] and the center position of the image area 763 may or may not coincide with each other. The image area 763 is an image area corresponding to the shooting area SR3 of the telephoto camera CM3 (see FIG. 4), and objects outside the image area 763 are likely to protrude from the shooting area SR3. For this reason, the above prohibition is performed to prevent the target object TG from going out of the frame. The position, shape, and size of the image area 763 are set in advance based on the positional relationship and size relationship between the shooting areas SR2 and SR3, and this relationship is specified by the camera installation information, camera parameters, etc. of the cameras CM2 and CM3.
[0087] In the method shown in Fig. 7, the image captured by the target camera at each time (target image I TG ) is used to perform the risk assessment process. However, a modification method may be used in which the risk assessment process is performed based on an image captured by a predetermined fixed camera (hereinafter referred to as the reference camera) among the cameras CM1 to CM3 at each time. When the modification method is used, a step of executing image recognition processing on the image captured by the reference camera is added between steps S11 and S12. Then, the risk assessment process of step S12 can be performed using the result of the image recognition processing on the image captured by the reference camera (recognition result information) instead of the result of the image recognition processing on the image captured by the target camera (recognition result information). It is preferable that the reference camera be wide-angle camera CM1 in order to evaluate the risk of objects over a wide range.
[0088] <<Example EX2_A>> Example EX2_A will be described. In Example EX2_A and Examples EX2_B and EX2_C described later, it is assumed that camera CM4 is provided in camera block CB, as shown in FIG. 15 . Although cameras other than camera CM4 may be provided in camera block CB, in Example EX2_A and Examples EX2_B and EX2_C described later, it is assumed that only camera CM4 is provided in camera block CB. Camera CM4 is a variable-angle camera configured to change the angle of view during shooting (hereinafter referred to as the angle of view of camera CM4). Camera CM4 has an optical system consisting of multiple lenses and an imaging element, and the focal length during shooting of camera CM4 can be changed by adjusting the positional relationship of each lens within the optical system. Changing the focal length changes the angle of view of camera CM4. The angle of view of camera CM4 can be changed in any number of stages, but here it is assumed that the angle of view of camera CM4 can be variably set to three stages: wide angle of view, standard angle of view, and telephoto angle of view. The telephoto angle of view may also be referred to as narrow angle of view.
[0089] The wide angle of view is larger than the standard angle of view, and the standard angle of view is larger than the telephoto angle of view. The angle of view of camera CM4 includes the angle of view in the water direction (horizontal angle of view) and the angle of view in the vertical direction (vertical direction). The horizontal angle of view corresponds to the angle of view in the WX axis direction, and the horizontal angle of view corresponds to the angle of view in the WZ axis direction. When the angle of view of camera CM4 is set to the wide angle of view, both the horizontal and vertical angles of view in shooting with camera CM4 are larger than when the angle of view of camera CM4 is set to the standard angle of view. When the angle of view of camera CM4 is set to the standard angle of view, both the horizontal and vertical angles of view in shooting with camera CM4 are larger than when the angle of view of camera CM4 is set to the telephoto angle of view.
[0090] In Example EX2_A, a wide-angle camera image refers to an image captured by camera CM4 when the angle of view of camera CM4 is set to a wide angle. In Example EX2_A, a standard camera image refers to an image captured by camera CM4 when the angle of view of camera CM4 is set to a standard angle. In Example EX2_A, a telephoto camera image refers to an image captured by camera CM4 when the angle of view of camera CM4 is set to a telephoto angle. The same applies to other examples in which camera CM4 is provided in camera block CB.
[0091] The shooting area (in other words, the field of view) of camera CM4 changes in conjunction with a change in the angle of view of camera CM4. In Example EX2_A, shooting area SR1 refers to the shooting area of camera CM4 when the angle of view of camera CM4 is set to a wide angle of view (see FIGS. 3(a) and 4). In Example EX2_A, shooting area SR2 refers to the shooting area of camera CM4 when the angle of view of camera CM4 is set to a standard angle of view (see FIGS. 3(b) and 4). In Example EX2_A, shooting area SR3 refers to the shooting area of camera CM4 when the angle of view of camera CM4 is set to a telephoto angle of view (see FIGS. 3(c) and 4). The same applies to other examples in which camera CM4 is provided in camera block CB. If a landscape 600 is shot by camera CM4 when the angle of view of camera CM4 is set to a wide angle of view, a wide-angle camera image 610 is obtained by camera CM4 (see FIG. 4). If the camera CM4 captures a scene 600 with its angle of view set to the standard angle of view, a standard camera image 620 will be obtained by the camera CM4 (see FIG. 4). If the camera CM4 captures a scene 600 with its angle of view set to the telephoto angle of view, a telephoto camera image 630 will be obtained by the camera CM4 (see FIG. 4). The relationship between the shooting areas SR1 to SR3 and the relationship between the images 610, 620, and 630 has been described above with reference to FIG. 4.
[0092] In each embodiment where it is assumed that the camera CM4 is provided in the camera block CB, the target camera is fixed to the camera CM4, and therefore the target image I TG is always an image captured by the camera CM4. Since the camera CM4 captures the area surrounding the vehicle V1, the target image I TG The objects in are objects located around the vehicle V1 (objects around the vehicle).
[0093] FIG. 16 is an operation flowchart of the controller 11 according to Example EX2_A. Each process of steps S20 to S27 shown in FIG. 16 is executed by the controller 11. When the vehicle V1 starts, the controller 11 also starts, and the operation of the controller 11 starts from the process of step S20. In step S20, the controller 11 executes an initialization process. In the initialization process of step S20, the controller 11 sets the angle of view of the camera CM4 to a predetermined initial angle of view. In order to enable detection of objects over a wider range in the image recognition process, in step S20, a wide angle of view is set as the initial angle of view. However, the initial angle of view may be a standard angle of view or a telephoto angle of view. Also, in step S20, the controller 11 assigns 1 to a variable i managed by itself. After step S20, the process proceeds to step S21.
[0094] After proceeding from step S20 to step S21, the controller 11 repeatedly executes a series of loop processes consisting of steps S21 to S27. Steps S21 to S27 correspond to steps S11 to S17 in FIG. 7. Unless otherwise specified in this embodiment, the processing contents of steps S21 to S27 are the same as the processing contents of steps S11 to S17, respectively, and the description of embodiment EX1_A also applies to embodiment EX2_A. In this application, the symbols "S11 to S17" indicating the step numbers are replaced with symbols "S21 to S27" in this embodiment. The processing contents of each step shown in FIG. 16 are further explained below.
[0095] The processing content in step S21 is the same as the processing content in step S11. However, the target camera in step S21 is always the camera CM4. In step S21, the image captured by the target camera at time t[i] (target image I TG [i]) is set to the input image IN, and then image recognition processing is performed on the input image IN. The recognition result information generated thereby is stored in the memory 12. After step S21, the process proceeds to step S22.
[0096] The processing content in step S22 is the same as the processing content in step S12. That is, in step S22, the controller 11 performs a risk determination process to determine and derive the risk of each object around the vehicle based on the recognition result information generated in step S21. It is assumed that the first to m-th objects are detected as recognition target objects in the object detection process in step S21. The risk derived for the j-th object is risk DG[j], and the significance of risk DG[j] is as described in embodiment EX1_A. In step S22, the controller 11 generates risk determination information including the derived risk levels DG[1] to DG[m], and stores the generated risk determination information in the memory 12 in association with the recognition result information. After step S22, the process proceeds to step S23.
[0097] The processing content in step S23 is the same as the processing content in step S13. That is, in step S23, the controller 11 selects the maximum risk DG among the risk levels DG[1] to DG[m]. MAX Identify the maximum risk DG in the input image IN. MAX Check whether there is one object corresponding to the maximum risk DG in the input image IN. MAX If there is only one object corresponding to the maximum risk DG in the input image IN (Y in step S23), proceed to step S24. MAX If there are multiple objects that correspond to the object (N in step S23), the process proceeds to step S25.
[0098] In step S24, the controller 11 determines the maximum risk level DG MAX After step S24, the controller 11 sets the object corresponding to the maximum danger level DG as the target TG. MAX One object is selected and set as the target object TG from among a plurality of objects corresponding to the target object TG. The selection in step S25 is performed according to a predetermined selection rule. After step S25, the process proceeds to step S26.
[0099] In step S26, the controller 11 calculates the distance d between the object TG and the vehicle V1. TG In step S26, the angle of view of the camera CM4 is sometimes not changed and sometimes changed. TG The method of deriving is as shown in Example EX1_A.
[0100] Referring again to Figure 12, distance d TG The method for adjusting the angle of view of the camera CM4 according to the above will be described. TG ≦d TH1 If the condition "d" is satisfied, the angle of view of the camera CM4 is set to the wide angle of view (maintain the wide angle of view or change it to the wide angle of view). TH1 <d TG ≦d TH2 If the condition "d" is satisfied, the angle of view of the camera CM4 is set to the standard angle of view (maintain the standard angle of view or change it to the standard angle of view). TH2 <d TG When the condition "is satisfied," the angle of view of the camera CM4 is set to the telephoto angle of view (maintain the telephoto angle of view or change it to the telephoto angle of view). TH1 and d TH2 is a predetermined threshold distance, and "0 <d TH1 <d TH2 " holds true.
[0101] After step S26, the process proceeds to step S27. In step S27, the controller 11 adds 1 to the variable i. After step S27, the process returns to step S21, and the processes from step S21 onwards are repeated.
[0102] Consider case CS3, in which input image IN1 in Fig. 8 is the input image IN in step S21 and table TBL1 in Fig. 11 is used as the risk determination table. In case CS3, the risks DG[1], DG[2], and DG[3] for objects 711, 712, and 713 are determined to be 1, 3, and 5, respectively (step S22).
[0103] In case CS3, the maximum risk DG MAX is "5" in the risk level DG[3], the maximum risk level DG MAX Therefore, the process proceeds from step S23 to step S24, and the object 713 is set as the target object TG. As a result, the distance between the object 713 and the vehicle V1 is set as the distance d TG In step S26, the angle of view of the camera CM4 is adjusted in accordance with the distance between the object 713 and the vehicle V1.
[0104] In case CS3, at time t[i A ], in the case where the angle of view of the camera CM4 at the stage of step S21 is a wide angle of view (i.e., the target image I TG [i A ] is a wide-angle camera image) is called case CS3a. A represents any natural number and has the value of variable i at a certain timing. A ] distance d TG is the time t[i A ] is the distance (distance in real space) between the object 713 and the vehicle V1.
[0105] Case CS3a is the case where the time t[i A ] distance d TG In case CS3a, the process branches to the following cases CS3a_1, CS3a_2, or CS3a_3 depending on the time t[i A ] distance d TG "d TG ≦d TH1 In case CS3a_1 that satisfies ", the angle of view of the camera CM4 is maintained at the wide angle of view in step S26. In case CS3a_1, at time t[i A +1], the input image IN (i.e., the target image I TG [i A +1]) is time t[i A +1]. In case CS3a, the image is taken by the wide-angle camera CM4 at time t[i A ] distance d TG "d TH1 <d TG ≦dTH2 In case CS3a_2 that satisfies ", the angle of view of the camera CM4 is changed from the wide angle of view to the standard angle of view in step S26. In case CS3a_2, A +1], the input image IN (i.e., the target image I TG [i A +1]) is time t[i A +1] is the standard camera image taken by camera CM4. In case CS3a, at time t[i A ] distance d TG "d TH2 <d TG In case CS3a_3 that satisfies ", the angle of view of the camera CM4 is changed from the wide angle of view to the telephoto angle of view in step S26. In case CS3a_3, at time t[i A +1], the input image IN (i.e., the target image I TG [i A +1]) is time t[i A +1]. A ], the case CS3 where the angle of view of the camera CM4 at step S21 is a wide angle of view has been considered, but the same applies when the angle of view of the camera CM4 is a standard angle of view or a wide angle of view.
[0106] In case CS3a, at time t[i A ], the object 713 continues to be set as the target TG and the distance d TG "d TH1 <d TG ≦d TH2 Target image I when continuing to satisfy " TG The column of is the same as that of FIG. 13 already referred to. In the example of FIG. 13, A ] and t[i A +1], the angle of view of the camera CM4 is switched from the wide angle to the standard angle, and thereafter, at least at time t[i A +2] and t[i A +3], the angle of view of camera CM4 is maintained at the standard angle of view.
[0107] As shown in the embodiment EX1_A, it is possible to perform recognition response processing that contributes to ensuring the safety of the vehicle V1 or pedestrians, etc., based on the recognition result information obtained by the image recognition processing.
[0108] In the embodiment EX2_A, as described above, the risk level of each object around the vehicle V1 is calculated based on the image captured by the camera CM4 (S21 and S22). MAX The object corresponding to the distance d between the object TG and the vehicle V1 is set as the object TG (S23 to S25). TG Depending on the target image I TG Therefore, the same effects and advantages as those achieved by the angle of view adjustment in the embodiment EX1_A can be obtained.
[0109] Specifically, the controller 11 receives an image (I) captured by the target camera (CM4) at a first time point. TG ) and derive and identify the risk level (DG[1] to DG[m]) of each object around the vehicle based on the risk level (DG[1] to DG[m]) to set the object TG (S22 to S25). Then, the distance d between the object TG and the vehicle V1 is calculated. TG The angle of view of the target camera at a second time point, which is later than the first time point, is set in accordance with the above (S26).
[0110] This allows the angle of view of the target camera at the second time to be appropriate for recognizing the target object TG. Therefore, it is expected that the accuracy of the image recognition process for the highly dangerous target object TG will be improved (the accuracy improvement at the second time), thereby promoting the safety of the vehicle V1 or pedestrians, etc. Note that the first time and the second time are, for example, the time t[i A ] and t[i A +1] (see Figure 13).
[0111] More specifically, the controller 11 calculates the distance d TG is the threshold distance (d TH1 or d TH2), the angle of view of the target camera (CM4) at a second time point after the first time point is set to the first angle of view. TG is the threshold distance (d TH1 or d TH2 ), the angle of view of the target camera (CM4) at the second time is set to a second angle of view narrower than the first angle of view.
[0112] This allows the distance d TG If the distance d is relatively small, the object TG is photographed at a relatively large angle of view, and the size of the object TG on the target image can be optimized for image recognition. TG If is relatively large, the object TG is photographed at a relatively small angle of view, and the size of the object TG on the target image can be optimized for image recognition. In Example EX2_A, the first angle of view and the second angle of view are two of a wide angle of view, a standard angle of view, and a telephoto angle of view, for example, a wide angle of view and a standard angle of view.
[0113] <<Example EX2_B>> The example EX2_B will be described. The selection rule in step S25 of FIG. 16 is the same as the selection rule in step S15 of FIG. 7. That is, in the example EX2_A, the maximum risk DG MAX When there are two or more objects corresponding to the target TG, the method for selecting the target TG from the two or more objects is as shown in the embodiment EX1_B.
[0114] <<Example EX2_C>> Example EX2_C will be described. In Example EX2_C, several modified techniques applicable to Examples EX2_A and EX2_B will be described.
[0115] Referring again to Figures 14(a) and (b), the case where the angle of view of the camera CM4 at time t[i] is a wide angle of view is referred to as case CS4a. TG [i] is a wide-angle camera image.
[0116] In case CS4a, the controller 11TG Based on the BBOX information for [i], target image I TG It is determined whether the entire object TG (the entire image of the object TG) on [i] is located within the image area 761 (see FIG. 14(a)). TG Only when the entire object TG on [i] is located within the image area 761, the controller 11 in the case CS4a permits the camera CM4 to switch the angle of view to the telephoto angle of view. TG When all or part of the object TG on [i] is located outside the image area 761, the controller 11 according to the case CS4a determines the distance d TG When this prohibition is performed, the controller 11 in the case CS4a prohibits the camera CM4 from switching the angle of view to the telephoto angle of view regardless of the TG ≦d TH1 ", set the angle of view of camera CM4 to wide angle, and "d TH1 <d TG ", then the angle of view of camera CM4 is set to the standard angle of view.
[0117] In case CS4a, the controller 11 TG Based on the BBOX information for [i], target image I TG It is determined whether the entire object TG (the entire image of the object TG) on [i] is located within the image area 762 (see FIG. 14(b)). TG Only when the entire object TG on [i] is located within the image area 762, the controller 11 in the case CS4a permits the camera CM4 to switch the angle of view to the standard angle of view. TG [i] When all or part of the object TG is located outside the image area 762, the controller 11 according to the case CS4a determines the distance d TG When this prohibition is performed, the controller 11 in case CS4a prohibits the camera CM4 from switching the angle of view to the standard angle of view or the telephoto angle of view regardless of the distance d TG Regardless of the angle of view of the camera CM4, the angle of view of the camera CM4 is maintained as wide.
[0118] The image areas 761 and 762 in the case CS4a are the target image I captured by the camera CM4 when the angle of view of the camera CM4 is set to the wide angle. TG In case CS4a, the image area 762 is larger than the image area 761 in both the X-axis and Y-axis directions, and the entire image area 761 is included in the image area 762. TG The center position of the entire image area [i], the center position of the image area 761, and the center position of the image area 762 may coincide with one another. The image area 761 is an image area corresponding to the shooting area SR3 with a telephoto angle of view (see FIG. 4), and objects outside the image area 761 are likely to protrude from or be outside the shooting area SR3. The image area 762 is an image area corresponding to the shooting area SR2 with a standard angle of view (see FIG. 4), and objects outside the image area 762 are likely to protrude from or be outside the shooting area SR2. For this reason, the above prohibition is performed to prevent the object TG from going out of the frame. The positions, shapes, and sizes of the image areas 761 and 762 are set in advance based on the positional and dimensional relationships between the shooting areas SR1 to SR3, and these relationships are specified by the differences between the wide angle of view, the standard angle of view, and the telephoto angle of view.
[0119] Referring again to FIG. 14(c), the case where the angle of view of the camera CM4 at time t[i] is the standard angle of view is referred to as case CS4b. TG [i] is a standard camera image.
[0120] In case CS4b, the controller 11 TG Based on the BBOX information for [i], target image I TG It is determined whether the entire object TG (the entire image of the object TG) on [i] is located within the image area 763 (see FIG. 14(c)). TG Only when the entire object TG on [i] is located within the image area 763, the controller 11 in case CS4b permits the camera CM4 to switch the angle of view to the telephoto angle of view. TGWhen all or part of the object TG on [i] is located outside the image area 763, the controller 11 according to the case CS4b determines the distance d TG When this prohibition is performed, the controller 11 in case CS4b prohibits the camera CM4 from switching the angle of view to the telephoto angle of view regardless of the TG ≦d TH1 ", set the angle of view of camera CM4 to wide angle, and "d TH1 <d TG ", then the angle of view of camera CM4 is set to the standard angle of view.
[0121] The image area 763 in case CS4b is a target image I captured by the camera CM4 when the angle of view of the camera CM4 is the standard angle of view. TG In case CS4b, the target image I TG The center position of the entire image area [i] and the center position of the image area 763 may coincide with each other. The image area 763 is an image area corresponding to the shooting area SR3 of the telephoto angle of view (see FIG. 4), and an object outside the image area 763 is likely to protrude from the shooting area SR3. For this reason, the above prohibition is performed to prevent the object TG from going out of the frame. The position, shape, and size of the image area 763 are set in advance based on the positional and dimensional relationships between the shooting areas SR2 and SR3, and this relationship is specified by the difference between the standard angle of view and the telephoto angle of view.
[0122] <<Example EX3>> Example EX3 will be described. Example EX3 provides a specific method for deriving and identifying the risk level. Example EX3 is implemented in combination with each of the above-mentioned examples. Although multiple people often exist within the shooting area of the target camera, for the sake of concrete explanation, any one person existing within the shooting area of the target camera will be referred to as a person of interest, and a method for deriving and identifying the risk level for that person of interest will be described. A person of interest is one of the objects around the vehicle, and is a person detected by object detection processing.
[0123] 17 shows a sequence 800 of images captured by a target camera. The sequence 800 of images captured by a target camera is a sequence of a plurality of captured images (a plurality of target images I) generated by the target camera sequentially capturing images during a period of interest 810. TG The end time of the attention period 810 is the time t[i A Therefore, one captured image in the captured image sequence 800 is the target image I TG [i A ] (see step S11 in FIG. 7, etc.), and the target image I TG [i A ] is acquired by the latest capture of the group of captured images that make up the captured image sequence 800. The length of the period of interest 810 is an integer multiple of the frame rate of the target camera. Each captured image that makes up the captured image sequence 800 includes an image of a person of interest 820. The controller 11 can perform image recognition processing on the captured image sequence 800. The image recognition processing on the captured image sequence 800 includes image recognition processing on each captured image that makes up the captured image sequence 800.
[0124] In the embodiment EX3, the risk determination process shown below is carried out at time t[i A ]. Alternatively, the risk determination process shown below in the embodiment EX3 refers to the risk determination process performed at step S12 (see FIG. 7) at time t[i A ] refers to the risk level determination process performed in step S22 (see FIG. 16). The person of interest 820 is one of the first to m-th objects in step S12 or S22. In the embodiment EX3, the risk level of the person of interest 820 shown below is A ] refers to the risk level of the person of interest 820 identified in step S12 or S22.
[0125] The controller 11 receives at least the target image I TG [i A ], the degree to which the person of interest 820 complies with traffic rules or traffic manners (hereinafter referred to as the rule compliance degree) is derived and specified. In the embodiment EX3, the rule compliance degree of the person of interest 820 shown below is calculated at time t[iA ] refers to the degree of rule compliance of the person of interest 820 in [Image Sequence 800]. In order to ensure the accuracy or validity of the derivation of the degree of rule compliance, the controller 11 may derive and identify the degree of rule compliance of the person of interest 820 based on the captured image sequence 800.
[0126] Traffic rules are traffic laws and regulations established in the region or country in which the vehicle V1 travels, and the person of interest 820 is required to comply with the traffic rules under these laws and regulations. Traffic etiquette is a matter that is required to be complied with according to social conventions in the region or country in which the vehicle V1 travels. Traffic etiquette is not something that is required to be complied with under laws and regulations. The degree of compliance with rules may be the degree of compliance with traffic rules, the degree of compliance with traffic etiquette, or the degree of compliance with both traffic rules and traffic etiquette. Hereinafter, traffic rules and traffic etiquette will be collectively referred to as traffic rules, etc. However, traffic rules, etc. may be interpreted as a term that refers to only either traffic rules or traffic etiquette.
[0127] When the controller 11 determines that the person of interest 820 is complying with traffic rules, etc., the controller 11 sets the rule compliance level of the person of interest 820 higher than when the controller 11 determines that the person of interest 820 is not complying with traffic rules, etc. The controller 11 determines that the lower the rule compliance level of the person of interest 820, the higher the risk level of the person of interest 820. Here, for the sake of concrete explanation, the rule compliance level and risk level are each expressed as integer values between 1 and 5. When the rule compliance level of the person of interest 820 is 1, 2, 3, 4, or 5, the controller 11 sets the risk levels of the person of interest 820 to 5, 4, 3, 2, and 1, respectively.
[0128] When the controller 11 determines that the person of interest 820 is complying with traffic rules, etc., it sets the rule compliance level of the person of interest 820 to 5. When the controller 11 determines that the person of interest 820 is not complying with traffic rules, etc., it sets the rule compliance level of the person of interest 820 to 1, 2, 3, or 4. In other words, the controller 11 increases the danger level of the person of interest 820 (the danger level set by the controller 11 itself) when the person of interest 820 is not complying with traffic rules, etc., compared to when the person of interest 820 is complying with traffic rules, etc.
[0129] By setting a relatively high risk level for a person who does not follow traffic rules, etc., the person who does not follow traffic rules, etc. is more likely to be set as an object TG. When comparing a person who follows traffic rules, etc. with a person who does not follow traffic rules, etc., it is thought that it will contribute to ensuring the safety of the vehicle V1 or pedestrians, etc. to set the latter person as an object TG with priority over the former person and improve the accuracy of image recognition processing for the latter person.
[0130] When the controller 11 determines that the person of interest 820 is not complying with traffic rules, etc., it sets the rule compliance level of the person of interest 820 to 1, 2, 3 or 4 depending on the seriousness of the violation of traffic rules, etc. by the person of interest 820.
[0131] A specific example will be given with reference to Figure 18. Figure 18 shows the time t[i A ] is a diagram showing the scenery ahead of the vehicle V1 at time t[i A Consider a case CS5 in which a vehicle V1 is positioned on the above-mentioned roadway RD at time t[i] and a crosswalk PX and a traffic light TL are provided ahead of the roadway V1. The crosswalk PX is an area formed within the roadway RD and is provided for pedestrians to cross the roadway RD. The traffic light TL is a traffic light for pedestrians associated with the crosswalk PX. The state of the traffic light TL is one of several signal states including a green signal state and a red signal state, which are indicated by light or sound. The person of interest 820 is AThe person of interest 820 moves or remains stationary within the shooting area of the target camera during a period including [time period]. The movement of the person of interest 820 may be walking or running, or may be moving while riding a bicycle or a kickboard. A state in which the person of interest 820 is violating some traffic rule is referred to as a violation state.
[0132] When the traffic light TL is green, pedestrians, including the person of interest 820, are permitted to move on the crosswalk PX under traffic rules, etc. The state in which the person of interest 820 moves on the crosswalk PX when the traffic light TL is green is called a normal crossing state. When the traffic light TL is red, pedestrians, including the person of interest 820, are prohibited from moving on the crosswalk PX under traffic rules, etc. The state in which the person of interest 820 moves on the crosswalk PX when the traffic light TL is red is called a first violation state. Furthermore, pedestrians, including the person of interest 820, are prohibited from crossing the roadway RD through any part of the roadway RD other than the crosswalk PX under traffic rules, etc., regardless of the state of the traffic light TL. The state in which the person of interest 820 crosses the roadway RD through any part of the roadway RD other than the crosswalk PX is called a second violation state.
[0133] The controller 11 receives the target image I TG [i A ] or based on the result of image recognition processing on the captured image sequence 800, at time t[i A ], it is determined whether the person of interest 820 is in a normal crossing state, a first violation state, or a second violation state. A ], when it is determined that the person of interest 820 is in a normal crossing state and does not belong to any violation state, the controller 11 sets the rule compliance level of the person of interest 820 to 5. A ], when it is determined that the person of interest 820 is in the first or second violation state, the controller 11 sets 1, 2 or 3 as the rule compliance level of the person of interest 820.
[0134] Pedestrians, including the person of interest 820, are prohibited or discouraged by traffic rules and the like from moving while looking at an information terminal (smartphone, etc.). A state in which the person of interest 820 moves while looking at the information terminal he or she is holding in his or her hand is referred to as a third violation state. The controller 11 controls the target image I TG [i A ] or based on the result of image recognition processing on the captured image sequence 800, at time t[i A ], it is determined whether the person of interest 820 is in the third violation state. A ], the controller 11 sets the rule compliance level of the person of interest 820 to 1, 2, 3 or 4. At this time, the controller 11 may determine the specific value of the rule compliance level of the person of interest 820 by taking into consideration whether the person of interest 820 is in another violation state and where the person of interest 820 is moving. For example, the controller 11 may determine the specific value of the rule compliance level of the person of interest 820 by taking into consideration whether the person of interest 820 is in another violation state and where the person of interest 820 is moving. A ], if it is determined that the person of interest 820 is in a normal crossing state and in a third violation state, the controller 11 sets the rule compliance level of the person of interest 820 to 4. A ], if it is determined that the person of interest 820 is in the first or second violation state and also in the third violation state, the rule compliance level of the person of interest 820 is set to 1.
[0135] Other states that fall under the violation state include a state in which the person of interest 820 is riding a bicycle while holding an umbrella, a state in which the person of interest 820 is coasting while riding a bicycle, and a state in which the person of interest 820 is walking unsteadily. In any case, when the controller 11 determines that the person of interest 820 is in some kind of violation state, it lowers the rule compliance level of the person of interest 820 compared to when it is determined that there is no violation state at all. Although the method for setting the rule compliance level and risk level for the person of interest 820 has been explained, in reality, the rule compliance level and risk level are derived for each person who is present in the shooting area of the target camera and who has been detected by object detection processing.
[0136] The risk level may be set taking into consideration factors other than the degree of compliance with traffic rules, etc. For example, when a person of interest 820 moves from behind an object in front of the vehicle V1, the controller 11 may set the risk level of the person of interest 820 to 2 or more. This will be explained in more detail below. Consider case CS6 in which the person of interest 820 moves toward the area in front of the vehicle V1 during the period of interest 810. Case CS6 is subdivided into case CS6a and CS6b.
[0137] In case CS6a, an obstruction exists between vehicle V1 and person of interest 820 during the period of interest 810, and a portion of the person of interest 820 is obscured by the obstruction and not captured in each image in the sequence of captured images 800 (i.e., the image data of that portion is not included in each captured image). Examples of the obstruction include a parked vehicle other than vehicle V1 or a block wall. In case CS6b, no obstruction exists between vehicle V1 and person of interest 820 during the period of interest 810, and the entire body of the person of interest 820 is captured in each image in the sequence of captured images 800. In other words, in case CS6a, a blind spot occurs in capturing the person of interest 820 due to the obstruction, but in case CS6b, no blind spot occurs. Furthermore, in case CS6a, the obstruction makes it more difficult for the person of interest 820 to notice the presence of vehicle V1 than in case CS6b.
[0138] The controller 11 receives the target image I TG [i A ] or based on the result of image recognition processing on the captured image sequence 800, at time t[i A ] corresponds to either case CS6a or CS6b. Then, the controller 11 determines whether the situation at time t[i A ] corresponds to case CS6a, the controller 11 sets the risk level of the person of interest 820 higher than when it is determined that the situation corresponds to case CS6b. That is, in case CS6, when an obstruction exists between the vehicle V1 and the person of interest 820, the controller 11 sets the risk level of the person of interest 820 higher than when the obstruction does not exist.
[0139] Furthermore, in the image recognition process, the controller 11 may estimate the age group of each person detected in the object detection process. When the controller 11 estimates that the person of interest 820 is in an age group corresponding to a child, the controller 11 may increase the risk level of the person of interest 820 compared to when the controller 11 estimates that the person of interest 820 is in an age group corresponding to an adult (the former age group is lower than the latter age group). This is because, statistically, children often behave more unpredictably or erratically than adults.
[0140] <<Example EX4>> An example EX4 will be described. Fig. 19 shows a functional block diagram of the controller 11. The controller 11 is provided with functional blocks F1 to F4. The functions of the functional blocks F1 to F4 may be realized by causing the controller 11 to execute a program recorded in the memory 12, the database 20, or any other recording medium (not shown). The relationship between each functional block and the flowchart of Fig. 7 or Fig. 16 will be described.
[0141] Functional block F1 is an image recognition unit that performs image recognition processing in step S11 or S21. Functional block F2 is a risk determination unit that performs risk determination processing in step S12 or S22. Functional block F3 is an object setting unit that sets the object TG by performing processing in steps S13 to S15 or processing in steps S23 to S25. Functional block F4 is an angle of view adjustment unit that adjusts the angle of view of the target camera by performing processing in step S16 or S26.
[0142] <<Example EX5>> Example EX5 will be described. The controller 11 receives the target image I TG Image data and target image I TG BBOX information of each object in the target image I TG The controller 11 may generate a set of data including information on the danger level (DG[1] to DG[m]) of each object in the target image I. TG Each time a set of data is obtained, multiple target images I TGA training data set having a plurality of sets of data for the above can be generated. The controller 11 can store the training data set in an arbitrary database. The database is an arbitrary recording medium and may be the memory 12.
[0143] Using a training dataset, we can generate an AI that can detect objects and detect danger levels. That is, a DNN (Deep Neural Network) is formed in a machine learning device, which is an arbitrary computer device, and the DNN is trained using the training dataset. The learning here is supervised machine learning. In learning, for each set of data, the target image I is input to the DNN. TG The image data of the object is input as the problem data, and the BBOX information of each object and the danger level information of each object are input as the correct answer data. TG From the image data of the target image I TG The type and position of each object in the target image I is estimated. TG The machine learning device then estimates the danger level of each object in the image. The machine learning device then updates the DNN parameters based on the error between the DNN's estimation results and the correct data. By repeating this process, the DNN's estimation accuracy improves, and the DNN functions as a trained model. The trained model becomes an inference device that can detect objects and the danger level for any two-dimensional image.
[0144] <<Example EX6>> Example EX6 will be described.
[0145] In the above-described embodiments, people are mainly cited as objects around the vehicle and objects detected by the object detection process (recognition target objects). However, the types of objects around the vehicle and objects detected by the object detection process (recognition target objects) are arbitrary, and they may be animals other than humans, or artificial objects such as robots or vehicles.
[0146] A program that causes a computer device to execute any of the methods described in the embodiments of the present invention, and a non-volatile recording medium on which the program is recorded, are included within the scope of the embodiments of the present invention. The program that causes a computer device to execute any of the methods described in the embodiments of the present invention may be a subprogram incorporated into any main program or called by any main program. The image processing device 10 is a type of computer device. Any of the processes in the embodiments of the present invention may be realized by hardware such as a semiconductor integrated circuit, software equivalent to the program, or a combination of hardware and software.
[0147] The embodiments of the present invention can be modified in various ways as appropriate within the scope of the technical ideas set forth in the claims. The above-described embodiments are merely examples of the present invention, and the meanings of the terms of the present invention and each constituent element are not limited to those described in the above-described embodiments. The specific numerical values shown in the above description are merely examples, and as a matter of course, they can be changed to various numerical values. [Explanation of symbols]
[0148] SYS In-vehicle system V1 vehicle U1 user ST1 seat CB Camera Block CM1 wide-angle camera CM2 Standard Camera CM3 Telephoto Camera CM4 variable angle camera 10 Image recognition device 11 Controller 12 Memory 13 Communications Department 20 Vehicle control device 30 Actuator section 40 Vehicle sensor unit 50 HMI 51 Display device 52 Speaker SR1~SR3 shooting area 600 Landscape 610 wide-angle camera images 620 standard camera images 630 telephoto camera images 700 Recognition result information 702 Risk Assessment Information TBL1 table (risk assessment table) RD Roadway SW_L, SW_R Sidewalk PX crosswalk F1 Image Recognition Unit F2 Risk assessment section F3 Object setting section F4 angle of view adjustment section
Claims
1. An image recognition device that performs image recognition processing on a target image that belongs to an image captured by a camera installed in a vehicle, a controller that identifies a degree of danger of each object around the vehicle in relation to the vehicle based on the captured image, and adjusts the angle of view of the camera for obtaining the target image in accordance with the distance between the vehicle and an object that corresponds to the greatest degree of danger; ,Image recognition device.
2. A plurality of cameras having different angles of view are installed on the vehicle, the controller sets one of the plurality of cameras as a target camera and uses an image captured by the target camera as the target image, thereby achieving the adjustment; The controller sets the target object by identifying a degree of danger of each object around the vehicle based on an image captured by the target camera at a first time, and selects the target camera at a second time after the first time from the plurality of cameras according to a distance between the target object and the vehicle. The image recognition device according to claim 1 .
3. the plurality of cameras includes a first camera having a first angle of view and a second camera having a second angle of view narrower than the first angle of view, The controller sets the target camera at the second time to the first camera when the distance between the object and the vehicle at the first time is shorter than a threshold distance, and sets the target camera at the second time to the second camera when the distance between the object and the vehicle at the first time is longer than the threshold distance. The image recognition device according to claim 2 .
4. the camera is a single target camera configured to be able to change the angle of view, and an image captured by the target camera is used as the target image; The controller sets the target object by identifying a degree of danger of each object around the vehicle based on an image captured by the target camera at a first time, and sets an angle of view of the target camera at a second time after the first time according to a distance between the target object and the vehicle. The image recognition device according to claim 1 .
5. The controller sets the angle of view of the target camera at the second time to a first angle of view when the distance between the object and the vehicle at the first time is shorter than a threshold distance, and sets the angle of view of the target camera at the second time to a second angle of view narrower than the first angle of view when the distance between the object and the vehicle at the first time is longer than the threshold distance. The image recognition device according to claim 4 .
6. Each object around the vehicle is a person around the vehicle, the controller determines a degree of danger to each person around the vehicle depending on whether each person around the vehicle is complying with traffic rules or traffic manners; The controller determines the risk level of a person of interest included in the people around the vehicle to be higher when the person of interest does not comply with the traffic rules or the traffic manners than when the person of interest complies with the traffic rules or the traffic manners.
6. The image recognition device according to claim 1.
7. When there are a plurality of candidate objects corresponding to the highest risk level, the controller selects the target object from the plurality of candidate objects based on the distance between the vehicle and each candidate object.
6. The image recognition device according to claim 1.
8. The image recognition device according to claim 2 or 3; the plurality of cameras; ,In-vehicle systems.
9. The image recognition device according to claim 4 or 5; the single camera; ,In-vehicle systems.
10. A method for adjusting an angle of view executed by an image recognition device that performs image recognition processing on a target image included in an image captured by a camera installed in a vehicle, comprising: Based on the captured image, a degree of danger of each object around the vehicle is identified in relation to the vehicle, and the angle of view of the camera for obtaining the target image is adjusted according to the distance between the vehicle and the object that corresponds to the greatest degree of danger. ,How to adjust the angle of view.
11. A field angle adjusting program that causes a computer device to perform the field angle adjusting method according to claim 10.
Citation Information
Patent Citations
Image processing system for vehicles and image processor
JP2006024120A