Direct vehicle detection as 3D bounding box by using neural network image processing

CN110678872BActive Publication Date: 2026-08-28ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201880036861.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-04-04
Filing Date
2018-03-29
Publication Date
2026-08-28
Estimated Expiration
2038-03-29

Smart Images

  • Figure CN110678872B_ABST
    Figure CN110678872B_ABST
Patent Text Reader

Abstract

Systems and methods for detecting and tracking one or more vehicles in a field of view of an imaging system using neural network processing. An electronic controller receives input images from a camera mounted on a host vehicle. The electronic controller applies a neural network configured to output a definition of a three-dimensional bounding box based at least in part on the input images. The three-dimensional bounding box indicates a size and a location of a detected vehicle in a field of view of the input images. The three-dimensional bounding box includes a first quadrilateral shape and a second quadrilateral shape, the first quadrilateral shape delineating a back or front of the detected vehicle, the second quadrilateral shape delineating a side of the detected vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 481,346, filed April 4, 2017, entitled “DIRECT VEHICLE DETECTION AS 3DBOUNDING BOXES USING NEURAL NETWORK IMAGE PROCESSING”, the entire contents of which are incorporated herein by reference. Background Technology

[0003] This invention relates to the detection of the presence of other vehicles. Vehicle detection is useful for various systems, including, for example, fully or partially automated driving systems. Summary of the Invention

[0004] In one embodiment, the present invention provides a system and method for detecting and tracking a vehicle using a convolutional neural network. An image of a region adjacent to a main vehicle is captured. An electronic processor applies a convolutional neural network to process the captured image as input to the convolutional neural network and directly outputs a three-dimensional bounding box (or “bounding box”) indicating the position of the detected vehicle in the captured image. In some embodiments, the output of the convolutional neural network defines the three-dimensional bounding box as a first quadrilateral and a second quadrilateral, the first quadrilateral indicating the rear or front of the detected vehicle, and the second quadrilateral indicating the side of the detected vehicle. In some embodiments, the output of the convolutional neural network defines the first quadrilateral and the second quadrilateral as a set of six points. Furthermore, in some embodiments, the system is configured to display an image captured by a camera and a three-dimensional bounding box superimposed on the detected vehicle. However, in other embodiments, the system is configured to utilize information related to the size, location, orientation, etc., of the detected vehicle as indicated by the three-dimensional bounding box, without displaying the bounding box to the operator of the vehicle on a screen.

[0005] In another embodiment, the present invention provides a method for detecting and tracking a vehicle near a host vehicle. An electronic controller receives input images from a camera mounted on the host vehicle. The electronic controller applies a neural network configured to output a three-dimensional bounding box based at least in part on the input images. The three-dimensional bounding box indicates the size and location of the detected vehicle within the field of view of the input images. The three-dimensional bounding box includes a first quadrilateral shape and a second quadrilateral shape, the first quadrilateral shape depicting the rear or front of the detected vehicle, and the second quadrilateral shape depicting the side of the detected vehicle.

[0006] In another embodiment, the present invention provides a vehicle detection system. The system includes: a camera positioned on a main vehicle, a display screen, a vehicle system configured to control movement of the main vehicle, an electronic processor, and a memory. The memory stores instructions that, when executed by the processor, provide a certain functionality of the vehicle detection system. Specifically, the instructions cause the system to receive an input image from the camera, the input image having a field of view including a road surface on which the main vehicle operates, and a neural network is applied to the input image. The neural network is configured to provide outputs that define a plurality of three-dimensional bounding boxes, each of which corresponds to a different vehicle among a plurality of vehicles detected in the field of view of the input image. Each three-dimensional bounding box is defined by the output of the neural network as a structured set of points defining a first quadrilateral shape and a second quadrilateral shape, the first quadrilateral shape being positioned around the rear or front of the detected vehicle, and the second quadrilateral shape being positioned around the side of the detected vehicle. The first quadrilateral shape and the second quadrilateral shape are adjacent such that the first quadrilateral shape and the second quadrilateral shape share an edge. The system is further configured to display an output image on a display screen. The displayed output image includes at least a portion of the input image and each of the plurality of three-dimensional bounding boxes superimposed on the input image. The system is also configured to operate the vehicle system to automatically control the movement of the main vehicle relative to the plurality of vehicles, at least in part based on the plurality of three-dimensional bounding boxes.

[0007] Other aspects of the invention will become apparent from consideration of the detailed description and accompanying drawings. Attached Figure Description

[0008] Figure 1 It is a screenshot of an image of a road scene captured by a camera mounted on the main vehicle.

[0009] Figure 2 yes Figure 1 A screenshot of an image showing vehicles operating on a road being detected and indicated on a display screen using two-dimensional bounding boxes.

[0010] Figure 3 yes Figure 1 A screenshot of an image showing a vehicle operating on a road being detected and indicated on a display screen using a general polygon.

[0011] Figure 4 yes Figure 1 A screenshot of an image showing vehicles operating on the road being detected and indicated on a display screen using pixel-level markings.

[0012] Figure 5 yes Figure 1 A screenshot of the image, in which vehicles operating on the road are detected and indicated on the display using a combination of two quadrilaterals as a three-dimensional bounding box.

[0013] Figure 6 This is a block diagram of a system for detecting vehicles in camera image data.

[0014] Figure 7 This is a flowchart of a method for detecting and labeling vehicles using neural network processing.

[0015] Figure 8A and 8B It is processed using a convolutional neural network. Figure 7 A schematic flowchart of the method.

[0016] Figure 9 It is a block diagram of a system for detecting vehicles in camera image data by using neural network processing and for retraining the neural network.

[0017] Figure 10 It is used to detect vehicles in camera image data and for use with... Figure 9 The flowchart shows a method for retraining neural networks using a system.

[0018] Figure 11 It is a block diagram of a system for detecting vehicles in camera image data and for retraining neural networks using a remote server computer. Detailed Implementation

[0019] Before explaining any embodiments of the invention in detail, it is to be understood that the invention is not limited in its application to the details of the construction and arrangement of the components set forth in the following description or illustrated in the following drawings. The invention can have other embodiments and can be practiced or implemented in various ways.

[0020] Figure 1 until Figure 5 The illustrations show different examples of approaches for detecting a vehicle in images captured by cameras mounted to the main vehicle. Figure 1 The illustration shows images of a road scene captured by a camera mounted on a main vehicle. The images in this example include straight-line views from the main vehicle's perspective, and also include several other vehicles operating on the same road as the main vehicle within the camera's field of view. Although Figure 1The example illustrates a straight-line image captured by a single camera, but in other implementations, the system may include a camera system configured to capture “fisheye” images of the road, and / or may include multiple cameras to capture images of the road surface using different viewpoints and / or fields of view. For example, in some implementations, the camera system may include multiple cameras configured and positioned to have a field of view that at least partially overlaps with other cameras in the camera system in order to capture / compute three-dimensional image data of other vehicles operating on the road.

[0021] As described further below, the system is configured to analyze images (or images) captured by a camera (or multiple cameras) to detect the positions of other vehicles operating on the same road as the main vehicle. In some implementations, the system is configured to detect vehicles and define the shape and position of the detected vehicles by defining the location of the shape corresponding to the detected vehicle in three-dimensional space. In some implementations, the images (or images) captured by the camera (or multiple cameras) along with the defined "shape" overlaid on the images are output to a display screen to indicate the position of the detected vehicle(s) to the user.

[0022] Figure 2 The diagram illustrates the "bounded box" approach, where the image (e.g., as in...) Figure 1 In the example, the image captured by the camera is processed, and two-dimensional rectangles are placed around the vehicle detected in the image frame. As discussed in further detail below, the image detection algorithm applied by the vehicle detection system can be tuned, adjusted, and / or trained to place the rectangular boxes such that they completely surround the detected vehicle. In some implementations, the system is configured to use the location and size of the two-dimensional bounding boxes (which indicate the size and position of a particular vehicle operating on the road) as input data for operations such as, for example, distance estimation (e.g., between a primary vehicle and another detected vehicle), collision detection / warning, dynamic cruise control, and lane change assistance. For operations such as distance estimation and collision detection, Figure 2 The two-dimensional bounding box approach has relatively low computational cost. However, using two-dimensional bounding boxes results in a relatively large amount of "non-vehicle" space within the bounded area of ​​the image. This is true for both axis-aligned and non-axis-aligned rectangles. Furthermore, the rectangles do not provide any information about the orientation of the detected vehicle relative to the road or the main vehicle.

[0023] exist Figure 3In the example, general polygons are used to more closely delimit the region of a vehicle detected in the image. In this approach, the system is configured to determine the location and size of a vehicle detected in a captured image, but is also configured to identify general vehicle shapes that best correspond to the shape of the detected vehicle. For example, the system can be configured to store multiple general shapes, each indicating a different general vehicle type—e.g., light truck, SUV, minivan, etc. When a vehicle is detected in one or more captured images, the system is configured to identify the shape that best corresponds to the multiple general shapes of the detected vehicle. While this approach does model the detected vehicle more accurately, it also increases computational cost. (The last sentence appears to be incomplete and possibly refers to a different approach.) Figure 2 Compared to using 2D bounding boxes in the examples, each operation performed using a polygon-based model, including the detection itself, is computationally more expensive.

[0024] Figure 4 An example is illustrated where a vehicle is detected in a camera image using pixel-level annotation and detection. Figure 4 The example shows thick solid lines that delineate each group of pixels that have been identified as associated with the body of the vehicle. However, in some implementations, each individual pixel identified as associated with the body of the vehicle is highlighted with a different color (e.g., light gray), making each detected body of the vehicle appear to be shaded or highlighted with a different color. By detecting the boundaries and dimensions of the vehicle at the individual pixel level, this method for vehicle detection is superior to the corresponding approach in... Figure 2 and 3 The 2D bounding box or polygon-based approach illustrated in the diagram is more accurate. However, this approach again increases computational complexity. Instead of simply sizing and positioning a general shape to best fit the detected vehicle, the system is configured to analyze the image to detect and define the vehicle's actual, precise board at a pixel-by-pixel level. Separating individual objects at the pixel level can take up to several seconds per image, and after separation, the handling of individual vehicles, collision checks, heading calculations, etc., can also be computationally more expensive than the polygon-based approach.

[0025] Figure 5The diagram illustrates yet another vehicle detection mechanism that utilizes a combination of two quadrilaterals: a bounding box identifying the rear (or front) of the vehicle and a corresponding bounding box identifying a single side of the same vehicle. In some implementations, the quadrilaterals in this approach can be simplified to parallelograms, and in highway driving scenarios, the two quadrilaterals can be further simplified to a combination of an axis-aligned rectangle and a parallelogram. In some implementations, the system is also configured to determine whether the sides of the vehicle are visible in the captured image, and if not, to label the image captured by the camera using only the visible sides of the vehicle. With a fixed-frame camera, only any two sides of the vehicle are visible at any given time step.

[0026] Figure 5 The fixed model offers several advantages. It comprises only a few straight planes that can be interpreted as 3D bounding boxes. A range of computer graphics and computer vision algorithms can then be deployed in a computationally highly efficient manner. As a result, compared to using… Figure 2 Compared to the example 2D bounding boxes, processing images using these “3D bounding boxes” requires only slightly higher computational costs. The resulting 3D bounding boxes also provide information related to the 3D shape, size, and orientation of the detected vehicle. Furthermore, labeling images for training a system configured to use artificial intelligence (e.g., neural network processing) to detect and define the location of vehicles based on one or more captured images is only slightly more complex than labeling images using 2D bounding boxes.

[0027] Figure 6 It is used for using Figure 5 A block diagram illustrating an example image processing system that uses 3D bounding box technology to detect the presence of vehicles. While this example, and others, focus on 3D bounding box technology, these examples can be further adapted in some implementations to other vehicle identification and labeling techniques, including, for example, in... Figure 2 until Figure 4 Those shown in the diagram.

[0028] exist Figure 6 In the example, camera 501 is positioned on the main vehicle. Images captured by camera 501 are transmitted to electronic controller 503. In some implementations, electronic controller 503 is configured to include an electronic processor and a non-transitory, computer-readable storage device storing instructions that are executed by the electronic processor of electronic controller 503 to provide image processing and vehicle detection functionality as described herein. Images captured by camera 501 are processed (via electronic controller 503 or another computer system, as discussed further below) and displayed on the screen of display 505 along with labels identifying any detected vehicles overlaid on the images. Figure 7The diagram shows... Figure 6 An example of system operation. Image 601 is captured using a field of view that includes the road in front of the main vehicle. Neural network processing is applied to the captured image (at box 603) to determine the appropriate placement, size, and shape of a three-dimensional bounding box for any vehicle detected in the captured image. The output shown on display 505 (output image 605) includes at least a portion of the original image 601 captured by the camera and annotations superimposed on image 601 identifying any detected vehicles. Figure 7 In the example, output image 605 shows 3D bounding box annotations that identify the vehicle that has already been detected in the original image 601 captured by camera 501.

[0029] Figure 7 Examples and other examples presented herein discuss displaying labeled camera images on a screen that can be viewed by the vehicle's operator. However, in some implementations, the camera images and / or three-dimensional bounding boxes are not displayed on any screen within the vehicle, and instead, the system can be configured to use defined three-dimensional bounding boxes indicating the location, size, etc., of the detected vehicle solely as input data to other automated systems of the vehicle. These other automated systems can be configured, for example, to use the location and size of the three-dimensional bounding box (which indicates the size and position of a particular vehicle operating on the road) as input data for operations such as, for example, distance estimation (e.g., between the primary vehicle and another detected vehicle), collision detection / warning, dynamic cruise control, and lane change assistance. Furthermore, in some implementations, the detection of a vehicle in the field of view can be based on analysis of captured images(s) and additional information captured by one or more additional vehicle sensors (e.g., radar, sonar, etc.). Similarly, in some implementations, information about the detected vehicle, indicated by the placement of a three-dimensional bounding box, is used as input to one or more additional processing steps / operations, where detection and timing information from several sensors is combined (i.e., sensor fusion). In some implementations, the output of the sensor fusion process can then be displayed, or in others, reused by one or more vehicle systems without showing any information to the vehicle operator. Information derived from image analysis and / or a combination of image analysis and information from one or more additional vehicle sensors can, for example, be used by the main vehicle's trajectory planning system.

[0030] A 3D bounding box can be defined by a fixed number of structured points. For example, in some implementations, the 3D bounding box is defined by six points—four points define the corners of a 2D rectangle indicating the rear of the vehicle, and four points define the corners of a 2D quadrilateral indicating the sides of the vehicle (resulting in only six points, since the two quadrilaterals defining the detected vehicle share a side and therefore share two points). In other implementations, the two quadrilaterals defining the 3D bounding box are calculated / determined as eight structured points, which define four corners of each of the two quadrilaterals.

[0031] In some implementations, a fixed number of structured points defining the 3D bounding box are defined in two-dimensional space of the image, while in others, the structured points are defined in three-dimensional space. In some implementations, the structured points defining the 3D bounding box are defined both in three-dimensional space (for use as input data for an automated vehicle control system) and in two-dimensional space (for display to the user in the output image). In some implementations, the system is configured to provide a fixed number of structured points defining the 3D bounding box (e.g., 6 points defining the 3D bounding box in 2D space, 6 points defining the 3D bounding box in 3D space, or 12 points defining the 3D bounding box in both 2D and 3D spaces) as the direct output of a machine learning image processing routine. In other implementations, the system can be configured to provide a limited number of structured points as the output of a neural network processing only in 2D or 3D space, and then apply a transformation to determine the structured points for other coordinate systems (e.g., determining the 3D coordinates of the structured points based on the 2D coordinates output by the neural network). In other implementations, the system can be configured to apply two separate neural network processing routines to separately determine the structured points in 2D and 3D space.

[0032] In some implementations, the system may further be configured to determine a set of eight structured points defining a 3D bounding box for a vehicle. These eight structured points collectively define the four corners of a quadrilateral on each of the four distinct side surfaces of the 3D bounding box (e.g., the two sides of the vehicle, the front of the vehicle, and the rear of the vehicle). In some implementations, the neural network is configured to output the entire set of eight structured points; in other configurations, the neural network outputs structured points defining either the rear or front of the vehicle, plus an additional side surface, and the controller is configured to compute these two additional structured points to define all eight corners of the 3D bounding box based on the set of six structured points output by the neural network. In some implementations, the system may further be configured to compute and output four additional visibility probabilities indicating which of the output points (and consequently which of the sides of the 3D bounding box) are visible and should be used for display or further processing.

[0033] The system can also be configured to apply other simplifications. For example, the system can be configured to assume that all detected vehicles are moving in the same direction during highway driving scenarios, and thus the rear and / or front of the vehicle can be estimated as a rectangle. In other cases (e.g., on fairly flat roads), the system can be configured to represent all sides and the rear / front of the detected vehicle as trapezoids (with only a small reduction in accuracy). By using a fixed structure, such as a trapezoid, fewer values ​​need to be calculated per point because some points will share values ​​(e.g., two corners of adjacent shapes sharing an edge).

[0034] As discussed above, in Figure 7 In the example, a neural network process is applied to capture images to directly generate 3D bounding boxes (e.g., points that define 3D bounding boxes). Figure 8A and 8B The illustration also shows an example using a convolutional neural network. A convolutional neural network is a machine learning image processing technique that analyzes images to detect patterns and features, and based on the detected patterns / features (and in some cases, contextual information), outputs information such as the identifiers of objects detected in the image. Figure 8A As shown, a convolutional neural network is trained to receive input images from a camera and output an output image comprising 3D bounding boxes for any vehicles detected in the original image. In some implementations, the neural network is configured to provide a dynamic number of outputs, each defining structured points of a separate bounding box corresponding to a different vehicle detected in the input image. Figure 8BA specific example is illustrated, in which a convolutional neural network 701 is configured to receive as input a raw input image 703 captured by a camera and additional input data 705, including, for example, sensor data from one or more other vehicle sensors (e.g., sonar or radar), any bounding boxes defining the vehicle detected in the previous image, vehicle speed (for the main vehicle and / or for one or more other detected vehicles), vehicle steering changes, and / or acceleration. Based on these inputs, the convolutional neural network outputs different sets 707, 709, 711, 713 of dynamically generated structured data points, each defining a different 3D bounding box corresponding to the vehicle detected in the input image 703. However, in other implementations, the convolutional neural network processing 701 is designed and trained to directly generate the size and placement of the 3D bounding boxes based solely on the input image 703 from the camera without any additional input data 705.

[0035] Neural networks—especially those such as those in… Figure 8A and 8B The convolutional neural networks illustrated in the examples are "supervised" machine learning techniques (i.e., they can be retrained and improved by user feedback on incorrect results). In some implementations, the neural network-based image processing system is developed and trained before being deployed in the vehicle. Therefore, in some implementations, Figure 6 The system configuration illustrated can be used to simply capture images, process the images to detect the presence of other vehicles, and output the results to a display without any user input device for supervised retraining. In fact, in some implementations, the output display (e.g., Figure 6 The display 505 may not be used or even included in the system; instead, images from the camera are processed to identify the vehicle, and vehicle detection data (e.g., one or more combinations of points defining each detected vehicle in three-dimensional space) are used by other automated vehicle systems (e.g., fully or partially automated driving systems) without any indication of the identified vehicle being graphically displayed to the user in the image.

[0036] However, in other implementations, the system is further configured to receive user input to continue retraining and improve the operation of the convolutional neural network. Figure 9An example of a system is illustrated, configured to apply a convolutional neural network to detect a vehicle as a combination of six points defining a 3D bounding box, and further configured to retrain the convolutional neural network based on input from a user. The system includes an electronic processor 801 and a non-transitory computer-readable storage 803 storing training data and instructions for neural network processing. A camera 805 is configured to periodically capture images and transmit the captured images to the electronic processor 801. The electronic processor 801 processes the captured images(s), applies neural network processing to detect vehicles in the captured images, and defines a 3D bounding box for any detected vehicle.

[0037] As discussed above, the defined three-dimensional bounding box and / or information determined at least in part based on the three-dimensional bounding box can be provided by the electronic processor 801 to one or more auxiliary vehicle systems 811, including, for example, vehicle systems configured to control the movement of a vehicle. For example, vehicle system 811 may include one or more of the following: an adaptive cruise control system, a lane change assist system, or other vehicle systems configured to automatically control or adjust the vehicle's steering, speed, acceleration, braking, etc. Vehicle system 811 may also include other systems, for example, that can be configured to calculate / monitor the distance between the main vehicle and other detected vehicles, including, for example, collision detection / warning systems.

[0038] The electronic processor 801 is also configured to generate an output image that includes at least a portion of the image captured by the camera 805 and any 3D bounding boxes indicating the vehicle detected by neural network image processing. The output image is then transmitted by the electronic processor 801 to a display 807, where it is displayed on the screen of the display 807. Figure 9 The system also includes an input device 809 configured to receive input from a user, which is then used to retrain the neural network. In some implementations, the display 807 and the input device 809 may be provided together as a touch-sensitive display.

[0039] In some implementations—including, for example, those utilizing a touch-sensitive display—the system can be configured to enable a user to retrain the neural network by identifying (e.g., by touching on a touch-sensitive display) any vehicles in the displayed image that are not automatically detected by the system and any displayed 3D bounding boxes that do not properly correspond to any vehicle. Figure 10An example of a method for providing this type of retraining, implemented by an electronic processor 801, is illustrated. An image is received from a camera (at box 901, "Receive image from camera"), and neural network processing is applied to determine the location and size of any 3D bounding boxes, each defined by a set of structured points in 2D and / or 3D space (at box 903, "Apply neural network processing to determine (multiple) 3D bounding boxes"). The image is then displayed on a monitor along with the 3D bounding boxes (if any) overlaid on the image (at box 905, "Display image with overlaid (multiple) 3D bounding boxes"). The system then monitors for any user input devices (e.g., touch-sensitive displays) (at box 907, "Has user input been received?"). If no user input is received, the system proceeds to process the next image received from the camera (repeated boxes 901, 903, and 905).

[0040] In this particular example, user input is received via a “touch” on a touch-sensitive display. Therefore, when a “touch” input is detected, the system determines whether the touch input was received within the 3D bounding box shown on the display (at box 909, “Is the user input within the bounding box?”). If so, the system determines that the user input indicates the displayed 3D bounding box has incorrectly or inaccurately indicated the detected vehicle (e.g., there is no vehicle in the image corresponding to the bounding box or the bounding box is not properly aligned with the vehicle in the image). The system continues to retrain the neural network based on this input (at box 911, “Update neural network: Incorrect vehicle detection”). Conversely, if touch input is received at any location outside the displayed 3D bounding box, the system determines that the user input identifies a vehicle shown in the image that was not detected by the neural network. The neural network is then retrained accordingly based on this user input (at box 913, “Update neural network: Vehicle not detected in the image”). The updated / retrained neural network is then used to process the next image received from the camera (repeated boxes 901, 903, and 905).

[0041] In some implementations, the system is further configured to apply additional processing to the user-selected location to automatically detect vehicles at the selected location corresponding to undetected vehicles. In other implementations, the system is configured to prompt the user to manually place a new 3D bounding box at the selected location corresponding to the undetected vehicle. In some implementations, the system is configured to display this prompt for manually placing a new 3D bounding box in real time. In other implementations, an image of the selected undetected vehicle is received and stored in memory, and the system outputs an image with a prompt for manually placing a new 3D bounding box at a later time (e.g., when the vehicle stops).

[0042] Furthermore, in some implementations, the system is additionally configured to provide retraining data as a refinement of the displayed / output 3D bounding box. For example, the system may be configured to allow a user to selectively and manually adjust the size of the 3D bounding box after it has been displayed on the screen for a detected vehicle. After the user adjusts the shape, positioning, and / or size of the bounding box to more accurately indicate the rear / front and sides of the vehicle, the system can use this refinement as additional retraining data to retrain the neural network.

[0043] In the examples discussed above, images are captured and processed, and the neural network is retrained locally by the electronic processor 801 and the user in the vehicle. However, in other implementations, the system can be configured to interact with a remote server. Figure 11 An example of such a system is illustrated, comprising an electronic processor 1001, a camera 1003, and a wireless transceiver 1005 configured to communicate wirelessly with a remote server computer 1007. In various implementations, the wireless transceiver 1005 may be configured to communicate with the remote server computer 1007 using one or more wireless modes, including, for example, a cellular communication network.

[0044] In various implementations, attached to or replacing the electronic processor 1001, the remote server computer 1007 can be configured to perform some or all of image processing and / or neural network retraining. For example... Figure 10The system can be configured to transmit image data captured by camera 1003 to remote server computer 1007, which is configured to process the image data using a neural network and transmit a six-point combination back to wireless transceiver 1005, the six-point combination identifying a 3D bounding box for any vehicle detected in the image. By performing image processing at remote server computer 1007, the computational load is transferred from the local electronic processor 1001 in the vehicle to remote server computer 1007.

[0045] Retraining of the neural network can also be transferred to a remote server computer 1007. For example, an employee can review images received and processed by the remote server computer 1007 in real time or at a later time to identify any false positives or missed vehicle detections in the captured images. This information is then used to retrain the neural network. In addition to reducing computational complexity and reducing (or completely removing) the retraining burden from vehicle operators, in implementations where multiple vehicles are configured to interact with the remote server computer, a larger amount of retraining data can be captured and processed, resulting in a more robustly trained neural network.

[0046] In the examples discussed above, the remote server computer is configured to perform image processing and retraining of the neural network. Therefore, in some implementations, the display 1009 and / or input device 1011 may not be included in the main vehicle. However, in other implementations, some or all of the image processing and / or retraining functionality, instead of being performed by the remote server computer 1007, is implemented by the local electronic processor 1001. For example, Figure 10 The system can be configured to apply neural network image processing at a local electronic processor 1001 and transmit image data with overlaid 3D bounding boxes to a remote server computer 1007. Instead of retraining the neural network locally by the user / driver of the vehicle, the retraining of the neural network is then performed remotely using the collected image / output data and image / output data received from any other host vehicle connected to the remote server computer 1007. The remote server computer 1007 then periodically or upon request updates the neural network stored and implemented by the electronic processor 1001 based on remote / collective retraining.

[0047] Similarly, in some implementations, the system may be configured to apply neural network image processing at a local electronic processor 1001 and receive retraining input from a user via a local input device 1011, the identifier of which is referenced above. Figure 10The discussion concerns false positives and undetected vehicles. However, instead of retraining the neural network locally, the retraining input and corresponding images received by input device 1011 are saved to memory. The system is further configured to periodically or upon request upload the retraining input and corresponding images to a remote server computer 1007, which in turn develops a retrained neural network based on the retraining input / images received from multiple different master vehicles connected to the remote server computer 1007, and transmits the updated / retrained neural network to electronic processor 1001 for use.

[0048] Finally, while some of the discussions in the above examples are based on the training and retraining of neural networks using manual user input (from the user operating the vehicle or through another person on a remote server computer), in other implementations, the training and retraining of the neural network can be achieved by using another vehicle detection algorithm to verify the correct presence / location of the vehicle. Furthermore, in some implementations, the system can be configured to automatically determine the confidence level of vehicle detection in an image and automatically forward images marked as "low confidence" to a remote server computer for further manual or automated processing and for retraining the neural network. Additionally, in some implementations, it is possible to apply some of the techniques and systems described herein to systems that use sensor input instead of a camera as input to the detection system.

[0049] Therefore, among other things, the present invention provides systems and methods for directly detecting and labeling other vehicles in images using neural network processing. Various features and advantages are set forth in the appended claims.

Claims

1. A method for detecting and tracking a vehicle near a main vehicle, the method comprising: The electronic controller receives input images from cameras mounted on the main vehicle; An electronic controller applies a neural network configured to provide outputs defining a plurality of three-dimensional bounding boxes, each indicating the size and location of a different vehicle among a plurality of vehicles detected in the field of view of an input image. Each three-dimensional bounding box is defined by the output of the neural network as a structured set of points defining a first quadrilateral shape and a second quadrilateral shape. The first quadrilateral shape depicts the rear or front of the detected vehicle, and the second quadrilateral shape depicts the side of the detected vehicle. The first quadrilateral shape is adjacent to the second quadrilateral shape. as well as Display an output image on a screen, the output image including at least a portion of the input image and an indication of a three-dimensional bounding box overlaid on the input image; The method further includes: Receive first user input, which indicates the selection of a position on the output image outside the three-dimensional bounding box; Based on the first user input, it is determined that there is an undetected vehicle in the field of the input image at the location corresponding to the user input; The user is prompted to manually locate another 3D bounding box relative to a vehicle that is not detected in the input image; Define another three-dimensional bounding box based on the second user input received in response to the prompt; and The neural network is retrained based on user input indicating an undetected vehicle, wherein retraining the neural network includes at least partially retraining the neural network based on the input image to output the other three-dimensional bounding box.

2. The method of claim 1, further comprising: Receive user input indicating the selection of a position on the output image within one of the three-dimensional bounding boxes; The user input is used to determine the incorrect detection of the vehicle by the three-dimensional bounding box indication via a neural network; and The neural network is retrained based on the incorrect detections.

3. The method of claim 1, further comprising automatically operating the vehicle system to control the movement of the master vehicle at least in part based on the three-dimensional bounding box.

4. The method of claim 3, further comprising determining, at least in part, the distance between the master vehicle and one of the detected vehicles based on a corresponding three-dimensional bounding box, and wherein automatically operating the vehicle system comprises automatically adjusting the speed of the master vehicle based at least in part on the determined distance between the master vehicle and the detected vehicles.

5. The method of claim 4, wherein automatically operating the vehicle system includes automatically adjusting at least one selected from the group consisting of: steering of the main vehicle, acceleration of the main vehicle, braking of the main vehicle, defined trajectory of the main vehicle, and signaling of the main vehicle, at least in part based on the positioning of the three-dimensional bounding box.

6. The method of claim 1, wherein the definition of each of the three-dimensional bounding boxes comprises a set of six structured points defining four corners of the first quadrilateral shape and four corners of the second quadrilateral shape, wherein two corners of the first quadrilateral shape are the same as two corners of the second quadrilateral shape.

7. The method of claim 6, wherein the set of six structured points defines the corners of the first quadrilateral shape and the corners of the second quadrilateral shape in a three-dimensional coordinate system.

8. The method of claim 6, wherein the set of six structured points defines the corners of a first quadrilateral shape and the corners of a second quadrilateral shape in a two-dimensional coordinate system of the input image.

9. The method of claim 8, wherein the definition of the three-dimensional bounding box further includes a second set of structured points that define the corners of the first quadrilateral shape and the corners of the second quadrilateral shape in a three-dimensional coordinate system.

10. The method of claim 1, wherein the neural network is further configured to output a plurality of three-dimensional bounding boxes based at least in part on the input image, wherein each of the plurality of three-dimensional bounding boxes indicates the size and location of a different one of a plurality of vehicles detected in the field of view of the input image.

11. A vehicle detection system, the system comprising: Cameras positioned on the main vehicle; Display screen; The vehicle system is configured to control the movement of the main vehicle; Electronic processor; as well as A memory storing instructions that, when executed by an electronic processor, cause the vehicle detection system to: The system receives an input image from a camera, the input image having a field of view that includes the road surface on which the main vehicle operates. A neural network is applied, configured to provide outputs that define a plurality of three-dimensional bounding boxes, each indicating the size and location of a different vehicle among a plurality of vehicles detected in the field of view of an input image. Each three-dimensional bounding box is defined by the output of the neural network as a structured set of points defining a first quadrilateral shape and a second quadrilateral shape. The first quadrilateral shape depicts the rear or front of the detected vehicle, and the second quadrilateral shape depicts the side of the detected vehicle. The first quadrilateral shape is adjacent to the second quadrilateral shape. The display screen shows an output image, which includes at least a portion of the input image and indications of the plurality of three-dimensional bounding boxes overlaid on the input image. When the aforementioned instructions are executed by the electronic processor, they cause the vehicle detection system to: Receive first user input, which indicates the selection of a position on the output image outside the three-dimensional bounding box; Based on the first user input, it is determined that there is an undetected vehicle in the field of the input image at the location corresponding to the user input; The user is prompted to manually locate another 3D bounding box relative to a vehicle that is not detected in the input image; Another three-dimensional bounding box is defined based on the second user input received in response to the prompt; as well as The neural network is retrained based on user input indicating an undetected vehicle, wherein retraining the neural network includes at least partially retraining the neural network based on the input image to output the other three-dimensional bounding box.

Citation Information

Patent Citations

  • Target recognition system and target recognition method executed by the target recognition system

    US20130322691A1

  • Retraining a machine classifier based on audited issue data

    WO2017027030A1