Image data alignment for mobile system sensors

By identifying and correcting color inconsistency, determining and applying color consistency parameters in the image signal processing system, the problem of degradation in prediction accuracy caused by image appearance mismatch in the prior art is solved, and the accuracy of object recognition and positioning is improved.

CN119946446APending Publication Date: 2025-05-06FORD GLOBAL TECH LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411509511.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-10-28
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and locate objects in the environment around the vehicle when acquiring images that do not match the appearance of the image in the training dataset, resulting in a decrease in prediction accuracy.

Method used

Color inconsistency is identified by a machine learning system based on predicted differences determined by multiple images, and statistical analysis is performed to determine color consistency parameters in the image signal processing system, enhancing color consistency to mitigate predicted differences.

Benefits of technology

Improve the consistency of image colors in image signal processing systems and enhance the prediction accuracy of machine learning systems when identifying and positioning objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946446A_ABST
    Figure CN119946446A_ABST
Patent Text Reader

Abstract

The present disclosure provides image data alignment for mobile system sensors. A computer includes a processor and a memory including instructions executable by the processor to determine, with a machine learning system, a first prediction based on receiving a first image from a first camera, and determine, with the machine learning system, a second prediction based on receiving a second image from a second camera. When the first prediction is not equal to the second prediction within a user-determined tolerance: determining a color consistency based on comparing pixel values from the first image to a threshold determined based on previously determined pixel values, determining a color correction parameter for inclusion in an image signal processing system by determining pixel statistics based on pixel values from the first image; and applying the color correction parameter to a third image by receiving the third image from the first camera at the image signal processing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to image data alignment for mobile system sensors. Background Art

[0002] Computers can operate systems and devices including vehicle, robot, drone, and / or object tracking systems. Data including images can be acquired by sensors and processed by a computer to determine the position of the system relative to the environment and relative to objects in the environment. The computer can use the position data to determine one or more trajectories and / or actions for operating the system or its components in the environment. Summary of the invention

[0003] Systems that are mobile and / or have moving parts (including vehicles, robots, drones, mobile phones, etc.) can be operated by acquiring sensor data (including data about the environment surrounding the system) and processing the sensor data to determine the location of objects in the environment surrounding the system. The determined location data can be processed to determine the operation of the system or part of the system. For example, a robot can determine the location of the arm of another nearby robot. The robot can use the determined robot arm position to determine a path on which to move a gripper to grab a workpiece without encountering the arm of another robot. In another example, a vehicle can determine the location of another vehicle traveling on a road. The vehicle can use the determined location of another vehicle to determine a path on which to operate while maintaining a predetermined distance from the other vehicle. Vehicle operation will be used as a non-limiting example of system position determination in the following description herein.

[0004] A machine learning system (e.g., a convolutional neural network) may be trained to determine the identity and location of one or more objects included in an environment (e.g., roads and vehicles). A convolutional neural network includes convolutional layers and fully connected layers, and may be trained to recognize and locate objects. Training a convolutional neural network may require a training data set, which may include thousands of video sequences, which may include millions of images. Additionally, training a machine learning system such as a convolutional neural network may require ground truth data for the images in the training data set. Ground truth includes annotation data about the identity and location of objects included in the training data set obtained from a source other than the machine learning system, such as user annotations of images in the training data set.

[0005] A trained machine learning system, such as a convolutional neural network, may be installed in a computing device in a vehicle to receive sensor data from sensors included in the vehicle. The machine learning system may determine predictions about the received sensor data to assist in operating the vehicle. For example, a trained convolutional neural network may be trained to receive images from a camera and determine predictions about the environment surrounding the vehicle. The predictions may include determining the position and motion of the vehicle relative to the environment and the position and motion of objects in the environment. These predictions may include determining a three-dimensional (3D) space-color coordinate system based on the acquired two-dimensional (2D) images.

[0006] Obtaining predictions that accurately identify and locate objects in the vehicle's surroundings may depend on acquiring images that match the appearance of images included in a training data set used to train a machine learning system that determines predictions. Matching appearance, as used herein, means having color consistency with images included in the training data set. Color consistency, as used herein, means that different images include the same color values ​​(e.g., the same RGB values) within a tolerance or margin (e.g., 10% or less) from the same physical location or object within a scene, regardless of which of one or more cameras acquires the image. In examples where color consistency is expected, color inconsistency means a lack of color consistency, for example, when different images include different color values ​​greater than a tolerance value. Color consistency is generally expected when a camera observes the same scene twice (camera consistency), when two cameras observe the same object in a portion of the same scene (spatial consistency), and / or when two cameras observe the same object at different times (temporal consistency). Color inconsistency can be caused by changes in camera electronics over time (such as automatic gain control circuits, image signal processing (ISP) calculations, etc.) and changes in camera optics (e.g., dirt on the lens, age-related changes in the protective cover, etc.). Color inconsistency can also be caused by differences in natural lighting (such as the angle of the sun) or differences in artificial lighting (such as headlights or streetlights). Differences in natural and artificial lighting between cameras can be caused by differences in the field of view of different cameras, which results in different viewing angles for parts of the scene. The techniques described herein for correcting color inconsistency can also be applied to images acquired by grayscale cameras, multispectral cameras, cameras with sensitivity in the near infrared (NIR), infrared (IR), or ultraviolet (UV) portions of the electromagnetic spectrum, and lidar. These techniques can also be used to correct color problems, such as incorrect assumptions about illumination-invariant imaging or ISP color correction errors.

[0007] A system such as a vehicle may have multiple cameras that acquire data about the environment around the system (e.g., around the vehicle). The multiple cameras may acquire images that are input into a machine learning system to determine the position of the vehicle relative to the environment, determine the motion of the vehicle, and determine the position of objects (such as other vehicles) in the environment around the vehicle. Techniques for color alignment described herein may determine color inconsistency using predicted differences determined by the machine learning system based on multiple images. When color inconsistency is determined, a statistical analysis may be performed on pixel data included in the color image to determine color consistency parameters included in the image signal processing system to enhance color consistency and mitigate the predicted differences determined by the machine learning system.

[0008] Color consistency may include component values ​​of camera color consistency, spatial color consistency, and temporal color consistency. An image with camera color consistency has RGB pixels included in objects such as vehicles and roads that have the same value within a tolerance as described above when output by a camera when observing the same type of scene illuminated by the same type of lighting. The same type of scene means that objects of the same or the same type or class (such as roads, street signs, and vehicles, etc.) are included within a specified margin (e.g., 10%) of the same position, size, and orientation in the scene. The same type of lighting means that the scene includes the same illumination within a specified margin (e.g., 10%) of the same sun angle, lighting intensity, lighting pattern, etc., such as direct sun, cloudy sky, street lights, vehicle headlights, etc. Camera color consistency can be determined by acquiring images of the same type of scene illuminated by the same type of lighting at different times and comparing the predictions output by the machine learning system to determine whether the same prediction is output for images acquired by the same camera at different times.

[0009] The techniques for color consistency described herein can be used to filter out or calibrate objects with strong color or reflectivity anisotropy. For example, a road surface (dry) will have a larger diffuse reflectivity component than a vehicle. The light source in the scene (e.g., sun position) can be calculated, and the object surface orientation can be compared to the camera position angle to perform a bidirectional reflectance distribution function (BRDF) calculation of the diffuse reflectivity and specular reflectivity components. Some objects (such as some painted cars) can include paint with high color anisotropy. Determining outliers that exhibit strong color or reflectivity anisotropy can enhance detection and reduce false positives. As discussed above, camera color consistency can degrade over time due to drift of electronics or degradation of optics. Camera color inconsistency is caused by variations within a single camera.

[0010] Spatial color consistency means that images including matching digital color values ​​as described above are output by two different cameras when viewing portions of the same scene from different directions. Spatial color consistency can be determined by comparing predictions output by a machine learning system for overlapping portions of images acquired by more than one camera at substantially similar times. The machine learning system should determine the same predictions for objects included in overlapping portions of images acquired at substantially the same time (e.g., within a fraction of a second). Stereo or multi-view cameras should predict the same color when seeing the same object.

[0011] Temporal color consistency means that images output by two different cameras when observing a scene at two different times include color values ​​that match as described above. Temporal color consistency can be determined by comparing predictions output by a machine learning system for images acquired by different cameras at different times. Temporal color consistency can be determined by using the predictions output by the machine learning system to determine a trajectory of an object, the trajectory including speed and direction in real-world coordinates. A computing device included in a vehicle can use the object trajectory and data about the fields of view of the cameras included in the vehicle to estimate when the object will leave the field of view of a first camera and enter the field of view of a second camera.

[0012] Determining when an object will leave the field of view of the first camera and enter the field of view of the second camera may include determining the 3D projection of each camera, as pixels may have varying parallax based on distance. In examples, distant objects will have zero or low parallax for an aligned camera, or at least fixed parallax for multi-view stereo. In other examples, a first pixel area in a first camera may be determined to be related to a second pixel area in a second camera by determining corrections for differences in view angle and camera orientation. The goal may be selected to reduce these factors below a level that requires correction and therefore affects the prediction.

[0013] The machine learning system may process the image acquired by the second camera at the estimated time to determine whether the machine learning system will correctly predict the identity and location of the object in the second image. Correctly predicting the object in the image acquired by the second camera at the estimated time may indicate temporal color consistency between the first camera and the second camera. For example, a vehicle traveling at night may acquire color images that include light emitting diode (LED) flicker effects. Acquiring color images at the appropriate estimated time may reduce color LED flicker effects.

[0014] The statistical analysis may include determining the mean and standard deviation of pixel values ​​in an image or portion of an image. The statistical analysis may be used to determine color consistency parameters for an image signal processing (ISP) system that may compensate for differences in image appearance. The ISP system may include hardware and software programs executable on a computing device included in the vehicle that inputs an image and transforms the image using a number of different transformations that may enhance color consistency, as described below with respect to Figure 3 A color consistency parameter is herein meant to refer to a value that determines a transformation performed on an image that may enhance color consistency by changing RGB pixel values ​​in a received color image.

[0015] By compensating for differences in image appearance color consistency, the techniques described herein can reduce differences in machine learning system predictions between images from the same or different cameras. The techniques disclosed herein can compare predictions determined by a machine learning system based on receiving first and second different images of a scene acquired by a first camera and a second camera, respectively, or first and second images acquired by a single camera at a first time and a second time. When the predictions differ, color data from the first image can be compared with color data from the second image. When the color data from the first image differs from the color data from the second image by more than an empirically determined threshold, the color data from the first image can be statistically analyzed to determine updated color consistency parameters that can be included in an image signal processing (ISP) system that processes images from the first camera. Updating the color consistency parameters in this manner can make the color data from the image acquired by the first camera more consistent in color with the color data from the training data set and the color data from the second camera. Making the image data from the first camera more consistent in color with the second camera can make the predictions determined by the machine learning system based on the image from the first camera more consistent with the predictions determined by the machine learning system based on the second camera.

[0016] A method is disclosed herein, the method comprising determining a first prediction using a machine learning system based on receiving a first image from a first camera, and determining a second prediction using the machine learning system based on receiving a second image from a second camera. When the first prediction is not equal to the second prediction within a user-determined tolerance: determining color consistency based on comparing pixel values ​​from the first image to a threshold determined based on previously determined pixel values; determining color consistency parameters for inclusion in an image signal processing system by determining pixel statistics based on pixel values ​​from the first image; and applying the color consistency parameters to the third image by receiving a third image from the first camera at the image signal processing system. The threshold may be determined by varying pixel values ​​in a training data set image input to the machine learning system to determine when a prediction output changes based on pixel values. The pixel statistics may include a pixel mean and a pixel standard deviation. The color consistency parameters may include one or more of lens shading, white balance, defective pixels, denoising, color interpolation, edge enhancement, color correction matrix, brightness / contrast, and gamma.

[0017] Color consistency may be based on determining one or more of camera color consistency, spatial color consistency, and temporal color consistency on the image. Camera color consistency may be determined by comparing images acquired by different cameras observing the same scene at the same illumination. Spatial color consistency may be determined by comparing overlapping images acquired by different cameras observing parts of the same scene at different illuminations. Temporal color consistency may be determined by comparing images acquired by the cameras observing the same scene at different times. The one or more first predictions and the second prediction may include one or more of an object identity and an object position. The object identity may include one or more of a road or a vehicle. The machine learning system may include a convolutional neural network including a convolutional layer and a fully connected layer. A red, green, blue (RGB) color space image may be converted into a luminance, red projection, blue projection (YUV) image and output. The machine learning system may be included in a mobile machine. The mobile machine may operate based on one or more predictions.

[0018] Also disclosed is a computer readable medium storing program instructions for performing some or all of the above method steps. Also disclosed is a computer programmed to perform some or all of the above method steps, including a computer device programmed to: determine a first prediction using a machine learning system based on receiving a first image from a first camera, and determine a second prediction using the machine learning system based on receiving a second image from a second camera. When the first prediction is not equal to the second prediction within a user-determined tolerance: determine color consistency based on comparing pixel values ​​from the first image with a threshold determined based on previously determined pixel values; determine color consistency parameters for inclusion in an image signal processing system by determining pixel statistics based on pixel values ​​from the first image; and apply the color consistency parameters to the third image by receiving a third image from the first camera at the image signal processing system. The threshold can be determined by varying the pixel values ​​in a training data set image input to the machine learning system to determine when the predicted output changes based on the pixel values. The pixel statistics can include a pixel mean and a pixel standard deviation. Color consistency parameters may include one or more of lens shading, white balance, defective pixels, denoising, color interpolation, edge enhancement, color correction matrix, brightness / contrast, and gamma.

[0019] The instructions may include other instructions, wherein color consistency may be based on determining one or more of camera color consistency, spatial color consistency, and temporal color consistency on the image. Camera color consistency may be determined by comparing images acquired by different cameras observing the same scene at the same illumination. Spatial color consistency may be determined by comparing overlapping images acquired by different cameras observing parts of the same scene at different illuminations. Temporal color consistency may be determined by comparing images acquired by the cameras observing the same scene at different times. The one or more first predictions and the second prediction may include one or more of an object identity and an object position. The object identity may include one or more of a road or a vehicle. The machine learning system may include a convolutional neural network including a convolutional layer and a fully connected layer. A red, green, blue (RGB) color space image may be converted into a luminance, red projection, blue projection (YUV) image and output. The machine learning system may be included in a mobile machine. The mobile machine may operate based on one or more predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a block diagram of an example vehicle system.

[0021] Figure 2 is an illustration of an example vehicle including a camera.

[0022] Figure 3 is a graphic representation of example vehicle image data.

[0023] Figure 4 is another illustration of example vehicle image data.

[0024] Figure 5 is another illustration of example vehicle image data.

[0025] Figure 6 is a diagram of an example image signal processing system.

[0026] Figure 7 is a flow chart of an example process for generating camera, spatial, and temporal correction data.

[0027] Figure 8 is a flow chart of an example process for operating a vehicle based on calibrated camera data. DETAILED DESCRIPTION

[0028] Figure 1 1 is a diagram of a vehicle computing system 100. The vehicle computing system 100 includes a vehicle 110, a computing device 115 included in the vehicle 110, and a server computer 120 remote from the vehicle 110. The computing device 115 of one or more vehicles 110 may receive data regarding the operation of the vehicle 110 from sensors 116. The computing device 115 may operate the vehicle 110 based on the data received from the sensors 116 and the data received from the remote server computer 120. The server computer 120 may communicate with the vehicle 110 via a network 130.

[0029] The computing device 115 includes a processor and memory such as are known. In addition, the memory includes one or more forms of computer-readable media and stores instructions that can be executed by the processor to perform various operations (including operations as disclosed herein). For example, the computing device 115 may include programming to operate one or more of vehicle propulsion (i.e., controlling the speed and / or changes in speed in the vehicle 110 by controlling one or more of an internal combustion engine, an electric motor, a hybrid engine, etc.), steering, climate control, interior lights, and exterior lights, etc., and determine whether and when the computing device 115 (rather than a human operator) controls such operations. The computing device 115 can also control the time alignment of lighting with sensor acquisition to take into account the color effects of vehicle lights or exterior lights.

[0030] The computing device 115 may include more than one computing device (i.e., controllers included in the vehicle 110 for monitoring and controlling various vehicle components, etc. (i.e., propulsion controller 112, steering controller 114, etc.)), or be communicatively coupled to the more than one computing device via a vehicle communication bus as further described below. The computing device 115 is typically arranged for communication over a vehicle communication network (i.e., including a bus in the vehicle 110, such as a controller area network (CAN), etc.); the vehicle 110 network may additionally or alternatively include, for example, known wired or wireless communication mechanisms, such as Ethernet or other communication protocols.

[0031] The computing device 115 may transmit and receive messages to and from various devices (i.e., controllers, actuators, sensors (including sensor 116), etc.) in the vehicle 110 via the vehicle network. Alternatively or additionally, where the computing device 115 actually includes multiple devices, the vehicle communication network may be used for communication between the devices represented as the computing device 115 in this disclosure. In addition, as mentioned below, various controllers or sensing elements (such as sensor 116) may provide data to the computing device 115 via the vehicle communication network.

[0032] Additionally, computing device 115 may be configured to communicate with remote server computer 120 (i.e., cloud server) via network 130 through vehicle-to-infrastructure (V2I) interface 111, as described below, which interface includes hardware, firmware, and software that permit computing device 115 to communicate with remote server computer 120 (i.e., cloud server) via network 130, such as wireless Internet. The V2X interface 111 may thus include a network 130 configured to communicate with the remote server computer 120 using various wired and wireless networking technologies (i.e., cellular, The computing device 115 may be configured to communicate with other vehicles 110 through a V2X (vehicle to the outside world) interface 111 using a vehicle-to-vehicle (V2V) network (i.e., based on cellular communication (C-V2X) wireless communication cellular, dedicated short-range communication (DSRC) and similar communications) formed between neighboring vehicles 110 on the basis of a mobile ad hoc network or through an infrastructure-based network. The computing device 115 also includes a non-volatile memory such as known. The computing device 115 can record data by storing the data in the non-volatile memory for later retrieval and transmission to the server computer 120 or the user mobile device 160 via the vehicle communication network and the vehicle to infrastructure (V2I) interface 111.

[0033] As already mentioned, programming for operating one or more vehicle 110 components (i.e., steering, propulsion, etc.) without human operator intervention is typically included in instructions stored in memory and executable by a processor of the computing device 115. Using data received in the computing device 115 (i.e., sensor data from sensors 116, data from server computer 120, etc.), the computing device 115 can make various determinations and control various vehicle 110 components and operations. For example, the computing device 115 may include programming to or control vehicle 110 operating behaviors (i.e., the physical manifestations of vehicle 110 operation), such as speed, changing speed, steering, etc., and strategic behaviors (i.e., controlling operating behaviors in a manner that is generally intended to achieve efficient traversal of a route), such as distance between vehicles and the amount of time between vehicles, lane changes, minimum gaps between vehicles, left turn crossing path minimums, arrival times at specific locations, and minimum times from arrival at an intersection to crossing an intersection (without a signal).

[0034] A controller, as that term is used herein, includes a computing device that is typically programmed to monitor and control specific vehicle subsystems. Examples include a propulsion controller 112 and a steering controller 114. The controller may be an electronic control unit (ECU), such as is known, and may include additional programming as described herein. The controller may be communicatively connected to a computing device 115 and receive instructions from the computing device to actuate the subsystems in accordance with the instructions.

[0035] One or more controllers 112, 113, 114 for the vehicle 110 may include known electronic control units (ECUs) and the like, including, as non-limiting examples, one or more propulsion controllers 112 and one or more steering controllers 114. Each of the controllers 112, 113, 114 may include a corresponding processor and memory and one or more actuators. The controllers 112, 113, 114 may be programmed and connected to a vehicle 110 communication bus, such as a controller area network (CAN) bus or a local interconnect network (LIN) bus, to receive instructions from a computing device 115 and control the actuators based on the instructions.

[0036] Sensors 116 may include a variety of devices such as are known to provide data via a vehicle communication bus. For example, a radar fixed to a front bumper (not shown) of vehicle 110 may provide a distance from vehicle 110 to the next vehicle in front of vehicle 110, or a global positioning system (GPS) sensor disposed in vehicle 110 may provide geographic coordinates of vehicle 110. For example, the distances provided by radar and other sensors 116 and the geographic coordinates provided by the GPS sensor may be used by computing device 115 to operate vehicle 110 autonomously or semi-autonomously.

[0037] The vehicle 110 is typically a ground-based vehicle 110 (i.e., a passenger car, a light truck, etc.) capable of autonomous and semi-autonomous operation and having three or more wheels. The vehicle 110 includes one or more sensors 116, a V2I interface 111, a computing device 115, and one or more controllers 112, 113, 114. The sensors 116 may collect data related to the vehicle 110 and the operating environment of the vehicle 110. By way of example and not limitation, the sensors 116 may include, i.e., an altimeter, a camera, a lidar, a radar, an ultrasonic sensor, an infrared sensor, a pressure sensor, an accelerometer, a gyroscope, a temperature sensor, a Hall sensor, an optical sensor, a voltage sensor, a current sensor, a mechanical sensor (such as a switch), etc. The sensors 116 may be used to sense the operating environment of the vehicle 110, i.e., the sensors 116 may detect phenomena such as weather conditions (rainfall, outside temperature, etc.), road slope, road position (i.e., using road edges, lane markings, etc.), or the position of a target object (such as a neighboring vehicle 110). The sensors 116 may also be used to collect data, including dynamic vehicle 110 data related to the operation of the vehicle 110 , such as speed, yaw rate, steering angle, engine speed, oil pressure, power levels applied to controllers 112 , 113 , 114 in the vehicle 110 , connectivity between components, and accurate and timely performance of components of the vehicle 110 .

[0038] The server computer 120 generally has features in common with the V2I interface 111 and computing device 115 of the vehicle 110, such as a computer processor and memory and configuration for communicating via the network 130, and therefore these features will not be further described to reduce redundancy. The server computer 120 can be used to develop and train software that can be transferred to the computing device 115 in the vehicle 110.

[0039] Figure 21 is an illustration of a vehicle 110 including a central camera 202, a left camera 206, and a right camera 210. The central camera 202, the left camera 206, and the right camera 210 have fields of view 204, 208, 212, respectively. The camera field of view is the space within which the camera can obtain an image. The fields of view 204, 208, 212 are areas within which the central camera 202, the left camera 206, and the right camera 210 can obtain a color image of the environment surrounding the vehicle 110. The acquired color image can be received by a machine learning system included in the computing device 115 and processed to determine a prediction about the environment surrounding the vehicle 110. A color image is defined as a multispectral image that samples the visible range of the electromagnetic spectrum. The color image can simulate the human response to full-color light, or be tuned to enhance a specific portion of the spectrum, such as red light for low-light detection. The prediction determined by the machine learning system can include predictions about the identity of objects included in the environment (such as roads or vehicles) and the location of the objects relative to the vehicle 110. By determining the position of the object relative to the vehicle in the successive images, the motion of the vehicle 110 relative to the road and the motion of the object relative to the vehicle 110 may be determined by the computing device 115 based on the predictions output by the machine learning system.

[0040] The fields of view 204, 208, 212 may include overlapping areas 214, 216 between the center camera 202 and the left camera 206 and between the center camera 202 and the right camera 210, respectively. In an example, the overlap of the fields of view 204, 208, 212 may be determined by the movement of the vehicle or object. The objects included in the overlapping areas 214, 216 may be included in two color images acquired during a certain period of time. The techniques described herein for color consistency can use machine learning system predictions about objects appearing in two images to determine color consistency parameters to be included in an image signal processing (ISP) system, which color consistency parameters can enhance image color consistency and allow the machine learning system to determine similar predictions about objects in the two color images. In this context, similar predictions mean determining the same identity and location within a user-determined tolerance.

[0041] Figure 3 6 is a flow chart of process 300 for updating color consistency parameters in image signal processing system 600 to enhance color consistency. Process 300 may be implemented by image signal processing hardware included in vehicle 115 and image signal processing software executed on computing device 115, for example.

[0042] Figure 31 is a flow chart of a process 300 for enhancing color consistency in a color image. The process 300 may be implemented as an image signal processing (ISP) system including hardware and software executed on a computing device 115 to enhance color consistency by changing RGB pixel values ​​in an image to approximate RGB pixel values ​​in a previously acquired image. The process 300 includes a number of blocks that may be performed in the order shown. Alternatively or in addition, the process 300 may include fewer blocks, and may include blocks that are performed in a different order.

[0043] The process 300 operates on RGB pixels included in an image acquired by a camera included in the vehicle 110. A separate process 300 may receive image data from each camera included in the vehicle 110 and transform the RGB pixel values ​​included in the image according to an image transform as described herein. The image transform may be adjusted by a color consistency parameter determined based on image statistics determined by comparing RGB pixel values ​​in an acquired image to previously acquired images, as follows and with respect to Figure 4-Figure 6 The color consistency parameters include automatic gain control parameters, lens shading parameters, white balance gain parameters, defect pixel parameters, denoising parameters, color interpolation parameters, edge enhancement parameters, color correction matrix parameters, brightness / contrast correction parameters, and gamma correction parameters as discussed below.

[0044] Process 300 begins in box 302, where an image is acquired by a sensor included in a camera (e.g., camera 202, 206, 210) included in vehicle 110. The image sensor may include automatic gain control parameters that can be applied to RGB pixels of the acquired image. Automatic gain control changes the multiplication value applied to each pixel of the image sensor to expand or reduce the dynamic range of the pixel data. In an example where the maximum pixel value in the acquired image is less than the maximum allowable value, the automatic gain control multiplies each pixel by a value greater than 1, which attempts to make the maximum pixel value closer to the maximum allowable value. In an example where the maximum pixel value exceeds the maximum allowable value, the automatic gain control multiplies each pixel by a value less than 1 to make the maximum pixel value closer to the maximum allowable value. The automatic gain control parameter determines how strongly the gain control adjusts the pixel value. For example, an automatic gain control parameter of 1.0 causes the pixel value to rise or fall to the maximum allowable level. In another example, an automatic gain control parameter of 0.5 causes the pixel value to go from an initial value to half of the maximum allowable value, and so on.

[0045] At block 304, the acquired image is corrected for lens shading. Lens shading is a difference in image brightness that is a function of pixel location in the image and is caused by differences in light transmission from different parts of the lens. One form of image shading differences is called "vignetting". Vignetting is a darkening of pixels included in the outer edges of an image, typically caused by lens design, focal length, and / or aperture selection. In vignetting, pixel values ​​decrease as a function of distance from the center of the image. Other forms of lens shading may be caused by lens or sensor misalignment or partial occlusion in front of the camera. Lens shading may be corrected by adding a correction image to the acquired image, which increases the RGB pixel values ​​in the dark portions of the image. The correction image includes a pattern of pixel values ​​in the same X, Y dimensions at the input image, the pattern compensating for brightness deviations in the pixel values. Image shading differences may be determined by acquiring an image of a uniformly illuminated solid color panel (whose color may be gray or white) and analyzing the resulting image. Deviations from the uniform pattern of light in the acquired image may be saved as lens shading parameters including the correction image. Lens shading parameters including the correction image may be stored in the lens shading and added to the acquired image to correct for differences in image brightness by adding the stored correction image to the acquired image.

[0046] At box 306, the acquired image is corrected for white balance. White balance assumes that an image of a pure white panel acquired by the image sensor will generate RGB pixels that, when recombined in an image display device, will produce pixels perceived as white by an observer. If the RGB pixels generated by the white panel acquired by the image sensor 602 are not recombined to form white pixels, correction values ​​can be applied to the R, G, and B channels individually for each pixel. White balance correction factors that cause the RGB channels to combine to form a pixel that appears white can be determined. These white balance correction values ​​can be stored as white balance parameters and applied to each image acquired by the image sensor to correct the white balance in the image by adding the stored R, G, and B values ​​to the R, G, and B channels of the acquired image.

[0047] Next, at box 308, defective pixels can be detected and repaired. A defective pixel is a pixel that includes an RGB value that deviates from neighboring pixels by more than a user-selected amount (e.g., more than 50%). Defective pixels can be "stuck off", meaning that they return a zero brightness value regardless of the input scene brightness, or "stuck on", meaning that they return a maximum brightness value regardless of the input scene brightness value. Defective pixels can be noticeable in the acquired image by appearing as isolated bright or dark pixels. Defective pixels can be determined by acquiring a uniformly dark image and a uniformly bright image and detecting pixels that deviate from uniform brightness or darkness. The location of the detected defective pixel can be stored as a defective pixel parameter and used to correct the defective pixel by replacing the defective RGB pixel value with the average of the R, G, or B pixel values ​​from the neighboring pixels.

[0048] Next, at box 310, electronic noise present in the image data may be reduced. The process of an image sensor capturing photons of light and converting the photons of light into digital RGB pixels may introduce variations in signal strength or electronic noise at several steps in the process. In some examples, the image data may be demosaiced before processing. De-mosaicing refers to preprocessing image data acquired by a sensor including adjacent red, green, and blue filters to acquire RGB image data using a single sensor. Noise may appear in the image data as random variations in RGB pixel values. This noise may be reduced by performing spatial and temporal filtering to smooth the image data, trading a small amount of spatial or temporal resolution for greater color fidelity. The filter may be a 2D neighborhood, such as a 3×3 or 5×5 pixel group around a pixel to be processed. The pixels in the neighborhood may be combined to determine a central tendency statistic that may be used to change the pixel to be processed. Typically, the neighborhood is moved row by row and pixel by pixel across the image to determine a new value for each pixel in the image. The filter parameters include the size and shape of the neighborhood and the type of central tendency filter to be determined. For example, a central tendency filter may include a mean, a mode, and a median. A noise filter can replace pixel values ​​with values ​​based on a central tendency filter calculated based on neighborhood pixels. For example, an intensity parameter determines how much of a difference between a pixel and its neighbors is applied to that pixel. Noise filter parameters can be determined by measuring the variance of color pixel data across an image.

[0049] Next, at block 312, the intermediate color values ​​between adjacent RGB pixels may be interpolated. Because color images are typically acquired using a color filter array placed in front of a CMOS image sensor, the color image may include artifacts based on the color filter array. Color interpolation 312 may reduce artifacts by interpolating color pixels between RGB pixels in the original input image. Color interpolation parameters may be determined by measuring the variance of each color individually. Color interpolation may be performed in the X direction, the Y direction, or both.

[0050] The processing included in the lens shading box 304, the white balance box 306, the defective pixel box 308, and the denoising box 310 is based on a gray world assumption of color, which assumes that each of the RGB color channels averages to gray across the image. Deviations from the gray world assumption at the macro level (entire or large image area) or the micro level (adjacent pixels or small areas) may result in changes in color consistency parameters in one or more of the lens shading box 304, the white balance box 306, the defective pixel box 308, and the denoising box 310, which will remedy the deviation from the gray world assumption. Changes in color consistency parameters based on deviations from the gray world assumption may be corrected by automatic gain control as described above at box 302, which may be returned to box 302 to be stored as automatic gain control parameters. Changes in color consistency parameters based on deviations from the gray world assumption may be returned to the automatic white balance gain box 306 to adjust the white balance gain, which deviation may be corrected by white balance, as described above with respect to box 306.

[0051] At box 314, the edge enhancement filter performs a spatial high pass filter to increase the edge strength in the color image. The high pass filter is a 2D filter from a series of filters that include simple X and Y directional derivatives, edge detectors (such as Sobel, Prewitt or Canny), or higher order filters (such as Laplace). The color consistency parameter applied to the edge enhancement filter can select the type of edge detector and can determine the amount of magnification applied to enhance the edges in the color image. The amount of magnification can be determined by processing the edge magnification image with a machine learning system to determine the amount of magnification that enhances object detection without generating false positive object detections.

[0052] At box 316, a color correction matrix can be applied to the acquired image. The color correction matrix includes parameters that can remap individual RGB color pixels in the image to correct incorrect color mapping of the image sensor. The remapping can use a lookup table, which can be implemented as an array that uses RGB pixel values ​​as indices into three arrays, one for each of the red channel, the green channel, and the blue channel. The lookup table outputs RGB pixel values ​​that can transform the input RGB pixels into a user-selected RGB value. Color correction can compensate for changes in color appearance based on illumination. For example, artificial light can change the appearance of colors in an image depending on the light source. Colors look different in artificial light depending on the technology used to generate the light. By analyzing pixel statistics, the type of lighting can be determined, and a color correction matrix that compensates for that type of lighting can be selected.

[0053] At block 318, RGB pixel values ​​included in the color image may be corrected for overall brightness and contrast. A brightness parameter is a single digital value that may be added to an RGB pixel, and a contrast parameter is a single digital value that may be multiplied by an RGB pixel to adjust the brightness and contrast of the color image. The brightness / contrast parameter may be determined based on RGB pixel statistics.

[0054] At box 320, the color image is passed to gamma correction, where the response of the camera is corrected. The response of the image sensor is the shape of the curve that transforms the energy of photons received by the sensor into electrons. Gamma is corrected by raising the input RGB pixel values ​​to a user-selected power (gamma) and multiplying by a user-selected constant to approximate the visual response of the human eye. Modifying the input RGB pixels using gamma correction will allow the machine learning system to determine predictions about the object that approximate the performance of the human eye. Gamma correction parameters can be determined based on RGB pixel statistics, which are determined based on separating dark (shadow) portions from bright portions and determining statistics for each portion separately. Gamma correction parameters provide image details in both dark and bright portions of the image.

[0055] At box 322, the color image is converted from RGB pixel values ​​to YUV pixel values. A linear function converts the red channel, the green channel, and the blue channel from the RGB color space to a YUV image. The YUV pixel includes an intensity or brightness channel (Y) that more closely approximates the response of the human eye, a first chrominance or red projection channel (U), and a second chrominance or blue projection channel (V), which is more sensitive to intensity than color for processing using a machine learning system. After box 322, the color image output from the ISP system described in process 300 can be output to a machine learning system to determine predictions about objects included in the color image. After box 322, process 300 ends.

[0056] Figure 41 is an illustration of three color images 402, 404, 406 from three cameras included in the vehicle 110. The color images 402, 404, 406 include overlapping views of a traffic scene including roads 408, 410, 412 and backgrounds 414, 416, 418 including buildings and foliage. The color images 402, 404, 406 may be acquired by the cameras as three 8-bit channels of red, green, and blue (RGB) color data, each channel being arranged as a 2D array of RGB pixels that may be accessed by their X, Y addresses in the 2D array. Without loss of generality, the techniques described herein may use color spaces other than RGB, such as luminance, blue difference, red difference (YC B C R ) and Hue, Saturation, Brightness (HSB) are color spaces that can be used to represent color pixels.

[0057] A machine learning system executing on a computing device 115 in the vehicle 110 may receive the color images 402, 404, 406 and process them to determine predictions about traffic scenes included in the color images 402, 404, 406. For example, the machine learning system may determine predictions about the locations of the roads 408, 410, 412 and determine the location of the vehicle 110 relative to the roads 408, 410, 412. The machine learning system may determine predictions about the locations of static objects (such as lane markings, traffic signals, and traffic signs) included in the color images 402, 404, 406. The machine learning system may also determine predictions about the identities and locations of movable objects (such as vehicles) included in the color images 402, 404, 406.

[0058] The computing device 115 may use the predictions output from the machine learning system to determine a vehicle path on which to operate the vehicle 110. The vehicle path may be a polynomial function that includes the time-varying position, direction, and velocity of the vehicle 110. The vehicle path may be analyzed by the computing device 115 to generate an operation to be applied to the vehicle 110 through one or more of vehicle propulsion or vehicle steering that may cause the vehicle 110 to travel along the vehicle path.

[0059] The computing device 115 can use the predictions output by the machine learning system to determine a vehicle path that guides the vehicle 110 to stay in the lane as indicated by the road edge and lane markings, follow the direction indicated by the traffic signal and traffic signs, and maintain the user-determined limit on the distance to be maintained relative to the movable object. The accuracy of the predictions output by the machine learning system in response to receiving the color images 402, 404, 406 can be affected by the consistency of the colors output from each camera. The accuracy of the predictions as used herein means that the predictions correctly identify the objects in the image and correctly locate the objects relative to the vehicle in real-world coordinates. The technology described herein can enhance the accuracy of the predictions output by the machine learning system by enhancing the color consistency in the images acquired by the camera included in the vehicle. The predictions determined based on the images acquired by the sensor included in the vehicle can be compared. When the predictions do not match, for example, when the identity and location of the object do not match, the RGB pixel values ​​included in the image including the unsuccessful prediction can be compared with the RGB values ​​of the image that has been successfully processed by the machine learning system to determine the difference between the image including the successful prediction and the image in which the machine learning system did not successfully predict the object. When the RGB pixel values ​​in an unsuccessfully predicted image differ from the RGB pixel values ​​in a successfully predicted image by more than a user-determined threshold, a statistical metric may be determined based on the RGB values, which may be used to determine color consistency parameters, which may be input into a transformation included in the ISP system to enhance color consistency and enhance the accuracy of predictions determined based on color images.

[0060] The color images 402, 404, 406 show the difference in color consistency between the three cameras included in the vehicle 110 observing the traffic scene. Even if the materials constituting the roads 408, 410, 412 are the same, the RGB pixel values ​​of the portions of the color images 402, 404, 406 included in the roads 408, 410, 412 may be different between the color images 402, 404, 406. In this example, the difference in RGB color values ​​in the roads 408, 410, 412 in the color images 402, 404, 406 may be caused by color inconsistency. As defined above, color inconsistency means that the RGB color values ​​in the two images differ by more than a threshold, such as a threshold determined by a user. If the three color images 402, 404, 406 are acquired at different times, the temporal color inconsistency may cause the RGB pixel values ​​to be different.

[0061] exist Figure 3In the example shown, color inconsistency between roads 408, 410, 412 in color images 402, 404, 406 may be caused by spatial color inconsistency. Color images 402, 404, 406 may be acquired by a central camera 202, a left camera 206, and a right camera 210 included in a vehicle 110. The central camera 202, the left camera 206, and the right camera 210 may be arranged to have three different fields of view 204, 208, 212 and point to three different directions relative to the vehicle 110 and the traffic scene. Due to the angle at which strong sunlight is reflected by the roads 408, 410, 412, the strong sunlight may be reflected directly by the road 408 into the lens of the left camera 206 of the vehicle 110, so that the road 4408 in the image 402 appears bright. Detecting strong direct sunlight may allow the system to reduce specular reflections and lens flare that may cause errors. Because the field of view 204 of the central camera 202 is different from that of the left camera 206, the angle at which the strong sunlight is reflected from the road 410 into the lens of the central camera 202 is different from the angle at which the strong sunlight is reflected into the lens of the left camera 204, making the image of the road 410 in the image 404 darker than the image of the road 408 in the image 402 of the left camera 204. In turn, the angle of the right camera 210 relative to the direction in which the strong sunlight is reflected from the road 412 in the image 406 is more inclined relative to the direction of the strong sunlight, and because the light reflected by the road 4412 in the image 406 is reflected from away from the lens of the right camera 210 rather than the lens of the central camera 202, the road 4412 in the image 406 is even darker than the road 4410 in the image 404. The differences in RGB color pixels in color images 402, 404, 406 due to spatial color inconsistency may cause the machine learning system to determine a prediction to identify and locate roads 408, 410, 412 in all three color images 402, 404, 406 to output different predictions about roads 408, 410, 412. The road predictions may include the locations of road edges, the locations of lane markings, and the centers of lanes, etc. For example, the machine learning system may accurately predict the location of road 4408 in image 402, miss a certain location of road 4410 in image 404, and not predict the identity and location of road 4412 in image 406 at all.

[0062] A threshold color consistency value may be determined based on empirical testing performed using test images and a machine learning system. The threshold color consistency value may be determined by varying the pixel values ​​in a training data set image input to the machine learning system to determine when a prediction output changes based on a pixel value change. For example, a color image including an object may be processed by a machine learning system to correctly predict the identity and location of the object. The image received by the machine learning system may be iteratively altered by changing the RGB pixel values ​​until the machine learning system fails to successfully predict the object. Other color spaces other than RGB may be used to determine the threshold. For example, when hue, saturation, and brightness are altered to determine the impact on prediction accuracy, the HSB color space may be used to determine the difference in hue, saturation, and brightness.

[0063] Techniques for color consistency as described herein use expected predictions to correct RGB color values ​​in a color image. During training, a machine learning system may be trained to make predictions about the identity and location of objects in the environment surrounding the vehicle 110. For example, a machine learning system may be trained to identify and locate roads 408, 410, 412, including road edges, lane markings, traffic signals, and signs. RGB pixel data included in a portion of an image including an object may be examined to determine one or more colors included in the object, and the portion of the image is processed by the machine learning system to determine the prediction that the object is correctly identified and located. For example, a stop sign may include both red and white. RGB pixel values ​​representing a stop sign included in an image may be examined by performing a statistical analysis to group the colors into groups that may be labeled "red" and "white." For example, a maximum likelihood technique is a technique for grouping similar colors together. The color groups may be further processed to generate statistical information about the distribution of colors within the group. Assuming a Gaussian distribution, the color groups may be analyzed to determine the pixel mean and pixel standard deviation for each color group.

[0064] RGB pixel values ​​included in a portion of an image that includes an object that can be identified and located by a machine learning system can be used to determine a nominal value and a threshold value for the RGB pixel values. The threshold value can be determined by incrementally changing the RGB values ​​included in the object and processing the image including the object with the machine learning system. The threshold value has been determined when the change in the RGB value is large enough to stop the machine learning system from identifying and locating the object. In practice, when the machine learning system does not successfully identify and locate the object, the RGB value included in the portion of the image that includes the object can be compared with the threshold value determined for the object using the machine learning system. If the RGB value included in the portion of the image that includes the object is less than the threshold value, the machine learning system may not successfully predict that the image includes the object because the RGB value in the portion of the image that includes the object is different from the RGB value included in the image that includes a similar object in the training data set.

[0065] The computing device 115 may determine that the machine learning system failed to predict an object when the computing device 115 determines that the machine learning system failed to successfully identify and locate an object by comparing the prediction from the first image with the prediction from the overlapping portion of the image that includes the same object. The computing device 115 may also determine that the machine learning system failed to successfully predict the identity and location of the object based on determining the object motion and estimating that the first image should include the object based on predicting the object motion in the second image. Comparing the predictions may also take into account rarely occurring examples, such as water droplets or dirt on the camera lenses. Rarely occurring or temporary examples such as these may affect the camera-to-camera prediction correlation. Comparing the predictions of each camera while the vehicle 110 is moving may indicate areas of the image that are unreliable due to rarely occurring or temporary problems, such as water droplets or dirt.

[0066] When the computing device 115 determines that the machine learning system has not successfully identified and located an object and the RGB values ​​included in the portion of the image including the object are less than a previously determined threshold, a statistical analysis may be performed on the RGB values. The statistical analysis may include determining the mean and standard deviation of the RGB values ​​in the image and comparing them to previously determined nominal mean and standard deviation values ​​to determine changes in color consistency parameters included in the ISP system. Changes in color consistency parameters included in the ISP system may transform the RGB values ​​acquired by the camera into RGB values ​​that are more likely to be successfully processed by the machine learning system to predict the identity and location of the object. The techniques described herein for color consistency processing may be used in conjunction with other image processing techniques, for example, illumination invariant imaging, where color consistency processing may be used as a preprocessing step.

[0067] In an example of a color consistency technique as discussed herein, a statistical analysis of RGB images greater than a threshold may be used to trigger a change in a color consistency parameter without determining a predicted difference output by a machine learning system. In an example of a color consistency technique that relies on a statistical analysis of RGB images without requiring a predicted difference to trigger a change in a color consistency parameter, the statistical analysis of the RGB images would be performed on all input images or periodically performed on a subset of the input images and tested to determine if the difference between the images exceeds a user determined threshold.

[0068] Figure 51 is an illustration of two color images 502, 504 acquired by two cameras included in vehicle 110, which illustrates the spatial color consistency technique. Due to the difference in illumination type and scene type as defined above, the RGB values ​​of the colors in color images 502, 504 may have different appearances. For example, the difference in illumination may be caused by differences in sunlight angles or shadows, differences in artificial illumination (such as headlights or street lights). Differences in the type and position of objects may also result in differences in RGB pixel values. Differences in image color may also be caused by differences in ISP color consistency parameters. Color images 502, 504 include overlapping portions, for example, the right-hand portion of color image 502 overlaps the left-hand portion of color image 504. Because the image of vehicle 506 is included in the overlapping portion of image 502, another image of vehicle 508 appears in image 504. Due to the difference in illumination between color images 502, 504, the RGB pixels included in the portion of image 502 including vehicle 506 may have different values ​​than the RGB pixels included in the portion of image 504 including vehicle 508.

[0069] The difference in RGB pixel values ​​between the images of vehicle 506 and vehicle 508 may cause the machine learning system to successfully predict the identity and location of vehicle 508, and unsuccessfully predict the identity and location of vehicle 506. When the machine learning system outputs a prediction that includes an overlapping portion of a first color image, the color consistency system may check the prediction output by the machine learning system for the overlapping portion of an adjacent color image to ensure that the same prediction is output. If the predictions in the adjacent color images are different or missing, the color consistency system may compare the RGB pixel values ​​in the adjacent image to a user-determined threshold. If one or more RGB pixels in the adjacent image are less than the threshold, a statistical analysis may be performed on the RGB pixels of the adjacent color images to determine changes in color consistency parameters included in an image signal processing (ISP) system included in computing device 115. Changes to the color consistency parameters may adjust the RGB values ​​in the adjacent color images to allow the machine learning system to correctly output a prediction for image 402 that matches the prediction included in color image 504. The above description of Figure 3 The ISP process for achieving color consistency is described.

[0070] Figure 66 is an illustration of four color images 602, 604, 606, 608 illustrating differences in temporal color consistency. A pair of color images 602, 604 are acquired by two cameras included in vehicle 110, and two additional color images 606, 608 are acquired by the same two cameras at a later time. Due to differences in illumination between color image 602 and color image 606, an image of vehicle 610 may have different RGB pixel values ​​than an image of vehicle 612 acquired later. Due to the differences in RGB pixel values, a machine learning system may predict the identity and location of vehicle 610 and fail to predict the identity and location of vehicle 612.

[0071] The temporal color consistency technique may detect the location of vehicle 610 in image 604 and determine a trajectory of vehicle 610 based on a plurality of predictions. Based on the relationship between the field of view of the camera that acquired images 602, 604, 606, 608 and the determined trajectory of vehicle 610, computing device 115 in vehicle 110 may predict that vehicle 610 in color image 604 will appear later in color image 606. If the machine learning system does not correctly predict the identity and location of vehicle 612, then the RGB pixel values ​​in the area predicted to include vehicle 612 may be compared to the above-described relationship when the trajectory of the object determined based on color image 604 would cause it to appear in color image 606 at an estimated time and location based on the field of view of the camera that acquired color images 604 and 606, respectively. Figure 3 The thresholds discussed are compared.

[0072] When RGB pixel values ​​from an image region expected to generate predictions about an object are less than a previously determined threshold, the RGB values ​​may be statistically analyzed to determine variations in color consistency parameters included in the ISP by determining, for example, a mean and a standard deviation. Processing variations in color consistency parameters in the ISP of color images acquired by a camera included in vehicle 110 may allow a machine learning system to correctly predict the identity and location of an object (such as vehicle 612) included in color image 606.

[0073] Figure 7 7 is a flow chart of a process 700 for updating color consistency parameters in an image signal processing system 700 to enhance color consistency. The process 700 may be implemented, for example, in a computing device 115. The process 700 includes a number of blocks that may be performed in the order shown. Alternatively or in addition, the process 700 may include fewer blocks, and may include blocks that are performed in a different order.

[0074] Process 700 begins at block 702, where computing device 115 acquires two or more color images from one or more cameras included in vehicle 110. The color images may be color images of a scene including the same object acquired by different cameras to determine camera consistency, color images of a scene including overlapping portions including the same object to determine spatial color consistency, or color images including objects at two locations in a scene acquired at two different times to determine temporal color consistency. Examples of situations that may affect color consistency include LED flickering or rapidly changing shadows caused by the motion of vehicle 110.

[0075] At box 704, the color image acquired at box 702 is received by a machine learning system to determine predictions about the identity and location of objects included in the color image. As described above, the predictions output from the machine learning system can identify the object categories and object locations included in the image. For example, a machine learning system can input an image and output a prediction, the prediction including an identity equal to "road" and location data specifying the edges of the road and the center of the lane visible in the image. Another machine learning system can output a prediction including an identity equal to "vehicle" and location data describing the location of the vehicle relative to the road. The predictions typically include object identity and location, but may include other properties of the object, such as size and movement. Objects to be predicted include roads and vehicles, etc. The machine learning system may be a trained neural network, and the trained neural network may be a convolutional neural network. The convolutional neural network can receive a color image and output predictions about the identity and location of objects in the color image. Objects included in the color image may include, for example, vehicles.

[0076] At box 706, the computing device 115 may compare predictions about objects that appear in the two color images. When the two color images are acquired by two cameras observing the same field of view from the same direction so that the illumination of the scene and the object are the same, any difference between the predictions output by the machine learning system may be determined to be caused by differences in camera color consistency. When the two color images are acquired by two cameras observing overlapping portions of the scene observed with different fields of view, it may be determined that the difference in predictions is caused by differences in color consistency due to differences in illumination included in the two color images. When the two color images are acquired by two cameras observing the same object at two different times, the difference in predictions may be caused by differences in temporal color consistency. When the first prediction from the first color image is the same as the second prediction from the second color image, the process 700 color images include color consistency, and the process 700 ends. When the first prediction from the first color image is different from the second prediction from the second color image, the process 700 proceeds to box 708.

[0077] At block 708, color image statistics based on RGB pixel values ​​are calculated by the computing device 115 for the first color image and the second color image. The color image statistics may include the mean and standard deviation of the RGB pixel values ​​in the respective images, color sensitivity (which is the number of different colors that can be distinguished), sensitivity measurement index (SMI) (which is the ability to reproduce accurate colors), white balance, color matrix, and overall color sensitivity, e.g., as discussed above. For example, the statistics may be determined for the first color image and the second color image or for portions of the first color image and the second color image (e.g., portions of the first color image and the second color image that include an object).

[0078] At block 710, image statistics from the first color image and the second color image are compared. When the difference between the image statistics from the first color image and the statistics from the second color image is less than a user-determined threshold, process 700 ends. When the difference between the image statistics from the first color image and the statistics from the second color image is greater than the user-determined threshold, process 700 proceeds to block 712.

[0079] At block 712, the image statistics calculated for the second color image are used to determine parameter updates for the ISP system that processes the color image of the second camera. For example, if the average pixel value of the second color image is less than the average pixel value of the first color image, a color consistency parameter may be updated to the brightness / contrast 522 of the ISP system 600 to increase the brightness of the color image acquired by the second camera. In another example, if the standard deviation of the pixel values ​​of the second color image is greater than the standard deviation of the first color image, a color consistency parameter that determines the strength of the noise filter at denoising 610 may be increased to reduce the variance of the pixel values ​​in the color image acquired by the second camera. After block 712, process 700 ends.

[0080] Figure 8 is used to exploit the above Figure 7 A flow chart of a process 800 for operating a vehicle 110 using a machine learning system for color images corrected for color consistency using an ISP system 600 is described. The process 800 may be implemented, for example, in a computing device 115 in the vehicle 110. The process 800 includes a number of blocks that may be performed in the order shown. Alternatively or in addition, the process 800 may include fewer blocks, and may include blocks that are performed in a different order.

[0081] Process 800 begins at block 802 , where the computing device 115 in the vehicle 110 performs camera, spatial, and temporal color consistency corrections by determining color consistency parameters to be included in the ISP system 600 .

[0082] At block 804, computing device 115 acquires color image data from one or more cameras included in vehicle 110 and uses the color image data determined at block 800 and as described with respect to FIG. Figure 6 The color consistency parameters described by the ISP system are used to process them.

[0083] At box 806, the machine learning system may receive the color image corrected for color consistency and process it to determine object identities and object locations of objects in the environment around the vehicle 110.

[0084] At block 808, computing device 115 may receive the identity and location of the object and determine a vehicle path that includes the object. For example, computing device 115 may determine a vehicle path that applies vehicle steering to direct vehicle 110 away from the object or slow vehicle 110. Computing device 115 may transmit commands to vehicle controllers 112, 113, 114 to control one or more of vehicle propulsion or vehicle steering to allow vehicle 110 to travel on the determined vehicle path. After block 808, process 800 ends.

[0085] Any actions taken by the vehicle or a user of the vehicle in response to one or more navigation prompts disclosed herein should comply with all rules and regulations specific to the location (e.g., federal, state, country, city, etc.) and operation of the vehicle. More importantly, any navigation prompts disclosed herein are for illustrative purposes only. Certain navigation prompts may be modified and omitted depending on the context, situation, and applicable rules and regulations. In addition, regardless of the navigation prompts, the user should use good judgment and common sense when operating the vehicle. That is, all navigation prompts (whether standard or "enhanced") should be considered suggestions and should only be followed when it is prudent to do so and consistent with any rules and regulations specific to the location and operation of the vehicle.

[0086] Computing devices such as those described herein typically each include commands that can be executed by one or more computing devices such as those identified above and used to implement the blocks or steps of the processes described above. For example, the process blocks described above can be embodied as computer executable commands.

[0087] The computer executable instructions may be compiled or interpreted by a computer program created using a variety of programming languages ​​and techniques, including but not limited to the following, either singly or in combination: Java TM, C, C++, Python, Julia, SCALA, Visual Basic, Java Script, Perl, HTML, etc. Typically, a processor (i.e., a microprocessor) receives commands from a memory, a computer-readable medium, etc., and executes these commands to perform one or more processes including one or more of the processes described herein. Such commands and other data can be stored in files and transmitted using a variety of computer-readable media. Files in a computing device are typically collections of data stored on a computer-readable medium such as a storage medium, a random access memory, etc.

[0088] Computer-readable media (also referred to as processor-readable media) include any non-transitory (i.e., tangible) media that participate in providing data (i.e., instructions) that can be read by a computer (i.e., by a processor of a computer). Such media can take many forms, including, but not limited to, non-volatile media and volatile media. Instructions can be transmitted via one or more transmission media, including optical fibers, wires, wireless communications, including internals that make up a system bus coupled to a processor of a computer. Common forms of computer-readable media include, for example, RAM, PROM, EPROM, FLASH-EEPROM, any other memory chip or cassette, or any other medium from which a computer can read.

[0089] Unless otherwise expressly indicated herein, all terms used in the claims are intended to be given their ordinary and customary meanings as understood by those skilled in the art. Specifically, unless a claim recites an express limitation to the contrary, the use of singular articles such as "a," "an," "the," and "said" should be interpreted as reciting one or more of the indicated elements.

[0090] The term "exemplary" is used herein in the sense of referring to an example, ie, reference to "exemplary widget" should be interpreted as referring merely to an example of a widget.

[0091] The adverb “approximately” modifying a value or result means that the shape, structure, measurement, value, determination, calculation, etc. may vary from the exactly described geometry, distance, measurement, value, determination, calculation, etc. due to imperfections in materials, machining, manufacturing, sensor measurement, calculation, processing time, communication time, etc.

[0092] In the accompanying drawings, like reference numerals indicate like elements. With respect to the media, processes, systems, methods, etc. described herein, it should be understood that although the steps or blocks of such processes, etc. have been described as occurring according to a sequence in a particular order, such processes may be practiced by the described steps being performed in an order other than the order described herein. It should also be understood that certain steps may be performed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the description of the processes herein is provided for the purpose of illustrating certain embodiments and should in no way be construed as limiting the claimed invention.

[0093] According to the present invention, a system is provided, the system having: a computer, the computer including a processor and a memory, the memory including instructions executable by the processor to perform the following operations: determine a first prediction using a machine learning system based on receiving a first image from a first camera; determine a second prediction using a machine learning system based on receiving a second image from a second camera; when the first prediction is not equal to the second prediction within a user-determined tolerance: determine color consistency based on comparing pixel values ​​from the first image with a threshold determined based on previously determined pixel values; determine color consistency parameters for inclusion in an image signal processing system by determining pixel statistics based on pixel values ​​from the first image; and apply the color consistency parameters to the second image from the second camera using the image signal processing system.

[0094] According to an embodiment, the present invention is also characterized by other instructions for determining a threshold based on a change in pixel value by varying pixel values ​​in a training data set image input to a machine learning system to determine when a predicted output is.

[0095] According to an embodiment, the pixel statistics include a pixel mean and a pixel standard deviation.

[0096] According to an embodiment, the color consistency parameters include one or more of lens shading, white balance, defective pixels, denoising, color interpolation, edge enhancement, color correction matrix, brightness / contrast, and gamma.

[0097] According to an embodiment, the color consistency is based on determining one or more of camera color consistency, spatial color consistency, and temporal color consistency across the images.

[0098] According to an embodiment, camera color consistency is determined by comparing images acquired by different cameras observing the same scene under the same illumination.

[0099] According to an embodiment, spatial color consistency is determined by comparing overlapping images acquired by different cameras observing portions of the same scene under different illuminations.

[0100] According to an embodiment, temporal color consistency is determined by comparing images acquired by said camera observing the same scene at different times.

[0101] According to an embodiment, the one or more first predictions and the second prediction comprise one or more of an object identity and an object location.

[0102] According to an embodiment, the object identity comprises one or more of a road or a vehicle.

[0103] According to an embodiment, the machine learning system comprises a convolutional neural network including a convolutional layer and a fully connected layer.

[0104] According to an embodiment, the present invention is also characterized by other instructions for converting a red, green, blue (RGB) color space image into a luminance, red projection, blue projection (YUV) image and outputting the image.

[0105] According to an embodiment, a machine learning system is included in a mobile machine.

[0106] According to an embodiment, the mobile machine operates based on one or more predictions.

[0107] According to an embodiment, the mobile machine is a vehicle, and the vehicle is operated by controlling one or more of vehicle propulsion, vehicle steering, and vehicle braking.

[0108] According to the present invention, a method includes: determining a first prediction using a machine learning system based on receiving a first image from a first camera; determining a second prediction using a machine learning system based on receiving a second image from a second camera; when the first prediction is not equal to the second prediction within a user-determined tolerance: determining color consistency based on comparing pixel values ​​from the first image with a threshold determined based on previously determined pixel values; determining color consistency parameters for inclusion in an image signal processing system by determining pixel statistics based on the pixel values ​​from the first image; and applying the color consistency parameters to the third image by receiving a third image from the first camera at the image signal processing system.

[0109] In one aspect of the invention, the method includes determining the threshold by varying pixel values ​​in a training data set image input to a machine learning system to determine when a predicted output changes based on the pixel value.

[0110] In one aspect of the invention, the pixel statistics include a pixel mean and a pixel standard deviation.

[0111] In one aspect of the invention, the color consistency parameters include one or more of lens shading, white balance, defective pixels, denoising, color interpolation, edge enhancement, color correction matrix, brightness / contrast, and gamma.

[0112] In one aspect of the invention, color consistency is based on determining one or more of camera color consistency, spatial color consistency, and temporal color consistency across the images.

Claims

1. A method, comprising: determining, with a machine learning system, a first prediction based on receiving a first image from a first camera; determining, with the machine learning system, a second prediction based on receiving a second image from a second camera; When the first prediction is not equal to the second prediction within a user-defined tolerance: determining color consistency based on comparing pixel values ​​from the first image to a threshold value determined based on previously determined pixel values; determining a color consistency parameter for inclusion in an image signal processing system by determining pixel statistics based on pixel values ​​from said first image; as well as The color consistency parameters are applied to the third image by receiving, at the image signal processing system, a third image from the first camera.

2. The method of claim 1 further comprising determining the threshold by varying pixel values ​​in a training data set image input to a machine learning system to determine when a predicted output changes based on the pixel values.

3. The method of claim 1, wherein the pixel statistics include a pixel mean and a pixel standard deviation.

4. The method of claim 1, wherein the color consistency parameters include one or more of lens shading, white balance, defective pixels, denoising, color interpolation, edge enhancement, color correction matrix, brightness / contrast, and gamma. 5 . The method of claim 1 , wherein the color consistency is based on determining one or more of camera color consistency, spatial color consistency, and temporal color consistency across the image.

6. The method of claim 5, wherein camera color consistency is determined by comparing images acquired by different cameras observing the same scene under the same illumination.

7. The method of claim 5, wherein spatial color consistency is determined by comparing overlapping images acquired by different cameras viewing portions of the same scene under different illuminations.

8. The method of claim 5, wherein temporal color consistency is determined by comparing images acquired by the camera observing the same scene at different times.

9. The method of claim 1, wherein the first prediction and the second prediction include one or more of an object identity and an object location.

10. The method of claim 9, wherein the object identity comprises one or more of a road, a vehicle, or a pedestrian.

11. The method of claim 1, wherein the machine learning system comprises a convolutional neural network comprising a convolutional layer and a fully connected layer.

12. The method of claim 1, further comprising converting a red, green, blue (RGB) color space image into a luminance, red projection, blue projection (YUV) image and outputting the image.

13. The method of claim 1, wherein the machine learning system is included in a mobile machine. The method of claim 13 , wherein the mobile machine operates based on the one or more predictions.

15. A method comprising a computer programmed to perform the method of any one of claims 1-14.