System and method for labeling images to train a machine learning model
The system corrects optical indicator errors in autonomous driving systems by user-verified image labeling, enhancing detection accuracy and improving vehicle control.
Patent Information
- Application Number
- JP2024568976
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-20
- Filing Date
- 2023-05-18
- Publication Date
- 2025-06-24
AI Technical Summary
Autonomous driving systems face challenges in accurately labeling images to detect optical indicators on vehicles, leading to incorrect predictions and errors in vehicle behavior analysis.
A system and method for labeling images by acquiring vehicle images, identifying vehicle positions, displaying graphic seals, and receiving user inputs to correct optical indicator statuses, enabling precise training of machine learning models to improve detection accuracy.
Enhances the accuracy of optical indicator detection in autonomous driving systems by correcting errors and updating machine learning models with correct data, leading to improved vehicle control and safety.
Smart Images

Figure 2025519086000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Patent Application No. 63 / 344,303, filed May 20, 2022, entitled "SYSTEMS AND METHODS FOR LABLING IMAGES FOR TRAINING A MACHINE LEARNING MODEL", the disclosure of which is hereby incorporated by reference in its entirety.
[0002] Embodiments of the present disclosure relate to systems and methods for labeling images to train a machine learning model. More specifically, embodiments of the present disclosure relate to systems and methods for training a machine learning model to detect optical indicators on one or more vehicles as part of an autonomous driving system.
Background Art
[0003] An autonomous driving system (e.g., a self - driving system) typically acquires images of road lanes and nearby vehicles and inputs those images into a trained machine learning model to control the vehicle without user input or with limited user input. Machine learning models used in such systems are generally first trained by capturing millions or billions of images and then labeling those images with feature labels that indicate features that will be recognized in the vehicle's surrounding environment. For example, features can include curbs, painted lines, other vehicles, cones, traffic signals, and other items found on the road lane. Once the machine learning model is trained to recognize these features, the machine learning model can be downloaded and stored in the vehicle's memory, thereby enabling the vehicle to be driven in autonomous or semi - autonomous mode.
Summary of the Invention
Means for Solving the Problems
[0004] The innovations recited in the claims each have several aspects, and only one of them alone bears its desirable attributes. Without limiting the claims, some prominent features of the present disclosure are briefly described here.
[0005] One aspect of the present disclosure includes a system for labeling images to train a machine learning model to detect optical indicators on a vehicle. The system includes acquiring an image of one or more vehicles on a road, identifying the position of each of the one or more vehicles, displaying a graphic seal on each of the one or more vehicles to indicate that the vehicle has been detected by the system, and receiving an indication of whether the optical indicator is active or inactive in each of the one or more vehicles to label the image for the machine learning model.
[0006] In this system, acquiring an image may include acquiring images from a plurality of vehicles having an autonomous driving system, and acquiring an image may further include acquiring images of the plurality of vehicles when the autonomous driving system determines that optical indicator detection has been inappropriately determined by the autonomous driving system.
[0007] In this system, identifying the position of each of the one or more vehicles may include identifying the vehicle in the image and determining the graphic coordinates of the vehicle in the image.
[0008] In this system, displaying a graphic seal on each of the one or more vehicles may include displaying a bounding box around each of the one or more vehicles in the acquired image.
[0009] In this system, identifying the position of each of the one or more vehicles may include performing image segmentation on the acquired image, and the image segmentation generates a region of each acquired image corresponding to the vehicle.
[0010] In this system, receiving an indication of whether the optical indicator is active or inactive may include receiving a mouse selection from a user who labels the vehicle as having an active or inactive optical indicator.
[0011] In this system, receiving an indication of whether the optical indicator is active or inactive may include receiving an indication of whether the brake light is active or inactive.
[0012] In this system, receiving an indication of whether the optical indicator is active or inactive may include receiving an indication of whether the turn signal is active or inactive.
[0013] Another aspect of the present disclosure includes a system for labeling images to train a machine learning model to detect optical indicators on a vehicle. The system includes acquiring an image of one or more vehicles on a road, identifying the position of each of the one or more vehicles in the acquired image, determining whether the optical indicator has been indicated as active or inactive by the autonomous driving system of each of the one or more vehicles, determining from the image of the one or more vehicles one or more vehicles having an incorrect prediction as to whether the optical indicator was active or inactive, and labeling an image having an incorrect prediction with the correct indication as to whether the optical indicator is active or inactive.
[0014] In this system, identifying the position of each of the one or more vehicles may include identifying the vehicle in the image and determining the graphic coordinates of the vehicle in the image.
[0015] In this system, acquiring images may include acquiring images from a plurality of vehicles having autonomous driving systems. Acquiring images may further include acquiring images from a plurality of vehicles when the autonomous driving system determines that the optical indicator detection has been inappropriately determined by the autonomous driving system.
[0016] In this system, a graphic seal is displayed on each of one or more vehicles to indicate that the vehicle can be detected by the system. Displaying a graphic seal on each of one or more vehicles may also include displaying a bounding box around each of the one or more vehicles in the acquired image.
[0017] In this system, an indication of whether the optical indicator of each of one or more vehicles is active or inactive can be predicted by the autonomous driving system of each vehicle.
[0018] In this system, an incorrect prediction may represent a discrepancy between the optical indicator and the position of the vehicle.
[0019] This system may further include receiving an updated optical indicator that receives a mouse selection from a user who labels a vehicle with an optical indicator based on the position of the vehicle.
[0020] In this system, an indication of whether the optical indicator is active or inactive may be an indication of whether the brake light is active or inactive.
[0021] In this system, an indication of whether the optical indicator is active or inactive may be an indication of whether the turn signal is active or inactive.
[0022] Another aspect of the present disclosure includes a method for labeling images to train a machine learning model to detect optical indicators on a vehicle. The method includes obtaining an image of one or more vehicles on a road, identifying the position of each of the one or more vehicles, labeling an indication of whether the optical indicator is active or inactive in each of the one or more vehicles, determining from the image of the one or more vehicles one or more vehicles having an incorrect prediction, and receiving an updated indication of whether the optical indicator is active or inactive on a vehicle having an incorrect prediction.
[0023] In this method, identifying the position of each of the one or more vehicles may include identifying the vehicles in the image and determining the graphic coordinates of the vehicles in the image.
[0024] In this method, obtaining the image may include obtaining the image from a plurality of vehicles having an autonomous driving system.
[0025] In this method, obtaining the image may include obtaining the image from one or more vehicles when optical indicator detection is inappropriately determined by the autonomous driving system of each vehicle.
[0026] For the purpose of summarizing the present disclosure, certain aspects, advantages, and novel features of the innovation are described herein. It should be understood that not all of such advantages may necessarily be achieved according to any particular embodiment. Thus, the innovation may be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein.
Brief Description of the Drawings
[0027] Embodiments of the present disclosure will be described by way of non-limiting example with reference to the accompanying drawings.
[0028]
Figure 1
[0029]
Figure 2
[0030]
Figure 3A
[0031]
Figure 3B
[0032]
Figure 4
[0033]
Figure 5
[0034]
Figure 6A
[0035]
Figure 6B
[0036]
Figure 7
[0037]
Figure 8
[0038]
Figure 9
[0039]
Figure 10A
Figure 10B
Figure 10C
Figure 10D
[0040]
Figure 11
[0041] Certain preferred embodiments and examples are disclosed below, but the subject matter of the present invention extends beyond the specifically disclosed embodiments to other alternative embodiments and / or uses, and to their modifications and equivalents. Accordingly, the claims appended hereto are not limited by any of the specific embodiments described below. For example, in any method or process disclosed herein, the acts or operations of the method or process may be performed in any suitable order and are not necessarily limited to any particular disclosed order. The various operations may be described serially as a plurality of individual operations in a manner that may be helpful in understanding the particular embodiments, but the order of description should not be construed as meaning that these operations are order-dependent. Further, the structures, systems, and / or devices described herein may be embodied as integrated components or as separate components. For purposes of comparing the various embodiments, certain aspects and advantages of these embodiments are described. Not all such aspects or advantages are necessarily achieved by any particular embodiment. Thus, for example, the various embodiments may be carried out in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as may be taught or suggested herein.
[0042] One or more aspects of the present application correspond to systems and methods for training a machine learning model associated with an autonomous driving system. Using an exemplary machine learning model, nearby vehicles can be detected and it can be determined whether the nearby vehicles have light indicators or signals that can be detected. In some embodiments, the light signal can be a brake light, a turn indicator, a headlight, or any other lit indicator on the vehicle. The light signal can also be a brake light, a turn indicator, etc. associated with a trailer connected to the vehicle. In some embodiments, detectable light indicators such as traffic signals, flashing stop signals, or other lit signals on a typical roadway can be on the roadway. Based on the determined light indicator, the autonomous driving system can predict the driving route, speed, etc. of nearby detected vehicles.
[0043] Embodiments of the disclosed technology correspond to systems and methods for training a machine learning model by more accurately labeling light indicators in captured images (e.g., images obtained from an image sensor of a camera disposed on a vehicle) from vehicle or roadway features. More specifically, the system and method are used to obtain images captured by a camera attached to a vehicle when the vehicle is traveling on a roadway. These captured images can then be uploaded to a server or external system so that various features in the images can be labeled. The uploaded images can be presented to a user (e.g., a human user, a software agent) so that the user can identify and label the state of the light indicators seen in the captured images for use as training data. For example, the captured image can be of a vehicle with a left turn signal illuminated. The user can select, via a user interface, that the left turn signal is illuminated and then store that label along with the figure for use in training an autonomous or semi-autonomous machine learning model such as a vision model. As another example, the image can be of a traffic signal and the user can label the figure to indicate that the traffic has a red light illuminated. The terms image and video clip are used interchangeably throughout this disclosure and these terms have similar meanings. For example, if a set of 300 continuously captured images is available, a 10-second video clip can be played at a rate of 30 fps. Thus, 300 captured images can have the same meaning as a 10-second video clip. The number of images and video clip rate are provided as merely examples and various numbers of images and rates can be used based on a particular application.
[0044] In some embodiments, the labeling system used by a user to label images may include specific elements to increase the accuracy of labeling. For example, the system can automatically outline each vehicle in the image with a graphic such as a bounding box so that the user can select a specific vehicle to be labeled. The user can select the bounding box around the vehicle (e.g., via an interactive user interface) and can then be presented with various options for labeling the optical indicators on that vehicle. The options can include a left turn signal, a right turn signal, a brake light, or similar features of the vehicle. This allows the user to increase the accuracy of the labeling process and improve the ability of the images to train a machine learning model to identify optical indicators of vehicles on the road, enabling the user to label multiple vehicles in a single captured image with different features.
[0045] In some embodiments, the vehicle that captures and uploads an image can upload only the images in which an error in the optical indicator prediction is detected. For example, the vehicle is running autonomous driving software and may identify in a captured image that the vehicle in front is not turning on its brake light. However, the vehicle may also detect that the vehicle in front is decelerating due to traffic. In that situation, since the brake light is likely to be on, the captured image identified as not having the brake light on can be uploaded to the server for manual labeling of the brake light to improve future models for autonomous driving.
[0046] In some embodiments, the vehicle uploading the image may be running autonomous software in stealth mode, where the vehicle is not driving in autonomous mode, but still, the vehicle captures the image as if the system were controlling the vehicle and determines the vehicle's actions. In this stealth mode, the vehicle can identify potential errors in the processing of optical indicators and upload images that led to potential errors to the server for user processing, review, and updated labeling.
[0047] To resolve errors in the autonomous driving system related to optical indicator determination, the machine learning model can be trained by updating the machine model with correct data by the methods described herein. Exemplarily, incorrect optical indicator data (e.g., an image or video clip) based on the machine learning model can be corrected by receiving correct optical indicator data. For example, the correct optical indicator data can be overlaid on the incorrect optical indicator data, and the overlaid data can be used to train the machine learning model. Training can include updating or modifying multiple parameters and attributes related to the machine learning model.
[0048] Here, various aspects of machine learning model training will be described with respect to specific examples and embodiments that are merely intended to illustrate. The examples and embodiments described herein focus on specific calculations and algorithms for illustrative purposes, but those skilled in the art will understand that the examples are merely illustrative and not intended to be limiting.
[0049] FIG. 1 is a block diagram showing an embodiment of system 100. System 100 may include a network connecting several vehicles 110, a machine learning training system 120, and a verification computing device 130. Exemplarily, various aspects associated with the machine learning training system 120 can be implemented as one or more components associated with one or more functions or services. These components may correspond to software modules implemented or executed by one or more external computing devices, which may be separate stand-alone external computing devices. Thus, the components of the machine learning training system 120 should be regarded as a logical representation of the service and do not require a specific implementation on one or more external computing devices.
[0050] As shown in FIG. 1, network 160 connects vehicle 110 and verification computing device 130 to machine learning training system 120. Network 160 can include any combination of wired and / or wireless networks such as, for example, one or more direct communication channels, local area networks, wide area networks, personal area networks, and / or the Internet. In some embodiments, network 160 can include one or more wireless networks such as a Global System for Mobile Communications (GSM) network, a Code Division Multiple Access (CDMA) network, a Long Term Evolution (LTE) network, 5G communications, or any other type of wireless network. Network 160 can use protocols and components for communicating via the Internet or any of the other aforementioned types of networks. For example, protocols used by network 160 can include the Hypertext Transfer Protocol (HTTP), HTTP Secure (HTTPS), Message Queuing Telemetry Transport (MQTT), Constrained Application Protocol (CoAP), and the like. Protocols and components for communicating via the Internet or any of the other aforementioned types of communication networks are well known to those of ordinary skill in the art and are therefore not described in more detail herein. In some embodiments, wireless communication via network 160 can be performed over one or more secure networks, such as communicating encrypted data via SSL (e.g., 256-bit, military-grade encryption). The various communication protocols described herein are merely examples and the present application is not limited thereto.
[0051] The vehicle 110 of FIG. 1 can be connected to the machine learning training system 120. The vehicle 110 can be a set of multiple vehicles. In some embodiments, each of the vehicles 110 is configured to capture its surrounding image including nearby vehicles, traffic signals, the surrounding environment, etc. The captured images are encoded as video files based on the resolution specifications of each camera and can be transmitted (e.g., uploaded) to the machine learning training system 120 via the network 160. In some embodiments, each vehicle 110 can include one or more microprocessors and circuits configured to establish a wireless communication channel for connecting to the network 160. To establish the wireless communication channel, each of the vehicles 110 can periodically (or continuously) scan and detect any nearby wireless signals. In another embodiment, the operator of the vehicle 110 can manually establish a wireless connection and connect to the network 160. For example, since the operator can access a nearby Wi-Fi router, the vehicle 110 is wirelessly connected to the network 160.
[0052] The machine learning training system 120 of FIG. 1 can train the machine learning model 124 and provide the model to the vehicle 110 for use in autonomous driving or semi-autonomous driving. Exemplarily, the machine learning training system 120 can include a machine learning model 124, a routing component 126, and a network server 128. The network server 128 is configured to store the captured images received from the vehicle 110.
[0053] The machine learning model 124 can be part of the machine learning training system 120 as shown in FIG. 1. In some embodiments, the machine learning model 124 is included in the machine learning training system 120. In other embodiments, the machine learning model 124 is a stand-alone component and is interconnected with other components within the machine learning training system such as the network server 128.
[0054] In some embodiments, the machine learning model 124 is configured to identify features in the captured images stored in the network server 128. For example, the features may include curbs, painted lines, other vehicles, cones, traffic signals, and other items found on roadways. Accordingly, the machine learning model 124 can be, or include, a vision-only model such as a convolutional neural network, a transformer network, a fully-connected network, combinations thereof, etc.
[0055] Among the features, the machine learning model 124 can be configured to identify the light indicators of surrounding vehicles (or surrounding vehicles captured by the front camera of the vehicle 110) located in front of the vehicle 110. The identified light indicators can be displayed on the vehicles included in the image.
[0056] In the example of FIG. 1, the verification computing device 130 is connected to the machine learning training system 120 via the network 160. In some embodiments, one or more authorized analysts, including managers, developers, supervisors, administrators, etc., can use the verification computing device 130 to access the network server 128. The verification computing device 130 can be any computing device such as a desktop, laptop, personal computer, tablet computer, wearable computer, server, personal digital assistant (PDA), hybrid PDA / cell phone, cell phone, smartphone, set-top box, voice command device, digital media player, etc. The verification computing device 130 can execute an application (e.g., a browser, a stand-alone application, etc.) that enables a user to access the interactive user interface described herein and view images, analyses, aggregated data, etc. Further, the verification computing device 130 can have a display and an input device through which a user can interact with the user interface components.
[0057] The verification computing device 130 can be configured to access the network server 128 via the network 160 and download one or more images or video clips stored in the network server 128. In some embodiments, the verification component device 130 is configured to identify (or predict) the light indicators of the vehicle in the downloaded image or video clip. When identifying the light indicators of the vehicle, the verification component device 130 can be configured to use one or more attributes or algorithms stored in the machine learning model 124. For example, an analyst can download a captured image from the network server 128 and execute instructions on the machine learning model to identify the light indicators on the vehicle included in the downloaded image.
[0058] In some embodiments, the verification computing device 130 is configured to determine whether the machine learning model correctly identifies the features in the captured image. In these embodiments, the analyst can use the verification computing device 130 to determine whether the machine learning model 124 correctly identifies the light indicators of the surrounding vehicles in the captured image. For example, the analyst can analyze the image or video clip to determine whether there is a discrepancy between the identified light indicators of the vehicle in the captured image and the actual light indicators and driving route of the vehicle.
[0059] In some embodiments, the validation computing device 130 may be configured to flag those images after determining that the optical indicators of one or more vehicles in the image have been misjudged. An analyst can correct the flagged images. In some embodiments, the corrected images can be uploaded to the network server 128. In some embodiments, the analyst may use the corrected images as training data for training the machine learning model 124. For example, the training data can be supplied to the machine learning model 124. In this example, the machine learning model can update or modify its algorithms or attributes associated with the trained machine learning model. The trained machine learning model can be provided to the vehicle 110 via the routing component 126. Thus, the vehicle 110 can execute the model, such as by calculating a forward path based on the input of the image.
[0060] FIG. 2 is a schematic diagram showing an example of the vehicle 110. FIG. 2 shows a top view of the vehicle 110 showing the arrangement of a plurality of image sensors or cameras 220, 230, 240 (for example, cameras configured to be mounted at any internal or external vehicle position). In some embodiments, the vehicle 110 is configured to capture surrounding images. In some embodiments, the vehicle 110 has an autonomous driving function (for example, automatic driving). In some embodiments, the cameras are arranged at various positions inside or outside the vehicle 110. Exemplarily, in FIG. 2, the front camera 220 is mounted on the front side of the vehicle 110, such as above the front windshield. The pillar camera 230 is mounted on both sides of the vehicle 110, such as on the pillar of the vehicle 110. For example, the pillar camera 230 can be mounted inside the pillar. The repeater camera 240 is mounted on both repeater sides of the vehicle 110.
[0061] In some embodiments, cameras 220, 230, 240 capture images of the vehicle surrounding the lane and vehicle 110. In these embodiments, the front camera 220 captures a front image of the vehicle 110. The pillar camera 230 is configured to capture images on both sides of the vehicle 110. The repeater camera 240 is configured to capture a rear image of the vehicle 110.
[0062] In some embodiments, vehicle 110 includes at least one controller having one or more microprocessors and circuits configured to establish a wireless communication channel connected to network 160. The controller may transmit (e.g., supply or upload) the captured images to network server 128 via network 160. Also, the captured images may be encoded as video files based on the resolution specifications of each camera and transmitted to network server 128.
[0063] In some embodiments, vehicle 110 includes a vehicle autonomous driving system 210. The vehicle autonomous driving system 210 can control the vehicle 110 for autonomous driving (e.g., automatic driving). The autonomous driving system 210 can access the captured images and identify surrounding features based on the machine learning model provided by the machine learning training system 120. For example, the features may include the light indicators of each surrounding vehicle displayed on the image captured by the front camera 220. The features may also include road information such as curbs, painted lines, cones, traffic signals, and other items found on the road. The communication configuration between cameras 220, 230, 240 and the autonomous driving system 210 may be either direct or indirect communication via a wired connection using a communication cable or bus. Various wired communication networks such as a controller area network (CAN) can be used, and the network protocol can be specified based on specific applications.
[0064] FIG. 3A is a block diagram showing an embodiment of the architecture of the autonomous driving system 210. The general architecture of the autonomous driving system 210 includes the configuration of computer hardware and software components that can be used to implement the embodiments of the present disclosure. As shown, the autonomous driving system 210 includes a processing unit 302, an input / output device interface 304, a computer-readable medium 306, and a network interface 308, all of which can communicate with each other via a communication bus. The components of the autonomous driving system 210 can be physical hardware components installed within the vehicle 110.
[0065] The input / output device interface 304 can provide connectivity to the cameras 220, 230, 240. Thus, the input / output device interface 304 can receive captured images or video files from the cameras 220, 230, 240. The received images or video files can be stored in the computer-readable medium 306. The computer-readable medium 306 can be an internal or external drive and can communicate with the memory 310.
[0066] Memory 310 may include computer program instructions that are executed by processing unit 302 to implement one or more embodiments. Memory 310 generally includes RAM, ROM, or other persistent or non-transitory memory. Memory 310 can store an operating system 314 that provides computer program instructions for use by processing unit 302 in the general management and operation of autonomous driving system 210. Memory 310 may further include computer program instructions and other information for implementing aspects of the present disclosure. For example, memory 310 includes a detected vehicle input component 316 configured to acquire captured images or video files from cameras 220, 230, 240. Memory 310 further includes an autonomous driving model 318 configured to provide an autonomous driving function by identifying features around vehicle 110. For example, features can include curbs, painted lines, other vehicles, cones, traffic signals, and other items found on roadways. In some embodiments, machine learning model 124 can be supplied to autonomous driving model 124 via network 160, such that autonomous driving model 318 uses the attributes, parameters, and algorithms implemented in machine learning model 124. In some embodiments, autonomous driving model 318 can be updated with a trained machine learning model.
[0067] In some embodiments, the processing unit 302 can also communicate with the memory 310 and further provide output information about the autonomous vehicle operation via the input / output device interface 304. Exemplarily, the process unit 302 can receive optical instructions for each vehicle identified by the autonomous driving model 318. In response to receiving the identified optical instructions of the vehicle, the process unit 302 can execute one or more commands to the autonomous driving system 210 to adapt its autonomous driving based on the optical instructions. For example, after obtaining the detected vehicle from the detected vehicle input component 316, the autonomous driving model 318 determines, based on a plurality of machine learning attributes, that one of the detected vehicles has turned on a right turn signal and has identified a right turn signal instruction. In this example, the processing unit 302 can execute commands to the autonomous system to reduce the speed of vehicle 110 or steer vehicle 110 in a specific direction.
[0068] The network interface 308 can provide connectivity to one or more networks or computing systems, such as network 160 of FIG. 1. In some embodiments, the processing unit 302 executes sending and receiving data to and from the network server 128 via the network interface 308.
[0069] Figure 3B shows an embodiment of the architecture of the verification computing device 130 (as shown in Figure 1). The general architecture of the verification computing device 130 includes the configuration of computer hardware and software components that can be used to implement the embodiments of the present disclosure. As shown, the verification computing device 130 includes a processing unit 322, an input / output device interface 324, a computer-readable medium 326, and a network interface 328, all of which can communicate with each other via a communication bus. One or more authorized analysts, including managers, developers, supervisors, administrators, etc., can use the verification computing device 130 to execute instructions related to one or more of the embodiments of the present disclosure.
[0070] The input / output device interface 324 can provide connectivity to the network server 128. Thus, the processing unit 322 can access the network server 128 to send and receive data via the input / output device interface 324. In some embodiments, the data received from the network server 128 is stored in the computer-readable medium 326. The computer-readable medium 326 can be an internal or external drive and can communicate with the memory 310.
[0071] Memory 330 may include computer program instructions that are executed by processing unit 322 to implement one or more embodiments. Memory 330 generally includes RAM, ROM, or other persistent or non-transitory memory. Memory 330 can store an operating system 334 that provides computer program instructions for use by processing unit 332 in the general management and training of machine learning model 124. Memory 330 may further include computer program instructions and other information for implementing aspects of the present disclosure. For example, memory 330 includes an input processing component 336, a graphic seal overlay component 338, a light indicator display component 340, a machine learning model verification component 342, and a machine learning model training component 344.
[0072] The input processing component 336 of FIG. 3B is configured to obtain a captured image around the vehicle 110. The input processing component 336 is configured to access the captured images of surrounding vehicles, where the images are stored in the network server 128. The surrounding images of the vehicle 110 can be captured using cameras attached to each of the vehicles 110 and transmitted to the network server 128. The captured images can be used by the machine learning model 124 for autonomous driving (e.g., self-driving) to identify one or more features in the surrounding environment of the vehicle. For example, the features can include curbs, painted lines, other vehicles, cones, traffic signals, and other items found on the roadway. In some embodiments, the captured images are used to train the machine learning model 124. For example, in response to determining that the machine learning model has incorrectly identified one of the features in the image, an analyst can use the verification computing device 130 to label the image to correct the identified feature and use the labeled image to train the machine learning model.
[0073] The graphic seal overlay component 338 can be configured to generate a graphic seal for each vehicle in the captured image. For example, the graphic seal may be included in or presented on an interactive user interface that presents the image captured by the vehicle. The graphic seal can be box-shaped and can be overlaid on top of the vehicle in the captured image. In some embodiments, the graphic seal represents one or more semantics associated with the vehicle. For example, the graphic seal can represent the identified light indicators of the vehicle in the captured image, such as whether the light indicators of the vehicle have been identified (by the machine learning model 124). In another example, the graphic seal can represent the type of light indicator identified by the machine learning model 124. The graphic indicator can also be used to label one or more semantics associated with the vehicle.
[0074] For example, an analyst can label the graphic seal associated with a vehicle with the analyzed light indicator information of the vehicle. The labeling can be achieved via user input to an interactive user interface presenting an image or video clip obtained from the vehicle. In some embodiments, the graphic seal overlay component 338 overlays a specific graphic representation on the graphic seal associated with a vehicle with a discrepancy. For example, if there is a discrepancy in the vehicle in the captured image such that the machine learning model has identified a flashing brake light for a vehicle whose actual brake light was off, the graphic seal overlay component 338 can overlay the graphic seal on the vehicle with a specific graphic representation. The specific graphic representation can be any representation, such as adding any annotation on or near the graphic seal based on the color or shape of the graphic seal.
[0075] In some embodiments, the graphic representation of the seal is two-dimensional. In these embodiments, the graphic seal overlay component 338 can identify image pixels related to the vehicle. Then, a two-dimensional box can be generated and overlaid on the image of the vehicle. In some embodiments, the graphic seal overlay component 338 can perform image segmentation on the captured image to identify the vehicle and any lights. For example, the graphic seal overlay component 338 can segment the captured image into regions associated with the vehicle (e.g., a group of pixels) and regions not corresponding to the vehicle. In some embodiments, the graphic seal overlay component 338 generates a two-dimensional box on the region associated with the vehicle. In some embodiments, upon determining the region associated with the vehicle, the graphic seal overlay component 338 can generate a three-dimensional volume and overlay it on the region associated with the vehicle.
[0076] The light indicator display component 340 of FIG. 3B can be configured to display the identified light indicators of the vehicle within the captured image. An analyst can execute instructions on the light indicator display component 340 to obtain a set of continuously captured images (or video clip), where each image includes a graphic seal. In some embodiments, the light indicator display component 340 can be configured to request a machine learning model to identify or predict the light indicators of the vehicle included in the captured image. After the machine learning model identifies the light indicators associated with the vehicle in the captured image, the light indicator display component 340 is configured to obtain the identified light indicators and display them on the graphic seal associated with each vehicle in the captured image.
[0077] The machine learning model verification component 342 of FIG. 3B can be configured to determine one or more captured images that have at least one discrepancy between the identified optical indicators of the vehicles in the image and the actual optical indicators. Using the machine learning model verification component 342, an analyst can detect discrepancies by comparing the identified optical indicators with the actual optical indicators in the image and the driving routes of the vehicles. Some examples of potential discrepancies are shown in Table I. [Table 1]
[0078] The memory 330 further includes a machine learning model training component 344 configured to train the machine learning model. The training can be based on discrepancy correction. In some embodiments, one or more vehicles in the captured images associated with the detected discrepancies are labeled with the correct optical indicators. The correct optical indicators can be determined by an analyst. The analyst can label the correct optical indicator information on the graphic seals of the vehicles associated with the discrepancies. After receiving the labels (correct optical indicator data), the machine learning model training component 348 can correct the captured images with one or more discrepancies and store them in the network server 128. In some embodiments, the processing unit 322 instructs the machine learning model 124 to execute instructions to update, modify, or add one or more attributes related to optical indicator identification based on the corrected images.
[0079] FIG. 4 shows an example of a vehicle having cameras for capturing images of other vehicles. For ease of explanation, FIG. 4 may be described with reference to specific components of FIGS. 1, 2, 3A, and 3B. For illustrative purposes, FIG. 4 shows a top view including vehicle 110 and surrounding vehicles captured by cameras 220, 230, 240. Vehicle 110 may include a front camera 220, a pillar camera 230, and a repeater camera 240. The front camera 220 disposed in front of vehicle 110 is configured to capture a front image so that vehicle 410 can be captured in the image. The pillar cameras 230 disposed on both sides of vehicle 110 are configured to capture side images so that vehicle 412 can be captured in the image. The repeater camera 240 disposed behind vehicle 110 is configured to capture a rear image so that vehicle 414 can be captured in the image.
[0080] FIG. 5 shows an example of a captured image including identified light indicators associated with each vehicle in the image. The illustrated example may be included in an interactive user interface used, for example, by a user associated with labeling training data. The user can access an image or video clip obtained from a group of vehicles by executing the machine learning model described above. Thus, the image forms part of a video clip obtained by vehicles within the group of vehicles. Advantageously, to reduce the burden associated with labeling, the image may include a bounding box indicating an object (e.g., a vehicle) along with a light indicator (e.g., light indicator 506). The light indicator 506 may be assigned by the machine learning model, and the user can provide user input (e.g., touch-based input, mouse / keyboard, voice input) to change the light indicator 506 (e.g., to reflect a left turn, brake light, hazard light, etc. on the indicator).
[0081] For ease of explanation, FIG. 5 may be described with reference to certain components of FIGS. 1, 2, 3A, and 3B. The optical indicators can be identified by the machine learning model 124. As shown in FIG. 5, an image 500 including a vehicle 502 is captured by a front camera attached to the vehicle 110. The front camera of the vehicle 110 can capture a front view image of the vehicle, and the vehicle 110 is configured to transmit the captured image to the network server 128. In some embodiments, a user (e.g., an analyst) can use the verification computing device 130 to view or overlay a graphic seal on each vehicle included in the captured image.
[0082] For example, as shown in FIG. 5, the graphic seal 504 has a box shape and is overlaid on each vehicle 502. In one embodiment, the analyst uses the verification computing device 130 to overlay the graphic seal 504 on selected vehicles, such as a specific number of vehicles closer to the vehicle 110. The number of vehicles to be selected can be determined based on a specific application. In some embodiments, the graphic seal 504 may include one or more semantics related to the vehicle. In this example, as shown in FIG. 5, each graphic seal 504 includes an identified optical indicator 506 associated with the vehicle. In these embodiments, the optical indicator of each vehicle can be identified by the machine learning model 124. The machine learning model 124 identifies the optical indicator of each vehicle based at least on its learning parameters, algorithms, or attributes related to the identification of the optical indicator. The optical indicator can include, for example, a turn signal, a brake light, an emergency light, etc.
[0083] Figures 6A - 6B show an example of the verification of a machine learning model when identifying the light indicators of one or more surrounding vehicles. The verification can be obtained based on comparing the identified light indicators of the vehicles in the captured images using the machine learning model and the actual light indicators and driving routes of the vehicles. In some embodiments, one or more vehicles in the captured images are verified to have a discrepancy between the identified light indicators and the actual light indicators and driving routes of the vehicles. In these embodiments, the discrepancy can be corrected by receiving one or more inputs from an analyst having the authority to verify the machine learning model. For ease of explanation, Figures 6A - 6B can be described with reference to the specific components of Figures 1, 2, 3A, and 3B.
[0084] Figure 6A shows an example of the verification of the identified light indicators of the vehicles in the captured image 600. In some embodiments, the identified light indicators based on the machine learning model can be verified by analyzing a set of continuously captured images. For example, to verify the identified light indicators of the vehicles in the image, 300 images captured immediately after the image (played for 10 seconds in a 30fps video clip) can be analyzed to determine the actual light indicators and driving routes of the vehicles. The number of images (or the playback time of the video clip) can be determined based on the specific application.
[0085] Furthermore, in FIG. 6A, for illustrative purposes, two captured images are overlaid and used to determine the actual light indicators and driving routes of the vehicles included in the images. For example, the first captured image includes the first group of vehicles 604, 606, 608, and the second captured image includes the second group of vehicles 614, 616, 620. The first and second captured images are captured continuously, and the second captured image is captured immediately after the first captured image. In some embodiments, the identified light indicators 610, 620, 630 of the vehicles 604, 606, 608 are overlaid on top of the graphic seals 602 associated with each vehicle. In these embodiments, the light indicators of the vehicles 604, 606, 608 are identified using the machine learning model 124. The identified light indicators of the vehicles 604, 606, 608 can be compared with the captured images including the actual light indicators.
[0086] In some embodiments, the actual light indicator can be determined by analyzing the driving routes of vehicles 604, 606, 608. For example, vehicle 614 included in the second captured image indicates that vehicle 604, identified as "brake light on", has moved forward without decelerating. In this example, vehicle 604 can be verified to have a discrepancy between the identified signal indicator 610 and the vehicle's actual signal indicator and driving route 614. In another example, vehicle 616 included in the second captured image can verify that vehicle 606, identified as "light indicator off", has steered into the right lane 616 with the "right turn signal blinking". In this example, vehicle 606 is determined to have a discrepancy between the identified signal indicator 620 and the vehicle's actual signal indicator and driving route 616. Finally, in another example, vehicle 608, identified as "left turn signal blinking", has moved forward in the forward direction 618 without the "left turn signal blinking". In this example, vehicle 608 is determined to have a discrepancy between the identified signal indicator 630 and the vehicle's actual signal indicator and driving route 618. Various types of discrepancies can be detected within the system, and various examples of the discrepancies are listed in Table I above.
[0087] FIG. 6B shows an example of an interactive graphical user interface 650 for labeling graphic stamps of vehicles in a captured image. For example, an analyst can use the verification computing device 130 to access the interface 640 and correct the optical indicators of vehicles 604, 606, 608 having optical indicators 610, 620, 630 that are inappropriately identified by the machine learning model. The interactive graphical user interface can be provided as an application stored and executed by a processor of the verification computing device. In another example, the interactive graphical user interface may be provided by one or more processors implemented in the network server 128 so that a user can correct vehicle optical indicator detection using network server resources. In some embodiments, the interface 604 displays whether each of the vehicle lights is active or inactive. In some embodiments, an authorized user who labels or modifies the label associated with a vehicle light can modify the active or inactive status by utilizing a user input interface such as a mouse, keyboard, etc.
[0088] FIG. 7 shows an example of labeling graphic stamps with correct optical indicators of vehicles in a captured image 700. As shown in FIG. 7, the optical indicators of vehicles 604, 606, 608 (as shown in FIG. 6A) are corrected with optical indicators 710, 720, 730 respectively. For example, the optical indicators 610, 620, 630 (as shown in FIG. 6A) are corrected by overlaying the optical indicators 710, 720, 730. In some embodiments, the corrected image with overlaid optical indicator labeling is supplied to the machine learning model for training the model. In some embodiments, the trained machine learning model can be supplied to the autonomous driving system 210 (as shown in FIG. 2).
[0089] FIG. 8 is a flowchart showing an exemplary process that a verification computing device may perform to detect one or more vehicles in a captured image that has a discrepancy between an identified light indicator and an actual light indicator of a vehicle. Depending on the embodiment, the process shown in FIG. 8 may include fewer or additional blocks and / or the blocks may be executed in an order different from that shown. For ease of explanation, the process of FIG. 8 may be described with reference to the specific components of FIGS. 1, 2, 3B, 4, and 5.
[0090] Starting at block 810, the verification computing device can acquire a captured image of the surrounding view of the vehicle. The verification computing device can access the captured image by accessing a network server. In some embodiments, the verification computing device downloads the captured image. In other embodiments, the verification computing device can access the captured image using virtual computing service resources provided by a machine learning training system. In some embodiments, the vehicle can capture a front view image of the vehicle using a front camera attached to the front side of the vehicle, such as a windshield. The captured image can be in the form of a set of multiple video clips by merging a set of continuously captured images. For example, every 300 continuously captured images are merged into a video file that can be played back at 30 fps for about 10 seconds. The specifications of the video clip, including the number of captured images, frame rate, and resolution, can be determined based on a specific application. In some embodiments, the vehicle is configured to wirelessly connect to a network and transmit the captured image to a network server via the network. In some embodiments, the wireless standard for connecting to the network is based on, for example, other wireless communication technologies such as high-speed 4G LTE or 5G communication. Thus, in some embodiments, the network can include one or more wireless networks such as a Global System for Mobile Communications (GSM) network, a Code Division Multiple Access (CDMA) network, a Long-Term Evolution (LTE) network, or any other type of wireless network. The network can use protocols and components for communicating via either the Internet or any of the aforementioned types of networks.For example, protocols used by a network can include the Hypertext Transfer Protocol (HTTP), HTTP Secure (HTTPS), Message Queuing Telemetry Transport (MQTT), Constrained Application Protocol (CoAP), and the like. Protocols and components for communicating via either the Internet or any of the other aforementioned types of communication networks are well known to those of ordinary skill in the art and thus are not described in more detail herein.
[0091] Upon moving to block 820, the verification computing device can generate a graphic seal for each vehicle and overlay the graphic seal on each associated vehicle. The graphic seal can be box-shaped and can be overlaid on top of the vehicle within the captured image. In one embodiment, the verification computing device overlays the graphic seal on a selected vehicle within the captured image, such as a specific number of vehicles closer to the vehicle capturing the image. The number of vehicles to be selected can be determined based on a particular application. In some embodiments, the graphic seal represents one or more semantics associated with the vehicle.
[0092] In some embodiments, the graphic representation of the seal is two-dimensional. In these embodiments, the verification computing device can identify image pixels related to the vehicle. Next, a two-dimensional box (e.g., a bounding box) can be generated and overlaid on the image of the vehicle. In some embodiments, the verification computing device can perform image segmentation on the captured image to identify the vehicle and any lights. For example, the verification computing device can segment the captured image into regions associated with the vehicle (e.g., groups of pixels) and regions not corresponding to the vehicle. In some embodiments, the verification computing device generates a two-dimensional box (e.g., a bounding box) over the region associated with the vehicle. In some embodiments, when the region associated with the vehicle is determined, the graphic seal overlay component 338 can generate a three-dimensional volume and overlay it on the region associated with the vehicle.
[0093] Moving to block 830, in some embodiments, the machine learning model identifies the light indicators of the vehicle in the captured image. The identified light indicators can be supplied to the verification computing device.
[0094] Moving to block 840, the verification computing device can display the identified light indicators associated with each vehicle on the graphic seal of the vehicle in the captured image. For example, the graphic seal can represent whether the machine learning model has identified the light indicators of the vehicle. In some embodiments, the graphic seal can represent the type of light indicators identified by the machine learning model.
[0095] Upon moving to block 850, the verification computing device can detect a discrepancy between the identified optical indicator of the vehicle and the actual optical indicator of the vehicle. The discrepancy can be detected by analyzing a set of continuously captured images. For example, to verify the identified optical indicator of the vehicle in the image, 300 images captured immediately after the image (played for 10 seconds in a 30fps video clip) can be analyzed to determine the actual optical indicator and driving route of the vehicle. The number of images (or the playback time of the video clip) can be determined based on a specific application. In some embodiments, the discrepancy type is a false positive or a false negative. However, various types of discrepancies can be detected within the system, and various examples of discrepancies are described in Table I above.
[0096] Upon moving to block 860, if, in block 850, it is determined that one or more vehicles in the captured image have been detected as having a discrepancy between the identified optical indicator and the actual optical indicator, the verification computing device can flag the image and store it in the network server. The verification computing device can also store the flagged image in its internal or external storage medium. In some embodiments, the stored flagged images containing discrepancies are used to train a machine learning model.
[0097] If, in block 850, it is determined that there is no discrepancy, the verification computing device can end the process of detecting a discrepancy.
[0098] FIG. 9 is a flowchart illustrating an exemplary process that a verification computing device may perform to train a machine learning model. Depending on the embodiment, the process shown in FIG. 9 may include fewer or additional blocks and / or the blocks may be executed in an order different from that shown. For ease of explanation, the process of FIG. 9 may be described with reference to specific components of FIGS. 1, 2, 3A, 3B, and 6A.
[0099] Starting at block 910, the verification computing device can obtain flagged images by accessing a network server. Each of the flagged images or video clips that include the flagged images can include one or more vehicles that have a discrepancy between an identified light indicator and an actual light indicator. In some embodiments, the discrepancy type is a false positive or a false negative.
[0100] Moving to block 920, the verification computing device can request an analyst to correct the light indicators of the vehicles with discrepancies. The analyst can be an authorized user having the authority to verify the machine learning model, including a manager, developer, supervisor, administrator, etc.
[0101] Moving to block 930, after receiving the request, the analyst can correct the light indicators of the vehicles with discrepancies. Moving to block 940, in some embodiments, the analyst can label the graphic seals of the vehicles with discrepancies. The labeling can include the correct light indicators of the vehicles. In some embodiments, the labeled images with the correct light indicators are overlaid on the original images having the vehicles with discrepancies. The labeled images can be stored in the network server.
[0102] Upon moving to block 950, the verification computing device can send an image including the label of the correct light indicator of the vehicle to the machine learning model. In these embodiments, the machine learning model receives the labeled image with the correct light indicator and trains the machine learning module. For example, the parameters of the machine learning model can be updated (e.g., via gradient descent).
[0103] Upon moving to block 960, in some embodiments, the trained machine learning model can be provided to the vehicle's autonomous driving system. The vehicle can access the network and download the trained machine learning model via the network.
[0104] Figures 10A - 10D show examples of light indicator detection in various environments. The exemplary user interface is provided for illustrative purposes to show various functions of the system. As described above, the surrounding image of the vehicle is captured using a camera attached to the vehicle. The autonomous driving system can utilize the machine learning model to determine the light indicator by displaying it on the detected vehicle included in the captured image.
[0105] Figure 10A is an example of light indicator detection on a high - density road. In a high - density road environment, the vehicle can detect the nearest vehicle and determine the light indicator of the detected vehicle located closer to the vehicle. For example, vehicle 110 can determine the light indicator of the nearer vehicle 1002.
[0106] Figure 10B is an example of light indicator detection based on priority when determining the light indicator of a vehicle. In some embodiments, the analyst can prioritize determining the light indicator of the vehicle. For example, the prioritization of light indicator determination is based on the nearer vehicle 1012 (e.g., the highest priority), the flowing traffic vehicle 1014, the oncoming vehicle 1016, and the parked vehicle 1018. The prioritization described herein is merely an example and is not limited thereto.
[0107] FIG. 10C is an example of detecting a light indicator in a parking lot. As shown in the example, the vehicle can capture an image of the parking lot and determine the light indicator of vehicle 1020 in the parking lot.
[0108] FIG. 10D is an example of detecting light indicators for various types of moving objects. As shown in FIG. 10D, the analyst can determine the light indicators of various types of moving objects, including but not limited to motorcycle 1032, bus 1034, and any type of vehicle 1036. In some embodiments, the analyst can determine the light indicator by detecting the light at the edge of vehicle 1038.
[0109] FIG. 11 shows an exemplary interactive user interface 1100 that can be used by a user (e.g., an analyst). The interactive user interface 1100 presents an image from an image sensor or camera disposed around the vehicle. As described herein, the vehicle can provide an image or video clip to the system to update a machine learning model. Thus, in the illustrated example, the image reflects an image from these cameras or image sensors. Although an image is shown, as can be understood, the image can form a video clip and the user can play the video clip or a selected portion thereof.
[0110] The user interface 1100 includes a first image 1102 having a bounding box 1104 around an object (e.g., a truck). As shown, the object is included in multiple images from different cameras. In this example, a light indicator 1106 (e.g., a graphic signature of the light indicator), which is a graphical icon (e.g., a left-pointing hand representing a left turn signal), is disposed proximate to the bounding box 1104. There can be a number of graphical icons that provide an easy and concise way for the user to understand whether the light indicator 1106 is a left turn signal, right turn signal, hazard light, brake light, etc.
[0111] As described herein, the optical indicator 1106 can be determined by a machine learning model running on a vehicle. For example, referring to FIGS. 1-10D, label information can be provided along with an image or video clip that at least indicates a label associated with a vehicle light. As another example, a system (e.g., presenting a user interface or analyzing an image or video clip to perform training) can run a machine learning model to determine a label.
[0112] The optical indicator 1106 can be presented proximate to the bounding box 1104 during the presentation of the video clip. For example, the optical indicator 1106 can be presented with a similar offset from the bounding box 1104 such that it is attached to the bounding box 1104. Similarly, if a trailer is attached to an object, the optical indicator 1106 can be presented with a similar offset from the trailer.
[0113] A user of the user interface 1102 can provide user input to update the optical indicator 1106. For example, the user can select the indicator 1106 and be presented with a drop-down menu or other user interface (such as that shown in FIG. 6B) to update the indicator 1106. In this way, the user can generate ground truth (e.g., the updated indicator 1106) for use in training a machine learning model.
[0114] The user interface 1100 further includes a progress bar 1108 that enables selection of different portions of the video clip. For example, the progress bar 1108 may extend from the first timestamp to the last timestamp. In some embodiments, a portion of the progress bar 1108 may be a first color (e.g., green) that does not indicate an error or problem associated with the optical indicator. The progress bar 1108 may be a second color (e.g., red) that indicates that the optical indicator has been updated by the user.
[0115] Various embodiments of the present disclosure may be a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to execute aspects of the present disclosure.
[0116] For example, the functions described herein may be performed by one or more hardware processors and / or any other suitable computing device when software instructions are executed and / or in response to the execution of software instructions. The software instructions and / or other executable code can be read from a computer-readable storage medium (or media).
[0117] A computer-readable storage medium can be a tangible device that can hold and store data and / or instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electronic storage device (including any volatile and / or non-volatile electronic storage device), a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes, hereinafter, namely, a portable computer diskette, a hard disk, a solid state drive, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read only memory (CD-ROM), digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves in which instructions are recorded, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0118] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.
[0119] Computer-readable program instructions for performing the operations of the present disclosure (also referred to herein as, for example, "code", "instructions", "modules", "applications", "software applications", etc.) can be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, integrated circuit configuration data, and object-oriented programming languages such as Java, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be callable from other instructions or from themselves and / or can be called in response to detected events or interrupts. Computer-readable program instructions configured to execute on a computing device can be provided on a computer-readable storage medium and / or as a digital download (originally stored in a compressed or installable format that requires installation, decompression, or decryption before execution) that can then be stored on a computer-readable storage medium. Such computer-readable program instructions can be stored partially or fully on the memory device (e.g., computer-readable storage medium) of the computing device during execution for execution by the computing device. The computer-readable program instructions can execute entirely on the user's computer (e.g., the computing device in execution), partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing the electronic circuit using the state information of the computer-readable program instructions to implement aspects of the present disclosure.
[0120] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0121] These computer-readable program instructions may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium storing the instructions contains an article of manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0122] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions execute the functions / acts specified in one or more blocks of the flowchart and / or block diagram. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions and / or modules into its dynamic memory and send the instructions via a telephone line, cable line, or optical line using a modem. A modem local to the server computing system can receive data on the telephone / cable / optical line and place the data on a bus using a converter device that includes appropriate circuitry. The bus can carry the data to memory, from which the processor can retrieve and execute the instructions. The instructions received by the memory may optionally be stored in a storage device (e.g., solid state drive) either before or after execution by the computer processor.
[0123] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions described in the blocks may be performed in an order different from that shown in the figures. For example, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the related functions. Further, in some implementations, certain blocks may be omitted. The methods and processes described herein are also not limited to any particular sequence, and the related blocks or states can be executed in other sequences that are appropriate.
[0124] It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or a combination of dedicated hardware and computer instructions. For example, any of the processes, methods, algorithms, elements, blocks, applications, or other functions (or portions of functions) described in the foregoing sections can be embodied in electronic hardware such as a dedicated processor (e.g., an application-specific integrated circuit (ASIC)), a programmable processor (e.g., a field-programmable gate array (FPGA)), an application-specific circuit, etc. (any of which can also achieve the technology by combining custom hardwired logic, logic circuits, ASICs, FPGAs, etc. with custom programming / execution of software instructions), and / or may be fully or partially automated thereby.
[0125] Any of the above processors and / or any devices incorporating any of the above processors may be referred to herein, for example, as a "computer", "computer device", "computing device", "hardware computing device", "hardware processor", "processing unit", etc. The computing devices of the above embodiments can generally be controlled and / or coordinated (but not necessarily) by operating system software such as Mac OS, iOS, Android, Chrome OS, Windows OS (e.g., Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows Server, etc.), Windows CE, Unix, Linux, SunOS, Solaris, Blackberry OS, VxWorks, or other suitable operating systems. In other embodiments, the computing device may be controlled by its own operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file systems, networking, I / O services, and in particular provide user interface functions such as a graphical user interface ("GUI").
[0126] As described above, in various embodiments, certain functions may be accessible to a user via a web-based viewer (such as a web browser) or other suitable software program. In such an implementation, the user interface may be generated by a server computing system and transmitted to the user's web browser (e.g., to be executed on the user's computing system). Alternatively, data necessary to generate the user interface (e.g., user interface data) may be provided to the browser by the server computing system, where the user interface may be generated (e.g., the user interface data may be executed by a browser accessing a web service and may be configured to render the user interface based on the user interface data). The user can then interact with the user interface via the web browser. The user interface of certain implementations may be accessible via one or more dedicated software applications. In certain embodiments, one or more of the computing devices and / or systems of the present disclosure can include a mobile computing device, and the user interface may be accessible via such a mobile computing device (e.g., a smartphone and / or a tablet).
[0127] Numerous variations and modifications can be made to the above-described embodiments, and its elements should be understood to be among other acceptable examples. All such modifications and variations are intended to be included within the scope of the present disclosure herein. The above description details specific embodiments. However, it will be understood that, however detailed the above may seem in text, the systems and methods can be implemented in many ways. Also, as noted above, the use of specific terms when describing specific features or aspects of the systems and methods should not be construed to mean that the term is redefined herein so as to be limited to any particular properties of the features or aspects of the systems and methods with which the term is associated.
[0128] In particular, conditional language such as "can", "could", "might", or "may" generally, unless specifically stated otherwise or understood in a different sense within the context in which it is used, is intended to convey that a particular embodiment includes a particular feature, element, and / or step, but other embodiments may or may not. Thus, such conditional language is generally not intended to mean that a feature, element, and / or step is required in any way in one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are included in, or are to be performed in, any particular embodiment, regardless of user input or prompt.
[0129] Connective words such as the phrase "at least one of X, Y, and Z" or "at least one of X, Y, or Z" should be understood in the context in which they are generally used to convey that items, terms, etc. can be any one of X, Y, or Z, or a combination thereof, unless otherwise specified. For example, the term "or" is used, for example, when connecting a list of elements, in an inclusive sense (not an exclusive sense) such that the term "or" means one, some, or all of the elements in the list. Thus, such connective words are generally not intended to mean that a particular embodiment requires the presence of at least one of each of at least one of X, at least one of Y, and at least one of Z.
[0130] The term "a" as used herein should be given an inclusive interpretation rather than an exclusive interpretation. For example, unless otherwise specified, the term "a" should not be understood to mean "exactly one" or "one and only one", but rather, the term "a", whether used in the claims or elsewhere in the specification, and regardless of the use of quantifiers such as "at least one", "one or more", or "a plurality" in the claims or elsewhere in the specification, means "one or more" or "at least one".
[0131] The term "comprising" as used herein should be given an inclusive interpretation rather than an exclusive interpretation. For example, a general-purpose computer comprising one or more processors should not be construed as excluding other computer components, and may in particular include components such as memory, input / output devices, and / or network interfaces.
[0132] The foregoing detailed description has shown, described, and pointed out novel features as applicable to various embodiments. It will be understood, however, that various omissions, substitutions, and changes in the form and detail of the devices or processes illustrated may be made without departing from the spirit of the present disclosure. As will be recognized, certain embodiments of the invention described herein may be embodied in forms that do not provide all of the features and benefits described herein, as some features may be used or practiced separately from other features. The scope of the particular inventions disclosed herein is indicated by the appended claims rather than the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1. A system for labeling images to train a machine learning model to detect optical indicators on a vehicle, the system comprising: one or more processors; and a non-transitory computer storage medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: obtaining images of one or more vehicles on a roadway; identifying the position of each of the one or more vehicles; displaying a graphic seal on each of the one or more vehicles via a user interface to indicate that the vehicle has been detected by the system; receiving, via the user interface, an indication of whether an optical indicator is active or non-active in each of the one or more vehicles for labeling the images for the machine learning model; A system comprising.
2. The system of claim 1, wherein obtaining images includes obtaining images from a plurality of vehicles having an autonomous driving system.
3. The system of claim 2, wherein obtaining images includes obtaining images of the plurality of vehicles when the autonomous driving system determines that optical indicator detection has been improperly determined by the autonomous driving system.
4. The system of claim 1, wherein identifying the position of each of the one or more vehicles includes identifying the vehicle within the image and determining the graphic coordinates of the vehicle within the image.
5. The system of claim 1, wherein displaying the graphic seal on each of the one or more vehicles includes displaying a bounding box around each of the one or more vehicles in the acquired image.
6. The system of claim 1, wherein identifying the position of each of the one or more vehicles includes performing image segmentation on the acquired image, the image segmentation generating a region of each acquired image corresponding to the vehicle.
7. Receiving an indication of whether the light indicator is active or inactive includes receiving a mouse selection from a user who labels the vehicle as having an active or inactive light indicator, the system of claim 1.
8. Receiving the indication of whether the light indicator is active or inactive includes receiving an indication of whether the brake light is active or inactive, the system of claim 1.
9. Receiving the indication of whether the light indicator is active or inactive includes receiving an indication of whether the turn signal is active or inactive, the system of claim 1.
10. A system for labeling images to train a machine learning model to detect light indicators on a vehicle, the system including one or more processors and a non-transitory computer storage medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including obtaining an image of one or more vehicles on a roadway, identifying the position of each of the one or more vehicles in the obtained image, determining whether the light indicator has been indicated as active or non-active by the autonomous driving system of each of the one or more vehicles, determining one or more vehicles having an incorrect prediction of whether the light indicator was active or non-active from the image of the one or more vehicles, labeling, via a user interface, the image having an incorrect prediction with the correct indication of whether the light indicator is active or non-active, including, a system.
11. Identifying the position of each of the one or more vehicles includes identifying the vehicle in the image and determining the graphic coordinates of the vehicle in the image, the system of claim 10.
12. Obtaining an image includes obtaining images from a plurality of vehicles having autonomous driving systems, the system of claim 10.
13. The system according to claim 12, wherein obtaining an image includes obtaining an image from the plurality of vehicles when the autonomous driving system determines that the optical indicator detection has been inappropriately determined by the autonomous driving system.
14. The system according to claim 10, further comprising displaying a graphic seal on each of the one or more vehicles to indicate that the vehicle has been detected by the system via the user interface.
15. The system according to claim 14, wherein displaying the graphic seal on each of the one or more vehicles includes displaying a bounding box around each of the one or more vehicles in the acquired image.
16. The system according to claim 10, wherein the indication of whether the optical indicator of each of the one or more vehicles is active or inactive is predicted by the autonomous driving system of each vehicle.
17. The system according to claim 10, wherein the incorrect prediction is a mismatch between the optical indicator and the position of the vehicle.
18. The system according to claim 10, further comprising receiving an updated optical indicator that receives a mouse selection from a user who labels the vehicle with the optical indicator based on the position of the vehicle.
19. The system according to claim 10, wherein the indication of whether the optical indicator is active or inactive is an indication of whether the brake light is active or inactive.
20. The system according to claim 10, wherein the indication of whether the optical indicator is active or inactive is an indication of whether the turn signal is active or inactive.
21. A method for labeling an image for training a machine learning model to detect an optical indicator on a vehicle, comprising: obtaining an image of one or more vehicles on a lane; identifying the position of each of the one or more vehicles; labeling an indication of whether the optical indicator is active or inactive in each of the one or more vehicles via a user interface; determining one or more vehicles having an incorrect prediction from the image of the one or more vehicles; Receiving an updated indication of whether the optical indicator is active or non - active on the vehicle having the incorrect prediction; A method comprising. **Claim 22** The method according to claim 21, wherein the step of identifying the position of each of the one or more vehicles includes identifying the vehicle in the image and determining the graphic coordinates of the vehicle in the image. **Claim 23** The method according to claim 21, wherein the step of acquiring an image includes acquiring images from a plurality of vehicles having an autonomous driving system. **Claim 24** The method according to claim 21, wherein the step of acquiring an image includes acquiring images from the one or more vehicles when the optical indicator detection is inappropriately determined by the autonomous driving system of each vehicle. **Claim 25** The method according to claim 21, further comprising the step of displaying a graphic stamp on each of the one or more vehicles to indicate that the vehicle has been detected by the machine learning model. **Claim 26** The method according to claim 25, wherein the step of displaying the graphic stamp on each of the one or more vehicles includes displaying a bounding box around each of the one or more vehicles in the acquired image. **Claim 27** The method according to claim 21, wherein the indication of whether the optical indicator of each of the one or more vehicles is active or non - active is predicted by the autonomous driving system of each vehicle. **Claim 28** The method according to claim 21, wherein the incorrect prediction is a mismatch between the optical indicator and the position of the vehicle. **Claim 29** The method according to claim 21, wherein the step of receiving the updated optical indicator includes receiving a mouse selection from a user who labels the vehicle with the optical indicator based on the position of the vehicle.