Object classification and related applications based on frame and event camera processing
By combining frame and event camera processing, utilizing feature points and pixel-level motion information, and incorporating machine learning models for object classification, the problem of error detection and classification in conventional cameras is solved, thus improving the accuracy of object detection and classification.
Patent Information
- Application Number
- CN202180029975.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-17
- Filing Date
- 2021-10-07
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2041-10-07
AI Technical Summary
Conventional frame-based cameras suffer from misdetection or misclassification in object detection and classification, which is unacceptable, especially in critical decision-making applications.
By combining frame and event camera processing, frame and event camera signals are acquired through the image sensor circuit system. Object classification is performed using feature points and pixel-level motion information to generate first and second object classification results. Pattern matching and difference analysis are then performed using a machine learning model to determine the final object classification label.
It improves the accuracy of object detection and classification, and can robustly detect and classify objects that might be misdetected or misclassified by conventional methods, thereby enhancing the reliability of critical decisions.
Smart Images

Figure CN115668309B_ABST
Abstract
Description
[0001] Cross-reference to related applications / incorporation by reference
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 089413, filed October 8, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0003] Various embodiments of this disclosure relate to image and event camera processing. More specifically, various embodiments of this disclosure relate to systems and methods for object classification and related applications based on frame and event camera processing. Background Technology
[0004] Advances in image processing have led to the development of various image processing devices and technologies that can perform object detection and classification using captured images. In many cases, there may be a small number of objects in an image that conventional image processing devices (such as frame-based cameras) might misdetect or misclassify. Inaccuracies in object detection or classification can be unacceptable in several applications, especially those that rely on object detection or classification to make critical decisions or actions.
[0005] By comparing the described system with some aspects of this disclosure as set forth in the remainder of this application and with reference to the accompanying drawings, the limitations and disadvantages of conventional and traditional methods will become clear to those skilled in the art. Summary of the Invention
[0006] As set forth more fully in the claims, and substantially as illustrated in at least one figure and / or described in conjunction with at least one figure, a system and method for object classification and related applications based on frame and event camera processing are provided.
[0007] These and other features and advantages of this disclosure can be understood by reading the following detailed description of this disclosure and the accompanying drawings, in which the same reference numerals always refer to the same parts. Attached Figure Description
[0008] Figure 1 This is a block diagram illustrating an exemplary environment for object classification based on frame and event camera processing according to embodiments of the present disclosure.
[0009] Figure 2 According to embodiments of this disclosure Figure 1 A block diagram of an exemplary system.
[0010] Figure 3A and 3B Together, they depict exemplary operations of object classification based on frame and event camera processing according to embodiments of the present disclosure.
[0011] Figure 4A and 4B collectively depict a block diagram illustrating example operations for determining a behavior pattern of a driver of a vehicle according to embodiments of the present disclosure.
[0012] Figure 5 is a diagram illustrating an example scenario for controlling one or more vehicle functions based on selection of virtual user interface (UI) elements according to embodiments of the present disclosure.
[0013] Figure 6A , 6B and 6C collectively depict a diagram illustrating an example observation sequence for determining a behavior pattern of a driver of a vehicle according to embodiments of the present disclosure.
[0014] Figure 7A and 7B collectively depict a diagram illustrating an example scenario for generating notification information based on a behavior pattern of a driver according to embodiments of the present disclosure.
[0015] Figure 8 is a diagram illustrating an example scenario for generating notification information based on a behavior pattern of a driver according to embodiments of the present disclosure.
[0016] Figure 9 is a flowchart diagram illustrating an example method for object classification based on frame and event camera processing according to embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] The implementations described below can be found in the disclosed systems and methods for object classification based on frame and event camera processing and related applications. Example aspects of the present disclosure provide a system for object classification based on frame and event camera processing. The system can be implemented on, for example, an imaging device, a smartphone, an edge computing device, a vehicle, an Internet of Things (IOT) device, a passenger drone, and the like.
[0018] At any instant, the system can be configured to acquire the output of the image sensor circuitry and determine object classification labels based on frame and event camera processing. The system can be configured to determine first object classification results based on feature points associated with one or more first objects (such as objects that can be in a stationary state or can be about to reach a stationary state) in a first frame of the acquired output. Further, the system can determine second object classification results based on performing event camera signal processing operations on at least two frames of the acquired output to generate event frames. The generated event frames can include pixel-level motion information associated with one or more second objects (such as objects that can be in motion). Based on the determined first object classification results and the determined second object classification results, the system can be configured to determine one or more object classification labels corresponding to the one or more objects, which can be included in the at least first object(s) and second object(s). By performing frame and event camera processing, the disclosed system can be able to robustly detect and classify the object(s), some of which can be falsely detected and / or falsely classified by typical frame-based object detection / classification methods.
[0019] Figure 1 is a block diagram illustrating an exemplary environment for object classification based on frame and event camera processing, in accordance with an embodiment of the present disclosure. Referring to Figure 1 , a diagram of a network environment 100 is shown. The network environment 100 can include a system 102, image sensor circuitry 104, a server 106, and a database 108. The system 102 can be coupled to the image sensor circuitry 104 and the server 106 via a communication network 110. A user 112 that can be associated with the system 102 is also shown.
[0020] The system 102 can include suitable logic, circuitry, code, and / or interfaces that can be configured to acquire the output of the image sensor circuitry 104. The system 102 can determine object classification results based on performing frame and event camera signal processing on the acquired output. Example implementations of the system 102 can include, but are not limited to, one or more of an imaging device (such as a camera), a vehicle electronic control unit (ECU), a smartphone, a cellular phone, a mobile phone, a wearable or head-mounted electronic device (e.g., an extended reality (XR) headset), a gaming device, a mainframe, a server, a computer workstation, an edge computing device, a passenger drone support system, a consumer electronic (CE) device, or any computing device with image processing capabilities.
[0021] Image sensor circuitry 104 can include suitable logic, circuitry, and / or interfaces that can be configured to acquire image signals based on light signals that can be incident on image sensor circuitry 104. The acquired image signals can correspond to one or more frames of a scene in a field of view (FOV) of an imaging unit that includes image sensor circuitry 104. Image sensor circuitry 104 can include the acquired image signals in its output. In embodiments, the acquired image signals can be raw image data in a Bayer space. The imaging unit can be a camera device that is separate from or integrated into system 102. Examples of image sensor circuitry 104 can include, but are not limited to, a passive pixel sensor, an active pixel sensor, a semiconductor charge-coupled device (CCD) based image sensor, a complementary metal-oxide-semiconductor (CMOS) based image sensor, a back-illuminated CMOS sensor with global shutter, a silicon-on-insulator (SOI) based single-chip image sensor, an N-type metal-oxide-semiconductor based image sensor, a flat panel detector, or other variants of image sensors.
[0022] Server 106 can include suitable logic, circuitry, and interfaces and / or code that can be configured to detect and / or classify one or more objects in image(s) and / or event frame(s). Server 106 can include a database for pattern matching, such as database 108, and a trained artificial intelligence (Al), machine learning, or deep learning model for correctly detecting and / or classifying objects in frame(s). In an example embodiment, server 106 can be implemented as a cloud server and can perform operations through a web application, a cloud application, an HTTP request, a repository operation, a file transfer, or the like. Other example implementations of server 106 can include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, or a cloud computing server.
[0023] In at least one embodiment, server 106 can be implemented as a plurality of distributed cloud-based resources by using several techniques well known to those of ordinary skill in the art. Those of ordinary skill in the art will understand that the scope of the present disclosure can not be limited to implementing server 106 and system 102 as two separate entities. In certain embodiments, the functionality of server 106 can be incorporated, in whole or at least in part, into system 102 without departing from the scope of the present disclosure.
[0024] The database 108 can include suitable logic, interfaces, and / or code that can be configured to store object pattern information, which can include, for example, example images of objects for each object class, object features, class annotations or labels, and the like. The database 108 can be a relational or non-relational database. In embodiments, the database 108 can be stored on a server, such as a cloud server, or can be cached and stored on the system 102.
[0025] The communication network 110 can include a communication medium through which the system 102 can communicate with the server 106 and other devices omitted from the disclosure for brevity. The communication network 110 can include a wired connection, a wireless connection, or a combination thereof. Examples of the communication network 110 can include, but are not limited to, the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a satellite communication network such as using a global navigation satellite system (GNSS) satellite constellation, a cellular or mobile wireless network such as a 4th generation long term evolution (LTE) or a 5th generation new radio (NR), a personal area network (PAN), a local area network (LAN), or a metropolitan area network (MAN). The various devices in the network environment 100 can be configured to connect to the communication network 110 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols can include, but are not limited to, at least one of a transmission control protocol and internet protocol (TCP / IP), a user datagram protocol (UDP), a hypertext transfer protocol (HTTP), a file transfer protocol (FTP), Zig Bee, EDGE, IEEE 802.11, light fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, a wireless access point (AP), device-to-device communication, a cellular communication protocol, a satellite communication protocol, and a Bluetooth (BT) communication protocol.
[0026] In operation, the image sensor circuitry 104 can produce an output that can include one or more frames of a scene. In embodiments, the system 102 can control the image sensor circuitry 104 to acquire an image signal that can be represented as one or more frames of a scene.
[0027] The system 102 can acquire the output of the image sensor circuitry 104. The acquired output can include one or more frames. For example, the acquired output can include a first frame 114, which can be associated with a scene having one or more stationary objects, one or more moving objects, or a combination thereof.
[0028] The system 102 can detect feature points in the first frame 114 of the acquired output. Each such feature point can correspond to information associated with a particular structure, such as a point, edge, or a basic object or shape component in the first frame 114. Additionally or alternatively, such a feature point can be a feature vector or a result of a general neighborhood operation or feature detection applied to the first frame 114. Based on the feature points associated with one or more first objects in the first frame 114 of the acquired output, the system 102 can determine a first object classification result. In embodiments, the first object classification result can include a bounding box around such first objects in the first frame 114 and a class label predicted for such first objects. In these or other embodiments, such first objects can be in a stationary state in the scene or can be about to reach a stationary state in the scene.
[0029] In some scenarios, the first object classification result can not detect and classify all objects of interest that can be present in the scene and the first frame 114. For example, some objects can be falsely detected or can still not be detected. Moreover, even if some objects are detected in the first frame 114, one or more such objects can be falsely classified. Accordingly, event camera signal processing can be performed separately to detect objects in the scene that can be in motion and information that can be captured in the acquired output of the image sensor circuitry 104.
[0030] The system 102 can perform event camera signal processing operations on at least two frames of the acquired output to generate event frames 116. The generated event frames 116 can include pixel-level motion information associated with one or more second objects. In contrast to standard frame-based camera processing, the event camera signal processing operations can output events, each of which can represent a change in luminance. For each pixel location in a frame of the acquired output, the circuitry 202 can record a log intensity each time it can detect an event, and can continuously monitor for a change of sufficient magnitude compared to the recorded intensity. When the change exceeds a threshold, the circuitry 202 can detect a new event, which can be recorded by the location of the pixel, the time at which the new event was detected, and the polarity of the change in luminance, such as 1 to indicate an increase in luminance or 0 to indicate a decrease in luminance.
[0031] The pixel-level movement information 308A can correspond to changes in luminance or intensity recorded in the pixels of the event frame 116. Based on the pixel-level movement information, the system 102 can determine a second object classification result. In embodiments, the second object classification result can include a bounding box around such second objects in the event frame 116 and a class label predicted for such second objects. In some scenarios, the second object classification result can detect and classify object(s) that can have been mis-detected, mis-classified, or can not have been included in the first object classification result of the first frame 114 at all. Thus, the second object classification result can supplement the classification processing of all object(s) of interest in the scene as captured in the output of the acquisition of the image sensor circuitry 104. Based on the determined first object classification result and the determined second object classification result, the system 102 can determine one or more object classification labels. Such object classification labels can correspond to one or more objects that can be included in the one or more first objects and the one or more second objects at least.
[0032] Figure 2 is an example system in accordance with an embodiment of the present disclosure Figure 1 A block diagram of an example system in accordance with an embodiment of the present disclosure is shown. The elements of the system 102 are explained in conjunction with Figure 1 Figure 2 Figure 2 A block diagram 200 of the system 102 is shown with reference to FIG. 2. The system 102 can include circuitry 202, a memory 204, input / output (I / O) devices 206, and a network interface 208. In embodiments, the system 102 can also include the image sensor circuitry 104. The circuitry 202 can be communicatively coupled to the memory 204, the I / O devices 206, and the network interface 208. The I / O devices 206 can include, for example, a display device 210.
[0033] The circuitry 202 can include suitable logic, circuitry, and / or interfaces that can be configured to execute program instructions associated with different operations to be performed by the system 102. The circuitry 202 can include one or more specialized processing units, which can be implemented as an integrated processor or a cluster of processors that collectively perform the functions of the one or more specialized processing units. The circuitry 202 can be implemented based on a variety of processor technologies known in the art. Examples of implementations of the circuitry 202 can be x86-based processors, graphics processing units (GPUs), reduced instruction set computing (RISC) processors, application-specific integrated circuit (ASIC) processors, complex instruction set computing (CISC) processors, microcontrollers, central processing units (CPUs), and / or other computing circuitry.
[0034] Memory 204 can include suitable logic, circuitry, and / or interfaces that can be configured to store program instructions to be executed by circuitry 202. In at least one embodiment, memory 204 can be configured to store first frame 114, event frame 116, first object classification results, second object classification results, and one or more object classification labels. Memory 204 can be configured to store object classes. Example implementations of memory 204 can include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), a hard disk drive (HDD), a solid-state drive (SSD), a CPU cache, and / or a secure digital (SD) card.
[0035] I / O devices 206 can include suitable logic, circuitry, interfaces, and / or code that can be configured to receive input and provide output based on the received input. I / O devices 206 can include various input and output devices that can be configured to communicate with circuitry 202. For example, system 102 can receive user input to select user interface elements and initiate frame and event camera processing via I / O devices 206. Examples of I / O devices 206 can include, but are not limited to, a touchscreen, a keyboard, a mouse, a joystick, a display device (e.g., display device 210), a microphone, or a speaker.
[0036] Display device 210 can include suitable logic, circuitry, and / or interfaces that can be configured to display user interface elements. In one embodiment, display device 210 can be a touch device that can be configured to receive user input from user 112 via display device 210. Display device 210 can include a display unit that can be implemented by several known technologies, such as but not limited to, a liquid crystal display (LCD) display, a light-emitting diode (LED) display, a plasma display, or an organic LED (OLED) display technology, or at least one of other display devices. According to embodiments, the display unit of display device 210 can refer to a display screen of a head-mounted device (HMD), a smart eyewear device, a see-through display, a projection-based display, an electrochromic display, or a transparent display.
[0037] The network interface 208 can include suitable logic, circuitry, interfaces and / or code that can be configured to facilitate the circuitry 202 in communicating with the image sensor circuitry 104 and / or other communication devices via the communication network 110. The network interface 208 can be implemented by using various known technologies to support wireless communication of the system 102 via the communication network 110. The network interface 208 can include, for example, an antenna, a radio-frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, local buffer circuitry, and / or the like.
[0038] The network interface 208 can be configured to communicate via wireless communication with a network such as the Internet, an intranet, a wireless local area network (LAN) or metropolitan area network (MAN), a cellular telephone network, and / or the like. The wireless communication can be configured to use one or more of a variety of communications standards, protocols, and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), code division multiple access (CDMA), time division multiple access (TDMA), satellite communications such as using a GNSS satellite constellation, Bluetooth, wireless fidelity (Wi-Fi) such as IEEE 802.11a, IEEE 802.1 lb, IEEE 802.1 lg, or IEEE 802.1 ln, voice over Internet Protocol (VoIP), ZigBee, and / or the like.
[0039] As described in Figure 1 the functions or operations performed by the system 102 can be performed by the circuitry 202. The operations performed by the circuitry 202 are described in detail in, for example Figure 3A and 3B , Figure 4A and Figure 4B , Figure 5 , Figure 6A , 6B and 6C, Figure 7A and 7B , Figure 8 and Figure 9 .
[0040] Figure 3A and 3B collectively depict a block diagram illustrating exemplary operations for object classification based on frame and event camera processing in accordance with embodiments of the present disclosure. The Figure 1 and Figure 2 are explained in conjunction with elements in Figure 3A and 3B . Reference is made to Figure 3A and 3B, a block diagram 300 is shown that illustrates example operations from 302 to 336 as described herein. The example operations shown in block diagram 300 can begin at 302 and can be performed by any computing system, device, or apparatus, such as by the system 102 of Figure 1 or Figure 2 Although illustrated with discrete blocks, depending on the implementation of the example operations, the example operations associated with one or more blocks of the block diagram 300 can be divided into additional blocks, combined into fewer blocks, or eliminated.
[0041] At 302, data acquisition can be accomplished. In an embodiment, the circuitry 202 can acquire an output of the image sensor circuitry 104. The image sensor circuitry 104 can be integrated into the system 102 or can be externally connected to the system 102 via a network interface or I / O interface. In an embodiment, the output can be acquired from a data source, which can be, for example, a permanent storage on the system 102, a cloud server, etc. In another embodiment, the output can be acquired directly from a frame buffer of the image sensor circuitry 104 (in its uncompressed or raw form). The acquired output can include one or more frames of a scene in the FOV of the image sensor circuitry 104. Each frame of the acquired output can include at least one object.
[0042] At 304, frame camera signal processing operations can be performed. In an embodiment, the circuitry 202 can be configured to perform frame camera signal processing operations on the acquired output to obtain a first frame 302A. As shown, for example, the first frame 302A depicts a scene of a busy road and includes a pedestrian and a stationary vehicle. In an embodiment, operations such as noise reduction and edge detection can be performed on the first frame 302A. Depending on the image sensor circuitry 104, these operations can be performed before an analog-to-digital (ADC) of the image sensor circuitry 104.
[0043] As part of the frame camera signal processing operations, the circuitry 202 can detect feature points 304A in the first frame 302A of the acquired output. The detected feature points 304A can be referred to as unique features associated with one or more first objects in the first frame 302A. In an embodiment, the feature points 304A can include information that can be needed to classify each such first object in a particular object class. For example, a first object can have feature points such as, but not limited to, headlight, window, and wheel, and a second object can have feature points such as, but not limited to, leg, arm, and face.
[0044] The circuitry 202 can determine a first pattern match between the detected feature points 304A and a first reference feature 304B associated with a known object class. The first reference feature 304B can include feature points or template features, which can be extracted from a training dataset of object images (e.g., which can belong to a known object class and can be stored in the database 108). The first pattern match can be determined based on an implementation of a pattern matching method. The pattern matching method can include a set of computational techniques by which the detected feature points 304A can be matched with the first reference feature 304B. Based on the matching, a group of the detected feature points 304A can be probabilistically classified as a particular object class. The pattern matching method can be, for example, a template matching method, a machine learning or deep learning based pattern matching method, a Haar cascade, or any feature matching method that can be based on features such as Haar features, Local Binary Pattern (LBP), or Scale-Invariant Feature Transform (SIFT) features. Detailed implementations of these example methods can be known to those skilled in the art, and therefore detailed descriptions of these methods are omitted in the present disclosure for brevity.
[0045] At 306, a first object classification result can be determined. In an embodiment, the circuitry 202 can be configured to determine the first object classification result based on the feature points 304A associated with the one or more first objects in the first frame 302A. Specifically, the first object classification result can be determined based on the first pattern match (determined at 304). The first object classification result can include the first frame 302A overlaid with first bounding box information associated with the one or more first objects. As shown, for example, the bounding box information can include a bounding box 306A and a bounding box 306B, which can be predicted to include the one or more first objects. In an embodiment, the first object classification result can further include one or more class labels (e.g., car, person, or bicycle) corresponding to the one or more first objects in the first frame 302A.
[0046] In some scenarios, the first object classification result can not detect and classify all the objects of interest that can be present in the scene and in the first frame 302A. For example, some objects can be falsely detected or can still not be detected. Further, even if some objects are detected in the first frame 302A, there can be a chance that one or more such objects are falsely classified. Therefore, event camera signal processing can be performed separately to help detect and classify all the objects of interest in the scene, information of which can be captured in the output of the acquisition by the image sensor circuitry 104.
[0047] At 308, an event camera signal processing operation can be performed. In an embodiment, the circuitry 202 can be configured to perform an event camera signal processing operation on at least two frames (or scans) of the acquired output to generate an event frame 302B. The generated event frame 302B can include pixel-level movement information 308A associated with one or more second objects. In an embodiment, there can be at least one possibly common object between the one or more second objects and the one or more first objects. In another embodiment, there can be no common object between the one or more first objects and the one or more second objects. In another embodiment, the one or more first objects can be the same as the one or more second objects.
[0048] The event camera signal processing operation can include measuring the change in luminance or intensity (asynchronously and independently) of each pixel from at least two scans or frames of the acquired output. The initial output of the event camera signal processing operation can include a sequence of digital events or spikes, where each event represents a change in luminance (e.g., in log intensity measurement). Changes in luminance or intensity that can be below a preset threshold can not be included in the sequence of digital events or spikes.
[0049] In comparison to standard frame-based camera processing, the event camera signal processing operation can output events, each of which can represent a change in luminance. For each pixel location in a frame of the acquired output, the circuitry 202 can record the log intensity each time it can detect an event, and can continuously monitor for changes of sufficient magnitude compared to that recorded value. When the change exceeds a threshold, the circuitry 202 can detect a new event, which can be recorded by the location of the pixel, the time at which the new event was detected, and the polarity of the change in luminance, such as 1 to represent an increase in luminance or 0 to represent a decrease in luminance. The pixel-level movement information 308A can correspond to the changes in luminance or intensity recorded in the pixels of the event frame 302B. For example, a moving object can cause a change in intensity of the light signal incident on a particular pixel location of the image sensor circuitry 104. This change can be recorded in the event frame 302B in a pixel corresponding to the particular pixel location of the image sensor circuitry 104.
[0050] At 310, a second object classification result can be determined. In an embodiment, the circuitry 202 can be configured to determine the second object classification result based on the pixel-level movement information 308A. The second object classification result can include the event frame 302B overlaid with second bounding box information associated with one or more second objects. The bounding box information can include one or more bounding boxes, such as the bounding box 310A and the bounding box 310B, which can be predicted to include one or more second objects. The first object classification result can also include one or more class labels (e.g., car, person, or bicycle) corresponding to the one or more second objects in the event frame 302B.
[0051] In an embodiment, the circuitry 202 can determine the second object classification result by pattern matching the pixel-level movement information 308A with second reference features 308B associated with a known object class. Similar to the first reference features 304B, the second reference features 308B can include feature points or template features, which can be extracted from a training dataset of object images (e.g., which can belong to a known object class and can be stored in the database 108). In an embodiment, the second reference features 308B can include pixel information associated with an object or object component of the known object class. The pattern matching can be based on a method, examples of which are mentioned in the foregoing description (at 304). Detailed implementations of these example methods are known to those skilled in the art, and therefore a detailed description of these methods is omitted in the present disclosure for the sake of brevity.
[0052] At 312, the first and second classification results can be compared. In an embodiment, the circuitry 202 can be configured to compare the determined first object classification result with the determined second object classification result. Based on the comparison, it can be determined whether there is a discrepancy between the first object classification result and the second object classification result. In the event that there is no discrepancy between the first object classification result and the second object classification result, control can pass to 314. Otherwise, control can pass to 316.
[0053] At 314, object classification labels can be determined for the one or more objects. The circuitry 202 can be configured to determine one or more object classification labels based on the determined first object classification results and the determined second object classification results. The determined one or more object classification labels can correspond to one or more objects that can be included in the one or more first objects and the one or more second objects. Each such object classification label can specify an object class that the object can belong to. For example, the first object classification results can include the first frame 302A with bounding boxes (i.e., bounding box 306A and bounding box 306B) overlaid on the car and the pedestrian (i.e., first objects), and the second object classification results can include the event frame 302B with bounding boxes (i.e., bounding box 310A and bounding box 310B) overlaid on the car and the pedestrian in the event frame 302B. For the car and the pedestrian, object classification labels can be determined as “car” and “person,” respectively.
[0054] At 316, a difference between the determined first object classification results and the determined second object classification results can be determined. In an embodiment, the circuitry 202 can be configured to determine the difference between the determined first object classification results and the determined second object classification results. For example, scenario a can include a moving vehicle and a pedestrian. The determined first object classification results can include only the bounding box information for the pedestrian. In this case, the frame-based camera processing can have missed detecting the moving vehicle. On the other hand, the determined second object classification results can include the bounding box information for the moving vehicle and the pedestrian. Based on the determined first object classification results and the determined second object classification results, a difference at the pixel location where the moving vehicle is present can be determined.
[0055] At 318, a first object that caused the first object classification results to be different from the second object classification results can be determined. In an embodiment, the circuitry 202 can be configured to determine the first object (e.g., the moving vehicle) that can be misclassified, not classified, detected but not classified, misdetected, or misdetected and misclassified. The first object can be determined based on the determined difference and from among the one or more first objects or the one or more second objects. For example, if an object is detected in both the first and the second object classification results, but the class of the first object in the first object classification results is different from the class of the object in the second object classification results, then the object can be included in the determined difference. Similarly, if the second object detection result detects an object that is not detected in the first object detection result, then the object can also be included in the determined difference.
[0056] At 320, a pattern matching operation can be performed. In an embodiment, the circuitry 202 can be configured to perform a pattern matching operation on the determined first object. Such an operation can be performed based on the determined difference, the first object classification result, and the second object classification result. In another embodiment, the circuitry 202 can transmit information such as the determined difference, the first object classification result, the second object classification result, the first frame 302A, and / or the event frame 302B to an edge node. An edge node such as an edge computing device can be configured to receive the information and perform a pattern matching operation based on the received information.
[0057] In an embodiment, the pattern matching operation can be a machine learning based pattern matching operation or a deep learning based pattern matching operation. The pattern matching operation can rely on a machine learning (ML) model or a deep learning model trained on a pattern matching task to detect and / or classify the object(s) in the first frame 302A and / or the event frame 302B. Such a model can be defined by its hyperparameters and topology / architecture. For example, a neural network based model can have the number of nodes (or neurons), the activation function(s), the number of weights, the cost function, the regularization function, the input size, the learning rate, the number of layers, etc. as its hyperparameters. Such a model can be referred to as a computational network or a system of nodes (e.g., artificial neurons). For deep learning implementations, the nodes of the deep learning model can be arranged in layers as defined in a neural network topology. These layers can include an input layer, one or more hidden layers, and an output layer. Each layer can include one or more nodes (or artificial neurons, e.g., represented by a circle). The output of all nodes in the input layer can be coupled to the input of at least one node of the hidden layer(s). Similarly, the input of each hidden layer can be coupled to the output of at least one node in other layers of the model. The output of each hidden layer can be coupled to the input of at least one node in other layers of the deep learning model. The node(s) in the final layer can receive input from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer can be determined from the hyperparameters, which can be set before, during, or after training the deep learning model on a training dataset.
[0058] Each node of the deep learning model can correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters that are tunable during training of the model. The set of parameters can include, for example, a weight parameter, a regularization parameter, etc. Each node can use the mathematical function to compute an output based on one or more inputs from nodes in other layer(s) (e.g., previous layer(s)) of the deep learning model. All or a portion of the nodes of the deep learning model can correspond to the same mathematical function, which can be the same or different.
[0059] In training of a deep learning model, one or more parameters of each node can be updated based on whether the output of the final layer for a given input (from a training dataset) matches a correct result based on a loss function of the deep learning model. The above process can be repeated for the same or different inputs until a minimum of the loss function is reached and training error is minimized. Several methods for training are known in the art, e.g., gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, metaheuristics, etc.
[0060] In embodiments, the ML model or deep learning model can include electronic data, which can be implemented as a software component of an application executable on the system 102, for example. The ML model or deep learning model can rely on libraries, external scripts, or other logic / instructions for execution by a processing device such as the system 102. The ML model or deep learning model can include computer executable code or routines to enable a computing device such as the system 102 to perform one or more operations, such as detecting and classifying objects in input images or event frames. Additionally or alternatively, the ML model or deep learning model can be implemented using hardware, including but not limited to a processor, microprocessor (e.g., to perform or control performance of one or more operations), a field programmable gate array (FPGA), or an application specific integrated circuit system (ASIC). For example, an inference accelerator chip can be included in the system 102 to accelerate computations of the ML model or deep learning model. In some embodiments, the ML model or deep learning model can be implemented using a combination of both hardware and software.
[0061] Examples of deep learning models can include, but are not limited to, artificial neural networks (ANN), convolutional neural networks (CNN), regions with CNNs (R-CNN), fast R-CNN, faster R-CNN, You Only Look Once (YOLO) networks, residual neural networks (Res-Net), feature pyramid networks (FPN), Retina-Net, Single Shot Detector (SSD), and / or combinations thereof.
[0062] At 322, a first object class can be determined. In embodiments, the circuitry 202 can be configured to classify the determined first object into a first object class based on performance of the pattern matching operation. The first object class can specify a class to which the first object can belong.
[0063] At 324, it can be determined whether the first object class is correct. In embodiments, the circuitry 202 can be configured to determine whether the first object class is correct. For example, the circuitry 202 can match the first object class to the object class initially identified in the first object classification result or the second object classification result. If the two object classes match, then it can be determined that the first object class is correct and control can pass to 326. Otherwise, control can pass to 328.
[0064] At 326, a first object classification label for the first object can be determined. In embodiments, the circuitry 202 can be configured to determine the first object classification label from the one or more object classification labels based on determining that the first object class is correct.
[0065] At 328, the event frame 302B or the first frame 302A can be transmitted to a server, such as the server 106. In embodiments, the circuitry 202 can transmit the event frame 302B or the first frame 302A to the server based on determining that the first object class is not correct.
[0066] At 330, a second object class for the first object can be received from the server. In embodiments, the circuitry 202 can receive the determined second object class for the first object from the server 106 based on the transmission. The second object class can be different from the determined first object class. The server can implement a pattern matching method with or without the support of a database of exhaustive object images and pattern matching information, such as template features.
[0067] At 332, it can be determined whether the second object class is correct. In embodiments, the circuitry 202 can be configured to determine whether the second object class is correct. The circuitry 202 can match the second object class to the object class initially identified in the first object classification result or the second object classification result. If the two object classes match, then it can be determined that the second object class is correct and control can pass to 334. Otherwise, control can pass to 336.
[0068] At 334, an object classification label for the first object can be determined. In embodiments, the circuitry 202 can determine the first object classification label for the first object based on determining that the second object class is correct. The first object classification label can be included in the one or more object classification labels determined at 314.
[0069] At 336, the circuitry 202 can determine an object classification label for the first object as unknown. The first object can be labeled with the object classification label. The label can be reported to a server, such as the server 106, to improve a database, such as the database 108, for pattern matching. Control can pass to end.
[0070] Figure 4A and 4B collectively depict a block diagram illustrating example operations for determining a behavior pattern of a driver of a vehicle according to embodiments of the present disclosure. The elements in Figure 1 , 2 , 3A, and 3B are explained in conjunction with Figure 4A and 4B . Referring to Figure 4A and 4B , a block diagram 400 is shown that illustrates example operations 402-422 as described herein. The example operations shown in block diagram 400 can begin at 402 and can be performed by any computing system, device, passenger drone, or apparatus, such as by system 102 of Figure 1 or Figure 2 . While illustrated with discrete blocks, depending on the implementation of the example operations, the example operations associated with one or more blocks of block diagram 400 can be divided into additional blocks, combined into fewer blocks, or eliminated.
[0071] At 402, data can be acquired. In embodiments, circuitry 202 can acquire the output of image sensor circuitry 104 (as described, for example, at 302 of Figure 3A ). Each of the first frames 114 and event frames 116 can capture scene information of an environment outside of a vehicle. Image sensor circuitry 104 can be mounted inside a vehicle (such as a dashcam), mounted outside of a vehicle body, or can be part of a roadside unit (such as a traffic camera or any device that can include a camera and / or can be part of a vehicle-to-everything (V2X) network). The vehicle can be an autonomous, semi-autonomous, or non-autonomous vehicle. Examples of vehicles can include, but are not limited to, electric vehicles, hybrid vehicles, and / or vehicles that use a combination of one or more different renewable or non-renewable power sources. The present disclosure can also apply to vehicles such as two-wheeled vehicles, three-wheeled vehicles, four-wheeled vehicles, or vehicles having more than four wheels. Descriptions of such vehicles have been omitted from the disclosure for brevity.
[0072] The scene information of the environment outside of the vehicle can include objects that are typically present on a road or in a driving environment. For example, the scene information of the environment outside of the vehicle can include information about a road and a road type, a lane, a traffic signal or sign, a pedestrian, roadside infrastructure, a nearby vehicle, etc.
[0073] At 404, frame camera signal processing operations can be performed. In embodiments, circuitry 202 can be configured to perform frame camera signal processing operations on the acquired first frames 114 of the output to detect feature points 304A in the acquired first frames 114 of the output (e.g., at 306 ofFigure 3A at 304.
[0074] At 406, a first object classification result can be determined. In an embodiment, the circuitry 202 can be configured to determine the first object classification result based on the feature points 304A associated with the one or more first objects in the acquired output’s first frame 114 (e.g., as described at 306 of FIG. 3). Figure 3A
[0075] At 408, an event camera signal processing operation can be performed. In an embodiment, the circuitry 202 can be configured to perform the event camera signal processing operation on at least two frames of the acquired output to generate event frames 116 (e.g., as described at 308 of FIG. 3). Figure 3A
[0076] At 410, a second object classification result can be determined. In an embodiment, the circuitry 202 can be configured to determine the second object classification result based on the pixel-level motion information included in the generated event frames 116 (e.g., as described at 310 of FIG. 3). Figure 3A
[0077] At 412, an object classification label can be determined. In an embodiment, the circuitry 202 can be configured to determine one or more object classification labels based on the determined first object classification result and the determined second object classification result (such as described at operations 312 to 336 of FIG. 3). Figure 3A and Figure 3B
[0078] At 414, an electronic control system of the vehicle can be controlled to detect an event including, but not limited to, an accident or a hazard in the environment. In an embodiment, the circuitry 202 can be configured to control the electronic control system of the vehicle to detect an event including, but not limited to, an accident or a hazard in the environment. As an example and not a limitation, the electronic control system can be one of an advanced driver assistance system (ADAS) or an autonomous driving (AD) system. Detailed implementations of ADAS or AD systems can be known to those skilled in the art, and therefore, detailed descriptions of such systems are omitted in the present disclosure for the sake of brevity.
[0079] At 416, a behavior pattern of the driver of the vehicle can be determined. In embodiments, the circuitry 202 can be configured to determine a behavior pattern of the driver of the vehicle based on the determined one or more object classification labels and the detected event. A behavior pattern can refer to a set of actions, observations, or judgments that can be related to the driver or the vehicle, and can be recorded around the time of the detected event, such as from some period before (or just before) the event is detected to some period after the event is detected. In embodiments, machine learning can be used on the system 102 to determine the behavior pattern of the driver. In addition, a scenario graph can also be built on the system 102, depicting relationships between various data points recorded around the time period in which the event is detected. Figure 6A and Figure 6B An example of a scenario graph for a risky event is provided in FIGS. 1-3. As an example and not by way of limitation, a behavior pattern can include a driving pattern, a measure of the driver’s attention to the environment outside the vehicle, activities performed by the driver while driving, etc. The driving pattern can be based on data such as the average speed of the vehicle before, at, or after the event is detected, the one or more object classification labels. For example, when a pedestrian (an object detected on the road) is about to cross the street, the driver can be driving at a low speed, while the driver can be driving fast when the road is empty. The driver can also be depicted in terms of driving skills that can be determined from the behavior pattern.
[0080] At 418, one or more operations of the vehicle can be controlled. In embodiments, the circuitry 202 can be configured to control one or more operations of the vehicle based on the determined behavior pattern. For example, based on the behavior pattern, it can be determined that the driver is not paying attention or not slowing down when approaching a traffic stop sign, then the circuitry 202 (e.g., as part of a vehicle ECU) can apply the brakes to reduce the speed of the vehicle. For example, details of such control of one or more operations are further described in FIGS. 4-6. Figure 6A and 6B
[0081] At 420, notification information can be generated. In embodiments, the circuitry 202 can generate notification information based on the determined behavior pattern. The notification information can include, for example, a message that can alert the driver about a nearby traffic-related event (such as an accident, a risky situation, or a turn), a suggestion or instruction to the driver to modify behavior or perform an operation of the vehicle (such as stopping near a traffic stop sign). For example, examples of notification information are provided in FIGS. 7-9. Figure 7A and 7B
[0082] At 422, the generated notification information can be transmitted to an electronic device 422A. In an embodiment, the circuitry 202 can transmit the generated notification information to the electronic device 422A, which can be associated with the vehicle or the driver of the vehicle. For example, the electronic device 422A can be a smartphone, any personal device (that can be carried by the driver) or can be a display screen of the vehicle. The electronic device 422A can receive and present the notification information.
[0083] Figure 5 is a diagram illustrating an exemplary scenario for controlling one or more vehicle functions based on selection of virtual user interface (UI) elements, in accordance with an embodiment of the present disclosure. The elements in Figure 1 , 2 , 3A, 3B, 4A, and 4B are explained in conjunction with Figure 5 . Referring to Figure 5 , an exemplary scenario 500 depicting a vehicle 502 is shown. Also shown are an image projection device 504, virtual UI elements 506A, 506B, and 506C, and a display device 508 (such as the display device 210). In Figure 5 , a user 510 (e.g., a driver) seated on a driver’s seat inside the vehicle 502 is further shown.
[0084] The image projection device 504 can include suitable logic, circuitry, and interfaces that can be configured to project the virtual UI elements 506A, 506B, and 506C onto any space (not in use or less in use) that can be accessible to the user 510 while driving the vehicle 502. The image projection device 504 (such as a projector) can be coupled with the in-vehicle system (i.e., the system 102) of the vehicle 502 to display the virtual UI elements 506A, 506B, and 506C.
[0085] In an embodiment, the circuitry 202 can control the image projection device 504 to project the image of the UI onto a projection surface inside the vehicle 502. The projection surface can be any surface inside the vehicle 502 that can be accessible to the user 510 and can be reliably used to display the virtual UI elements 506A, 506B, and 506C. As shown, for example, the projection surface is a portion of the steering wheel of the vehicle 502. It is to be noted that the projection surface and the virtual UI elements 506A, 506B, and 506C in Figure 5 are provided by way of example only. The present disclosure can be applicable to any projection surface inside the vehicle 502 or any other machine and virtual UI elements (in any style, number, or appearance) without deviating from the scope of the present disclosure.
[0086] The circuitry 202 can be configured to obtain the output of the image sensor circuitry 104. Based on the obtained output, the Figure 3Aand Figure 3B operations to generate the first frame 114 and the event frame 116 and determine one or more object classification labels of one or more objects that can be captured in the first frame 114 and the event frame 116. Each of the first frame 114 and the event frame 116 can capture scene information of an in-vehicle environment inside the vehicle 502. The in-vehicle environment can include the user 510.
[0087] At any instant, the circuitry 202 can detect a finger movement on one of the virtual UI elements 506A, 506B, and 506C of the UI included in the projected image. Such detection can be done based on a second object classification label of the determined one or more object classification labels. For example, the second object classification label can specify that the finger is above the UI, and the event frame 116 can indicate movement of the finger around the area in which the event frame 116 detects the finger.
[0088] In an embodiment, as part of the event camera signal processing operations, the circuitry 202 can detect edges around one or more objects, such as the finger, and can determine the presence of a steering wheel sink region based on the event frame 116 and the detected edges. The presence of the sink region can indicate a pressing action of the finger on one of the virtual UI elements 506A, 506B, and 506C. The finger movement can be detected based on the determination of the edge detection and the determination of the sink region.
[0089] From the virtual UI elements 506A, 506B, and 506C, the circuitry 202 selects a virtual UI element based on the detection of the finger movement. In an embodiment, such selection can be based on the detection of the edges and the determination that the sink region coincides with the area on the steering wheel where the selected virtual UI element is present. The circuitry 202 can control one or more vehicle functions based on the selection. For example, the virtual UI element 506A can turn on / off a Bluetooth function, the virtual UI element 506B can open a map application on an electronic device associated with the user 510 or the display device 508, and the virtual UI element 506C can open a phone application. The UI elements can also be used to control vehicle functions such as wiper control, headlight control, or turn indicator control.
[0090] Figure 6A 、 6B and 6C collectively depict a diagram illustrating an exemplary observation sequence to be used to determine a behavior pattern of a driver of a vehicle in accordance with an embodiment of the disclosure. The elements in Figure 1 、 2 , 3A, 3B, 4A, 4B, and 5 are explained in conjunction with Figure 6A 、 6B and 6C. Reference is made to Figure 6A 、 6Band 6C, showing a graph 600 illustrating a sequence of observations after detecting an event as described herein. Such observations can be collected by the system 102 and can be analyzed to determine a behavior pattern of the driver (e.g., as described in Figure 4A and 4B .
[0091] At any moment, the circuitry 202 or ADAS / AD can detect an event, such as an accident 602B or a hazard 602A in the environment. When there is a hazard 602A in the environment, the driver can take an action 604A or the ADAS / AD 604B can take an action according to the vehicle 502 (and vehicle features). In the case of an accident, the driver can apply the brakes 606A, adjust the steering wheel 606B, or can just honk the horn 606C. In the case of a moving object approaching 608A, a stationary object approaching 608B, or a stationary object approaching more 608C, the driver can apply the brakes 606A. The object can be a vehicle 610A, a pedestrian 610B, a motorcycle 610C, a bicycle 610D, a bus 610E, or an animal 610F. The object’s movement can be towards the left side 612A, the right side 612B, the front 612C, or the back 612D of the vehicle 502. Based on the strength of the object’s movement, the driver can press or release the accelerator 614 in order to merge onto a highway 616A, drive to a highway location 616B, or pass a crossroad 616C. The circuitry 202 can determine the type of traffic, whether it is a jam 618B, a slow traffic 618C, or no traffic 618D. The time of day can be recorded on the system 102 when an event is detected, such as evening 620A, morning 620B, night 620C, and day 620D. If it is evening, it can be determined whether there is sun glare 622A, or it is cloudy 622B, raining 622C, or foggy 622D. If there is sun glare in the evening, it can be determined whether the driver is on a hands-free call 624A, listening to music 626B or the radio 626C, or listening to non-music 626D. At 628, the sequence of observations from 602A to 626D can be recorded on the system 102. The circuitry 202 analyzes the recorded sequence of observations to determine a behavior pattern of the driver and helps the system 102 control one or more operations of the vehicle based on the behavior pattern.
[0092] Figure 7A and 7B collectively depict a graph illustrating an exemplary scenario for generating notification information based on a behavior pattern of a driver according to embodiments of the present disclosure. The elements in Figure 1 , 2 , 3A, 3B, 4A, 4B, 5, 6A, 6B, and 6C are explained in Figure 7A and 7B . Reference is made to Figure 7Aand 7B FIG. 700 is shown illustrating example operations 702-730 as described herein. The example operations shown in FIG. 700 can begin at 702 and can be performed by any computing system, device, or apparatus, such as by the system 102 of Figure 1 or Figure 2 .
[0093] At 702, a database can be retrieved. In an embodiment, the circuitry 202 can retrieve a database associated with a driver of a vehicle. The database can be retrieved from a server, such as the server 106, or a local storage device on the vehicle. Thereafter, the circuitry 202 can analyze information that can be included in the retrieved database and can be associated with an environment in the vehicle. Further, the circuitry 202 can analyze scene information of an environment outside the vehicle. In an embodiment, such analysis can also be made based on a classification of objects included in the in-vehicle environment and the environment outside the vehicle.
[0094] At 704, it can be determined that the driver is using a phone with a hands-free function while driving the vehicle. In an embodiment, the circuitry 202 can determine that the driver is using a phone while driving based on the analysis of the scene information and the information associated with the environment in the vehicle.
[0095] The circuitry 202 can record a sequence of observations based on the analysis of the scene information and the information associated with the environment in the vehicle. For example, the sequence of observations can include an environmental condition, such as sun glare 706, a time of day, such as evening 708, an action of merging onto a highway location (at 712) immediately after a jam (at 710), an action of stepping on a pedal (at 714). Based on such observations and the determination that the driver can be using a phone, the circuitry 202 can control an electronic device, such as a vehicle display or a user device, to display a notification 716, which can include a message 718, such as “Some risk when merging on the left, slow down.”
[0096] In an embodiment, the circuitry 202 can also be configured to transmit the notification 716 to the electronic device. The notification 716 can include, but is not limited to, an audible signal or a visual signal or any other notification signal. In an example, the notification 716 can be a voice signal that includes an audible message, such as “Hold the steering wheel.” In another example, the notification 716 can be a visual signal, such as a marking on a map and / or on a potential risk area on a see-through camera.
[0097] At 720, it can be determined whether the driver followed the displayed notification 716. In embodiments, the circuitry 202 can be configured to determine whether the driver followed the displayed notification 716. For example, the circuitry 202 can match the action taken by the driver with the action suggested in the displayed notification 716. If the two actions match, then control can pass to 722. Otherwise, control can pass to 730.
[0098] At 722, it can be determined whether the detected event (e.g., in Figure 6A embodiments, the circuitry 202 can determine a behavior classification label for the driver based on the detected event, such as the accident or hazard in the environment and the action taken by the driver in response to the detected event. The determined behavior classification label can correspond to a behavior pattern, which can include a set of actions, observations, or judgments related to the driver or the vehicle, and can be recorded before and after the time of the detected event.
[0099] In the case where the driver did not follow the displayed notification 716 and a detected event, such as an accident or hazard in the environment, the circuitry 202 can determine a behavior classification label for the driver as “dangerous” 724. A behavior classification label such as dangerous 724 can imply that the driver did not follow the displayed notification 716 and therefore took a risk while driving. For such a classification label, the circuitry 202 can enrich the warning that can be issued to the driver. For example, the warning can include specific penalties, fees, and / or rewards related to at least one of an insurance claim, a rental car cost, a driver or pilot license renewal system, a truck and ride sharing driver or pilot rating system including passengers, etc. The warning can also include specific instructions to avoid such an event in the future.
[0100] In another instance where the driver followed the displayed notification 716 and a detected event, such as an accident or hazard in the environment, the circuitry 202 can determine a behavior classification label for the behavior pattern of the driver as “prevented by warning” 728.
[0101] At 730, it can be determined whether the detected event (e.g., in Figure 6AWhether it is a hazard or an accident. In an embodiment, the circuit system 202 can detect the event as a hazard or accident in the environment. If the driver does not follow the displayed notification 716 and an event (such as an accident or hazard in the environment) is detected, the circuit system 202 can determine the behavior classification label of the driver's behavior pattern as "warning required" 732. In another case where the driver does not follow the displayed notification 716 and no event (such as an accident or hazard in the environment) is detected, the circuit system 202 can determine the behavior classification label of the driver's behavior pattern as "no label" 734.
[0102] Figure 8 This is a diagram illustrating an exemplary scenario for generating notification information based on driver behavior patterns according to embodiments of the present disclosure. (In conjunction with...) Figure 1 , 2 Explain using elements from 3A, 3B, 4A, 4B, 5, 6A, 6B, 6C, 7A, and 7B. Figure 8 . refer to Figure 8 Figure 800 illustrates exemplary operations from 802 to 812 as described herein. The exemplary operations shown in Figure 800 may begin at 802 and can be performed by any computing system, device, or apparatus, such as by [unclear text - likely a computer or device]. Figure 1 or Figure 2 The system 102 is executed.
[0103] At point 802, a notification may be displayed. In an embodiment, the circuitry 202 may be configured to display the notification based on analysis of scene information and information associated with the environment within the vehicle. The notification may include message 802A, such as “There is a hazard while merging to the left, slow down.”
[0104] In one embodiment, the circuit system 202 can evaluate the driver's behavioral patterns based on the determined behavioral classification labels (also...). Figure 4A , 4B (As described in 6A, 6B, and 6C). Assessment may include identifying driver behavior patterns, such as rewards and / or punishments.
[0105] At 804, it can be determined whether the driver has followed the displayed notification. In an embodiment, circuitry 202 can be configured to determine whether the driver has followed the displayed notification. For example, circuitry 202 can match the actions taken by the driver with the actions suggested in the displayed notification. If the two actions match, then control can be passed to 806. Otherwise, control can be passed to 812.
[0106] At position 806, it can be determined (for example, at...). Figure 6AThe system 202 determines whether the detected event constitutes a hazard or an accident. If the driver follows the displayed instructions and detects an event (such as an accident or hazard in the environment), the circuit system 202 can assess the effectiveness of supporting the driver's behavioral pattern 808 (such as the reasons for this behavior). Such an assessment can help insurance services provide more accurate and reliable judgments for insurance claims. In another scenario, if the driver follows the displayed instructions and no event (such as an accident or hazard in the environment) is detected, the circuit system 202 can assess whether a reward 810 (such as a discount or reward points) should be added to the driver's behavior.
[0107] At 812, it can be determined (for example, at...) Figure 6A The circuit system 202 can determine whether the detected event is a hazard or an accident. In an embodiment, the circuit system 202 can detect the event as a hazard or accident in the environment. If the driver fails to follow the displayed instructions and an event such as an accident or hazard in the environment is detected, the circuit system 202 can assess that a penalty 814 (such as a fine or penalty) should be added to the insurance claim or vehicle costs. In another case where the driver fails to follow the displayed instructions and no event (such as an accident or hazard in the environment) is detected, the circuit system 202 can assess that no action is required 816.
[0108] Figure 9 This is a flowchart illustrating an exemplary method for object classification based on frame and event camera processing according to embodiments of the present disclosure. (In conjunction with...) Figure 1 , 2 Explain using elements from 3A, 3B, 4A, 4B, 5, 6A, 6B, 6C, 7, and 8. Figure 9 . refer to Figure 9 The flowchart 900 is shown. The operations of flowchart 900 can be performed by a computing system such as system 102 or circuit system 202. The operations can begin at 902 and proceed to 904.
[0109] At 904, the output of the image sensor circuitry 104 can be acquired. In one or more embodiments, the circuitry 202 can be configured to acquire the output of the image sensor circuitry. For example, Figure 3A and 3B The details of the acquisition are described in the document.
[0110] At position 906, a first object classification result can be determined. In one or more embodiments, the circuit system 202 can be configured to determine the first object classification result based on feature points associated with one or more first objects in the first frame of the acquired output. For example, in Figure 3A and 3B The details of determining the classification result of the first object are described in the document.
[0111] At 908, event camera signal processing operations can be performed. In one or more embodiments, circuitry 202 can be configured to perform event camera signal processing operations on at least two frames of the acquired output to generate event frames. The generated event frames may include pixel-level motion information associated with one or more second objects. For example, in Figure 3A and 3B The document describes the execution details of the event camera signal processing operations.
[0112] At point 910, the second object classification result can be determined. In one or more embodiments, the circuit system 202 can be configured based on, for example, in Figure 3A and 3B The pixel-level movement information described in the text is used to determine the classification result of the second object.
[0113] At 912, one or more object classification labels can be determined. In one or more embodiments, circuit system 202 can be configured to determine one or more object classification labels based on a determined first object classification result and a determined second object classification result. The determined one or more object classification labels can correspond to one or more objects that are at least included in one or more first objects and one or more second objects. For example, in Figure 3A and 3B The document describes the details of determining the category labels for one or more objects. Control can be passed to the end.
[0114] Although flowchart 900 is shown as discrete operations, such as 904, 906, 908, 910, and 912, this disclosure is not limited thereto. Therefore, in some embodiments, depending on the particular implementation, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated without departing from the essence of the disclosed embodiments.
[0115] Various embodiments of the present disclosure can provide a non-transitory computer-readable medium and / or storage medium having stored thereon computer-executable instructions that are executable by a machine and / or computer (e.g., system 102). The computer-executable instructions can cause the machine and / or computer (e.g., system 102) to perform operations that can further include determining a first object classification result based on feature points associated with one or more first objects in a first frame of the acquired output. The operations can further include performing event camera signal processing operations on at least two frames of the acquired output to generate event frames. The generated event frames can include pixel-level movement information associated with one or more second objects. The operations can further include determining a second object classification result based on the pixel-level movement information. The operations can further include determining one or more object classification labels based on the determined first object classification result and the determined second object classification result. The determined one or more object classification labels can correspond to one or more objects included in at least the one or more first objects and the one or more second objects.
[0116] Exemplary aspects of the present disclosure can include a system (such as system 102) that can include circuitry (such as circuitry 202). The circuitry 202 can be configured to acquire output of image sensor circuitry. The circuitry 202 can be configured to determine a first object classification result based on feature points associated with one or more first objects in a first frame of the acquired output. The circuitry 202 can be configured to perform event camera signal processing operations on at least two frames of the acquired output to generate event frames. The generated event frames can include pixel-level movement information associated with one or more second objects. The circuitry 202 can be configured to determine a second object classification result based on the pixel-level movement information. The circuitry 202 can be configured to determine one or more object classification labels based on the determined first object classification result and the determined second object classification result. The determined one or more object classification labels can correspond to one or more objects included in at least the one or more first objects and the one or more second objects.
[0117] According to an embodiment, the one or more first objects can be in a stationary state or about to reach a stationary state, and the one or more second objects are in motion.
[0118] According to an embodiment, the circuitry 202 can be further configured to detect feature points in the first frame of the acquired output. The circuitry 202 can be further configured to determine a first pattern match between the detected feature points and first reference features associated with a known object class. The first object classification result can be determined based on the determined first pattern match.
[0119] According to an embodiment, the circuitry 202 can be further configured to determine the second object classification result based on pattern matching the pixel level movement information and the second reference features associated with the known object classes.
[0120] According to an embodiment, the first object classification result can comprise a first frame overlaid with first bounding box information associated with the one or more first objects, and the second object classification result can comprise an event frame overlaid with second bounding box information associated with the one or more second objects.
[0121] According to an embodiment, the circuitry 202 can be further configured to determine a difference between the determined first object classification result and the determined second object classification result. Thereafter, the circuitry 202 can be further configured to determine, from the one or more first objects or the one or more second objects, a first object that is misclassified, unclassified, detected but unclassified, misdetected, or misdetected and misclassified, wherein the first object is determined based on the difference.
[0122] According to an embodiment, the circuitry 202 can be further configured to perform a pattern matching operation for the determined first object based on the determined difference, the determined first object classification result, and the determined second object classification result. The circuitry 202 can classify the determined first object into a first object class based on the performance of the pattern matching operation.
[0123] According to an embodiment, the circuitry 202 can be further configured to determine, for the first object, a first object classification label of the one or more object classification labels based on determining that the first object class is correct.
[0124] According to an embodiment, the pattern matching operation can be a machine learning based pattern matching operation or a deep learning based pattern matching operation.
[0125] According to an embodiment, the circuitry 202 can be further configured to transmit, to a server, the event frame or the first frame based on determining that the first object class is incorrect. The circuitry 202 can be further configured to receive, from the server, a second object class of the determined first object based on the transmission. The circuitry 202 can be further configured to determine, for the first object, a first object classification label of the one or more object classification labels based on determining that the second object class is correct.
[0126] According to an embodiment, each of the first frame and the event frame captures scene information of an environment external to the vehicle.
[0127] According to an embodiment, the circuitry 202 can be further configured to control an electronic control system of the vehicle to detect an event including an accident or a dangerous situation in the environment. The circuitry 202 can be further configured to determine a behavior pattern of a driver of the vehicle based on the determined one or more object classification labels and the detected event.
[0128] According to an embodiment, the electronic control system can be one of an advanced driver assistance system (ADAS) and an autonomous driving (AD) system.
[0129] According to an embodiment, the circuitry 202 can be further configured to control one or more operations of the vehicle based on the determined behavior pattern.
[0130] According to an embodiment, the circuitry 202 can be further configured to generate notification information based on the determined behavior pattern. The circuitry 202 can be further configured to transmit the generated notification information to an electronic device associated with the vehicle or the driver of the vehicle.
[0131] According to an embodiment, the circuitry 202 can be further configured to control an image projection device inside the vehicle to project an image of a user interface (UI) onto a projection surface inside the vehicle. Each of the first frame and the event frame can capture scene information of an in-vehicle environment inside the vehicle. The circuitry 202 can be further configured to detect a finger movement on a virtual UI element of the UI included in the projected image based on a second object classification label of the determined one or more object classification labels. The circuitry 202 can be further configured to select the virtual UI element based on the detection and control one or more vehicle functions based on the selection.
[0132] The present disclosure can be realized in hardware, or a combination of hardware and software. The present disclosure can be realized in a centralized fashion in one computer system or in a distributed fashion where different elements are spread across several interconnected computer systems. A computer system suitable for the execution of a computer program can be used. A combination of hardware and software can be a general-purpose computer system with a computer program that, when being loaded and executed, can control the computer system such that it carries out the methods described herein. The present disclosure can be realized in a hardware portion of an integrated circuit that also realizes other functions.
[0133] The present disclosure can also be embedded in a computer program product, which comprises all the features enabling the implementation of the methods described herein, and which - when loaded in a computer system - is able to carry out these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system having information processing capability to perform a particular function either directly or after either a) conversion to another language, code or notation; b) reproduction in a different material form.
[0134] While the present disclosure has been described with reference to certain implementations, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt a particular situation or material to the teachings of the present disclosure without departing from the central inventive concept hereof. Therefore, it is intended that the present disclosure not be limited to the particular implementation disclosed, but that the present disclosure will include all implementations falling within the scope of the appended claims.
Claims
1. A system comprising: The circuit system is configured as follows: Acquire the output of the image sensor circuit system; The classification result of the first object is determined based on the feature points associated with one or more first objects in the first frame of the acquired output; Perform event camera signal processing operations on at least two of the acquired output frames to generate event frames. The generated event frames include pixel-level motion information associated with one or more second objects; The classification result of the second object is determined based on pixel-level motion information; Based on the determined first object classification result and the determined second object classification result, one or more object classification labels are determined. The determined object classification labels correspond to one or more objects that are at least included in the one or more first objects and the one or more second objects; Determine the difference between the determined first object classification result and the determined second object classification result; From the one or more first objects or the one or more second objects, determine the first object that is misclassified, unclassified, detected but not classified, misdetected, or misdetected and misclassified, wherein the first object is determined based on the difference; Based on the determined differences, the determined first object classification result, and the determined second object classification result, perform a pattern matching operation for the determined first object; as well as Based on the execution of pattern matching operations, the identified first object is classified into the first object category.
2. The system of claim 1, wherein the one or more first objects are in a stationary state or about to reach a stationary state, and the one or more second objects are in motion.
3. The system of claim 1, wherein the circuit system is further configured to: Detect feature points in the first frame of the acquired output; and Determine the first pattern match between the detected feature points and a first reference feature associated with a known object category. The classification result of the first object is determined based on the first pattern matching.
4. The system of claim 1, wherein the circuitry is further configured to determine a second object classification result based on pattern matching of pixel-level motion information and a second reference feature associated with a known object category.
5. The system according to claim 1, wherein The first object classification result includes a first frame overlaid with first bounding box information associated with the one or more first objects, and The second object classification result includes an event frame overlaid with second bounding box information associated with the one or more second objects.
6. The system of claim 1, wherein the circuit system is further configured to determine a first object classification label among the one or more object classification labels for the first object based on the determination that the first object category is correct.
7. The system according to claim 1, wherein the pattern matching operation is a machine learning-based pattern matching operation or a deep learning-based pattern matching operation.
8. The system of claim 1, wherein the circuit system is further configured to: Based on the determination that the first object category is incorrect, the event frame or the first frame is transmitted to the server; and Based on the transmission, the server receives a second object category for the first object determined by the transmission; and Based on the determination that the second object category is correct, a first object category label is determined for the first object from among the one or more object category labels.
9. The system of claim 1, wherein each of the first frame and the event frame captures scene information of the vehicle's external environment.
10. The system of claim 9, wherein the circuit system is further configured to: The vehicle's electronic control system detects events, including accidents or hazards in the environment; and The driver's behavior patterns are determined based on one or more object classification labels and detected events.
11. The system of claim 10, wherein the electronic control system is one of an Advanced Driver Assistance System (ADAS) and an Autonomous Driving System (AD).
12. The system of claim 10, wherein the circuitry is further configured to control one or more operations of the vehicle based on a determined behavior pattern.
13. The system of claim 10, wherein the circuit system is further configured to: Generate notification information based on the identified behavioral patterns; and The generated notification information is transmitted to electronic devices associated with the vehicle or its driver.
14. The system of claim 1, wherein the circuit system is further configured to: The image projection device is controlled inside the vehicle to project images of the user interface (UI) onto projection surfaces inside the vehicle. Each of the first frame and the event frame captures scene information about the environment inside the vehicle. Based on the second object classification label from one or more determined object classification labels, detect finger movement on virtual UI elements of the UI included in the projected image. Select virtual UI elements based on the aforementioned detection; as well as Based on the selection, control one or more vehicle functions.
15. A method comprising: Acquire the output of the image sensor circuit system; The classification result of the first object is determined based on the feature points associated with one or more first objects in the first frame of the acquired output; Perform event camera signal processing operations on at least two of the acquired output frames to generate event frames. The generated event frames include pixel-level motion information associated with one or more second objects; The classification result of the second object is determined based on pixel-level motion information; Based on the determined first object classification result and the determined second object classification result, one or more object classification labels are determined. The determined object classification labels correspond to one or more objects that are at least included in the one or more first objects and the one or more second objects; Determine the difference between the determined first object classification result and the determined second object classification result; From the one or more first objects or the one or more second objects, determine the first object that is misclassified, unclassified, detected but not classified, misdetected, or misdetected and misclassified, wherein the first object is determined based on the difference; Based on the determined differences, the determined first object classification result, and the determined second object classification result, perform a pattern matching operation for the determined first object; as well as Based on the execution of pattern matching operations, the identified first object is classified into the first object category.
16. The method of claim 15, further comprising: Detect feature points in the first frame of the acquired output; as well as Determine the first pattern match between the detected feature points and a first reference feature associated with a known object category. The classification result of the first object is determined based on the first pattern matching.
17. The method of claim 15, further comprising determining a second object classification result based on pattern matching of pixel-level motion information and a second reference feature associated with a known object category.
18. A non-transitory computer-readable medium having computer-executable instructions stored thereon, the computer-executable instructions causing the system to perform operations when executed by the system, the operations including: Acquire the output of the image sensor circuit system; The classification result of the first object is determined based on the feature points associated with one or more first objects in the first frame of the acquired output; Perform event camera signal processing operations on at least two of the acquired output frames to generate event frames. The generated event frames include pixel-level motion information associated with one or more second objects; The classification result of the second object is determined based on pixel-level motion information; Based on the determined first object classification result and the determined second object classification result, one or more object classification labels are determined. The determined object classification labels correspond to one or more objects that are at least included in the one or more first objects and the one or more second objects; Determine the difference between the determined first object classification result and the determined second object classification result; From the one or more first objects or the one or more second objects, determine the first object that is misclassified, unclassified, detected but not classified, misdetected, or misdetected and misclassified, wherein the first object is determined based on the difference; Based on the determined differences, the determined first object classification result, and the determined second object classification result, perform a pattern matching operation for the determined first object; as well as Based on the execution of pattern matching operations, the identified first object is classified into the first object category.
Citation Information
Patent Citations
Object classification in video images
US20070154100A1
Object detection method and system
US20190065885A1