Tracker-based security solution for camera systems
By performing various consistency checks on the camera system of the autonomous driving system, visual attacks are identified and mitigated, thus solving the security problem of image processing in the autonomous driving system and improving the system's security and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-09-09
- Publication Date
- 2026-04-10
AI Technical Summary
The camera system of an autonomous driving system is vulnerable to malicious visual attacks, which can lead to image processing confusion or misleading, affecting safe operation.
Visual attacks are identified by performing temporal consistency checks, inconsistency count checks, and historical consistency checks on multiple images, and mitigation actions such as removing malicious tracks, outputting attack indications, or disabling malicious features are taken.
Effectively identify and mitigate visual attacks, improve the operational safety of autonomous and semi-autonomous devices, and reduce the risk of being misled.
Smart Images

Figure CN121844541A_ABST
Abstract
Description
Related applications
[0001] This application claims the benefit of priority to U.S. nonprovisional application No. 18 / 470,924, filed September 20, 2023, the entire contents of which are incorporated herein by reference. Background Technology
[0002] With the emergence of autonomous and semi-autonomous vehicles, robotic vehicles, and other types of mobile devices using Advanced Driver Assistance Systems (ADAS) and Autonomous Driving Systems (ADS), devices with such systems are becoming vulnerable to new forms of malicious behavior and threats; namely, spoofing or otherwise attacking camera systems at the heart of autonomous vehicle navigation and object avoidance. While such attacks may be rare now, they are expected to become a significant problem in the future as devices with autonomous driving systems expand. Summary of the Invention
[0003] Various aspects include methods that can be implemented on the processing system of the device, and systems for implementing methods for identifying and reacting to inconsistencies in images that may be caused by malicious attacks. These aspects may include: receiving multiple images from one or more cameras of the device; performing multiple different processes on the multiple images to detect different types of image inconsistencies; using the results of the multiple different processes on the multiple images to identify visual attacks; and performing one or more mitigation actions in response to the identification of a visual attack.
[0004] In some aspects, performing multiple different processes on multiple images to detect different types of image inconsistencies may include performing a temporal consistency check on multiple images spanning a time period. In other aspects, performing multiple different processes on multiple images to detect different types of image inconsistencies may include performing an inconsistency counter check on multiple images, which determines whether the number of inconsistencies in the images meets a threshold.
[0005] In some aspects, performing multiple different processes on multiple images to detect different types of image inconsistencies may include performing a past history check on the multiple images, which compares previously identified objects in previously processed images with identified objects in the currently acquired image to identify changes in at least one of the following: objects, object locations, or object classifications. In some aspects, the results of using multiple different processes on multiple images may include one or more of the following: identifying a visual attack if any one of the different types of image inconsistencies is detected; identifying a visual attack if the number of detected different types of image inconsistencies exceeds a threshold; identifying a visual attack if a majority of detectors detect image inconsistencies; or identifying a visual attack if a weighted majority of detectors detect image inconsistencies, wherein the weights applied to each of the different detectors are predetermined.
[0006] In some aspects, mitigation actions performed in response to the identification of a visual attack may include one or more of the following: removing malicious traces from a tracking database, outputting an indication of the attack, or disabling malicious features associated with objects identified in an image. In some aspects, mitigation actions performed in response to the identification of a visual attack may include reporting the detected attack to a remote system.
[0007] Another aspect includes devices such as vehicles, which include memory and a processor configured to perform operations of any of the methods outlined above. Another aspect may include devices such as vehicles having various components for performing functions corresponding to any of the methods outlined above. Another aspect may include a non-transitory processor-readable storage medium having processor-executable instructions stored thereon, the processor-executable instructions being configured to cause one or more processors of the device processing system to perform various operations corresponding to any of the methods outlined above. Attached Figure Description
[0008] The accompanying drawings, incorporated herein and forming part of this specification, illustrate exemplary embodiments of the claims and, together with the general description given above and the detailed description given below, serve to explain the features of the claims.
[0009] Figures 1A to 1C This is a component block diagram illustrating a typical system of autonomous devices in the form of vehicles suitable for implementing various schemes.
[0010] Figure 2 It is a functional block diagram showing the functional elements or modules of an autonomous driving system suitable for implementing various implementation schemes.
[0011] Figure 3It is a component block diagram applicable to processing systems that implement various implementation schemes.
[0012] Figure 4 This is a block diagram illustrating various operations performed on multiple images as part of an autonomous driving system, as well as the processing block diagrams involved in implementing the operations in various implementation schemes.
[0013] Figure 5 It is a block diagram of operations performed according to various implementation schemes as part of the verification of visual attacks.
[0014] Figure 6 This is a block diagram illustrating the operations and data structures involved in execution time consistency checks based on some implementation schemes.
[0015] Figure 7 This is a block diagram illustrating the operations and data structures involved in the consistency counter check according to some implementation schemes.
[0016] Figure 8 This is a block diagram illustrating the operations and data structures involved in the execution of past historical checks based on some implementation schemes.
[0017] Figures 9 to 11 This is an illustration of an alternative decision-making algorithm based on some implementation schemes for identifying potential attacks on the device's camera based on multiple detection methods.
[0018] Figure 12 This is a flowchart illustrating an example method for processing multiple images to identify visual attacks or potential visual attacks, based on some implementation schemes.
[0019] Figure 13 This is a flowchart of example operations performed according to some implementation schemes as part of identifying visual attacks or potential visual attacks in multiple images. Detailed Implementation
[0020] Various embodiments will be described in detail with reference to the accompanying drawings. Where possible, the same reference numerals will be used throughout the drawings to refer to the same or similar parts. References to specific examples and embodiments are for illustrative purposes and are not intended to limit the scope of the claims.
[0021] Various implementations include methods and vehicle processing systems for identifying and responding to attacks on a device's (e.g., a vehicle's) camera (referred to herein as "visual attacks"). These implementations address the potential risks to devices (e.g., vehicles) that may be caused by malicious visual attacks and unintentional actions that make images acquired by the camera appear to include false objects or obstacles to be avoided, forged traffic signs, images that may interfere with depth and distance determination, and similarly misleading images that may interfere with the safe and autonomous operation of the device. Various implementations provide methods for identifying actual or potential visual attacks based on inconsistencies (e.g., unexpected or inappropriate shapes, shape shifts, object changes, etc.) and image processing (e.g., classification, labeling, etc.) between multiple images (e.g., multiple image frames from a camera image stream) and image processing. Various implementations may include identifying inconsistencies in multiple images that may result from a malicious attack. Specifically, performing multiple different types of processes on images received from the device's camera to identify or detect different types of image inconsistencies provides multiple ways to detect visual attacks, thereby overcoming vulnerabilities in any single detection method and implementing decision-making mechanisms that reduce false positive determinations. When a visual attack or potential attack is identified, some implementations include the processing system performing one or more mitigation actions to reduce the threat posed by the attack, outputting an indication of the visual attack, and / or reporting the detected attack to an external third party, such as law enforcement or highway maintenance organizations.
[0022] Various implementations can improve the operational safety of autonomous and semi-autonomous devices (e.g., vehicles) by providing effective methods and systems for detecting malicious attacks on camera systems and taking mitigation actions such as reducing the risk to vehicles, providing instructions, and / or reporting attacks to appropriate authorities.
[0023] The terms “airborne” or “within a vehicle” are used interchangeably herein to refer to equipment or components contained within, attached to, and / or carried by a device (e.g., a vehicle or equipment providing the functionality of a vehicle). Airborne equipment typically includes a processing system, which may include one or more processors, a System-on-a-Chip (SoC), and / or a System-on-Instrument (SIP), any of which may include one or more components, systems, units, and / or modules that implement functionality (collectively referred to herein as “processing system” for simplicity). Aspects of airborne equipment and functionality may be implemented in hardware components, software components, or a combination of hardware and software components.
[0024] The term "System-on-a-Chip" (SOC) is used herein to refer to a single integrated circuit (IC) chip containing multiple resources and / or processors integrated on a single substrate. A single SOC may contain circuitry for digital, analog, mixed-signal, and radio frequency functions. A single SOC may also include any number of general-purpose and / or special-purpose processors (digital signal processors, modem processors, video processors, etc.), memory blocks (e.g., ROM, RAM, flash memory, etc.), and resources (e.g., timers, voltage regulators, oscillators, etc.). A SOC may also include software for controlling the integrated resources and processors, as well as software for controlling peripheral devices.
[0025] The term "System-in-Package" (SIP) may be used herein to refer to a single module or package containing multiple resources, computing units, cores and / or processors on two or more IC chips, a substrate, or a System-on-a-Chip (SoC). For example, a SIP may include a single substrate on which multiple IC chips or semiconductor dies are stacked in a vertical configuration. Similarly, a SIP may include one or more multi-chip modules (MCMs) on which multiple ICs or semiconductor dies are packaged into a single substrate. A SIP may also include multiple independent SoCs coupled together and packaged in close proximity via high-speed communication circuitry, such as on a single motherboard or in a single wireless device. The proximity of the SoCs facilitates high-speed communication and the sharing of memory and resources.
[0026] The term "device" is used herein to refer to any of a variety of devices, systems, and equipment that can use camera vision systems and are therefore potentially vulnerable to visual attacks. Some non-limiting examples of devices to which various implementations can be applied include autonomous and semi-autonomous vehicles, mobile robots, mobile machinery, autonomous and semi-autonomous farm equipment, autonomous and semi-autonomous construction and paving equipment, autonomous and semi-autonomous military equipment, etc.
[0027] As used herein, the term "processing system" refers to one or more processors, including multi-core processors, that are organized and configured to perform various computational functions. Various implementation methods may be implemented in one or more of a plurality of processors within any of the various vehicle computers and processing systems described herein.
[0028] Camera systems and image processing play a crucial role in current and future autonomous and semi-autonomous devices, such as autonomous and semi-autonomous vehicles, mobile robots, mobile machinery, and autonomous and semi-autonomous farm equipment. Multiple cameras provide images of the road and surrounding landscape, thus providing data available for navigation (e.g., road following), object recognition, collision avoidance, and hazard detection. The processing of image data in modern autonomous systems has advanced far beyond basic object recognition and tracking, including understanding information posted on street signs, understanding road conditions, and navigating complex road situations (e.g., turning lanes, avoiding pedestrians and cyclists, maneuvering around traffic cones, etc.).
[0029] The processing of camera data fields involves multiple tasks (sometimes referred to as "visual tasks") that are crucial to the safe operation of autonomous devices such as vehicles. Visual tasks typically performed by camera systems include road tracing with depth estimation for path planning, object detection in three dimensions (3D), object identification or classification, traffic sign recognition (including temporary traffic signs and signs reflected in map data), and panoramic segmentation. In modern autonomous driving systems, camera images can be processed by multiple different analysis engines, including trained neural network / artificial intelligence (AI) analysis modules configured to perform a variety of analysis tasks. Analysis tasks can include decision-making tasks and segmentation tasks. Segmentation tasks can involve processing image frames to obtain a variety of different types of information available to the autonomous driving system. Decision-making tasks can involve processing multiple images by a 3D analysis module configured to identify road contours and the location of objects 3D-localized along the road, and analyzing image frames to identify and interpret objects in real time, such as understanding the meaning of traffic signs.
[0030] A crucial operation achieved through processing image data is object detection and recognition. Examples of objects that should be identified, categorized, and in some cases interpreted or understood include traffic signs, pedestrians, other vehicles, road obstacles, and road features that differ from the information included in detailed map data and observed during previous driving experiences.
[0031] Traffic signs are a type of object that needs to be identified, categorized, and displayed, and are understandable for autonomous vehicle applications, allowing the guidance and regulations indicated by the signs to be included in the decision-making of autonomous driving systems. Typically, traffic signs have recognizable shapes depending on the type of information displayed (e.g., stop, yield, speed limit, etc.). However, sometimes the information displayed differs from the meaning or classification corresponding to the shape, such as text in different languages or shapes that are not actually observable as traffic signs (e.g., advertisements, T-shirt designs, protest signs, etc.). Furthermore, traffic signs can indicate requirements and regulations that are inconsistent with information appearing in map data that autonomous driving systems may rely on.
[0032] Pedestrians and other vehicles are clearly important objects to identify, label, or classify, and to track closely to avoid collisions and properly plan vehicle paths. Classifying pedestrians and other vehicles can help predict their future location or trajectory, which is important for future planning performed by autonomous driving systems.
[0033] To achieve such object detection, recognition, and classification, multiple images can be processed through one or more complex processes. These processes may include receiving input frames from one or more camera systems and processing each frame to identify significant objects within each frame, such as traffic signs, vehicles, people, and other objects. Certain types of objects (such as traffic signs, pedestrians, and other vehicles) require identification and classification, allowing them to be given special recognition processing. When such types of objects are identified in multiple images, portions of the image frames containing the identified objects can be extracted and buffered in memory, allowing the trajectory of the object within the field of view of the device's cameras to be extracted. This information can be applied to a classifier that performs the function of classifying objects and obtaining information about the objects or within objects that is available to the navigation device. The recognition process may involve using a trained neural network / AI recognition model, already trained on one or more image databases (e.g., a database of traffic sign shapes, a database of vehicle model shapes, a database of animal shapes, etc.), to identify and classify objects based on features extracted from a series of images. The output of this process can be information available to the autonomous driving system, including location, classification, and, in the case of traffic signs, information displayed on the object.
[0034] In addition to identifying, classifying, and obtaining information about detected objects, image data needs to be processed in a manner that allows for frame-by-frame tracking of the positions of these objects, enabling the determination of the object's trajectory relative to the device (or the device's trajectory relative to the object) to support navigation and collision avoidance functions. This processing can also utilize a trained neural network / AI model that receives image frames and outputs various types of information, such as features, box size, intra-frame position (e.g., center offset), and importance or frequency (e.g., a "heatmap"). The frame-by-frame object tracking process may involve using an association algorithm that transforms information about bounding boxes and identified features (e.g., size, shape, frame position, color, etc.) into a data structure called a "track segment pool." Using such a data structure, the association algorithm can identify boxes and features appearing in an image frame and associate them with boxes and features appearing in previous image frames. In some implementation descriptions, this inter-frame association process is sometimes referred to as associating features in "frame t" with features in the next or subsequent frame, referred to as "frame (t=1)," where t is any time within the image frame sequence.
[0035] Visual attacks, as well as obfuscated or conflicting images that could mislead the image analysis process of autonomous driving systems, can originate from many different sources and involve a variety of different types of attacks. Visual attacks can target semantic segmentation operations, depth estimation, and / or object detection and recognition functions that are critical image processing capabilities of autonomous driving systems. Visual attacks can include projector attacks and patching attacks.
[0036] In projector vision attacks, images are projected onto a vehicle camera by a projector with the aim of creating false or misleading image data to interfere with autonomous driving systems. For example, a projector can be used to project an image onto a road so that, when viewed in the camera's two-dimensional visual plane, the image appears three-dimensional and resembles an object to be avoided. An example of this type of attack would be a projection of a picture or shape resembling a pedestrian (or other object) onto the road, which, when viewed from the vehicle camera's perspective, appears to be a pedestrian in the road. Another example is a projector that projects images onto structures along the road, such as projecting an image of a stop sign onto a building wall that would otherwise be blank. Yet another example is a projector that directly targets a device camera, injecting an image (e.g., an incorrect traffic sign) into the image.
[0037] Examples of patched visual attacks include images of recognizable objects, such as traffic signs, that are fake, inappropriate, or placed where such objects shouldn't be. For instance, a T-shirt with an image of a stop sign could interfere with an autonomous driving system, making it difficult to determine whether a vehicle should stop or ignore the sign, especially if the person wearing the shirt is walking or running rather than at or near an intersection. As another example, images of the rear of a vehicle or interfering shapes could interfere with image processing modules that estimate the depth and 3D localization of objects.
[0038] While some methods have been proposed to address image distortion and interference, a comprehensive multi-factor approach has not been identified. Therefore, camera-based autonomous driving systems remain vulnerable to many visual attacks.
[0039] Various implementations offer an integrated security solution to address threats posed by attacks on cameras of devices supporting autonomous driving and control systems. These implementations include the use of multiple different types of detection methods (referred to as detectors) that employ various techniques to identify inconsistencies within and across image frames. In this way, the system is able to identify different types of visual attacks without being easily limited by any single method. Implementations include methods based on temporal consistency checks (i.e., whether an object or feature changes significantly from one frame to the next), inconsistency count checks (i.e., whether the number of inconsistencies detected across multiple images exceeds a threshold), and historical consistency checks (i.e., identifying when features present in previous driving now appear in the current image at the same location or disappear from the current image). Some implementations include methods that use the results of the detection methods to determine whether inconsistencies across multiple images indicate or suggest a possible camera or visual attack. Some implementations include performing one or mitigation actions to protect the device, such as pruning, deleting, or ignoring suspicious information obtained from image processing to avoid misleading the autonomous driving system. Some implementations include reporting information about detected visual attacks to third parties, such as authorities that can take action to remove or suppress the source of the attack or image inconsistencies. Some implementations may include outputting instructions regarding visual attacks, such as informing operators of such attacks in displays or notifications.
[0040] Various implementation schemes can be implemented in various devices, Figure 1A and Figure 1B The example provided is a non-limiting illustration of a means of transport, form 100. (See reference.) Figure 1A and Figure 1BThe vehicle 100 may include a control unit 140 and a plurality of sensors 102 to 138, including a satellite geolocation system receiver 108, occupancy sensors 112, 116, 118, 126, 128, tire pressure sensors 114, 120, cameras 122, 136, microphones 124, 134, an impact sensor 130, radar 132, and lidar 138. These sensors 102 to 138, located in or on the vehicle, can be used for various purposes, such as autonomous and semi-autonomous navigation and control, collision avoidance, location determination, etc., and to provide sensor data about objects and people in or on the vehicle 100. Sensors 102 to 138 may include one or more of a wide variety of sensors capable of detecting various information available for navigation, collision avoidance, and autonomous and semi-autonomous navigation and control. Each of the sensors 102 to 138 may communicate wirelessly with the control unit 140 and with each other. Specifically, the sensors may include one or more cameras 122, 136, or other optical or photoelectric sensors. Cameras 122, 136, or other optical or photoelectric sensors may include outward-facing sensors for imaging objects outside the vehicle 100 and / or in-vehicle sensors for imaging objects (including passengers) inside the vehicle 100. In some embodiments, the number of cameras may be less than two or more than two. For example, more than two cameras may be present, such as two front-facing cameras with different fields of view (FOV), four side-facing cameras, and two rear-facing cameras. Sensors may also include other types of object detection and ranging sensors, such as radar 132, lidar 138, IR sensors, and ultrasonic sensors. The sensors may also include tire pressure sensors 114, 120, humidity sensors, temperature sensors, satellite geolocation sensors 108, accelerometers, vibration sensors, gyroscopes, gravimeters, impact sensors 130, force gauges, pressure gauges, strain sensors, fluid sensors, chemical sensors, gas content analyzers, hazardous substance sensors, microphones 124, 134 (inside or outside the vehicle 100), occupancy sensors 112, 116, 118, 126, 128, proximity sensors, and other sensors.
[0041] The vehicle control unit 140 may be configured with processor-executable instructions to perform operations in some embodiments using information received from various sensors, particularly cameras 122 and 136. In some embodiments, the control unit 140 may supplement the processing of multiple images with distance and relative positioning (e.g., relative azimuth) obtainable from radar 132 and / or lidar 138 sensors. The control unit 140 may also be configured to control the steering, braking, and speed of the vehicle 100 when operating in autonomous or semi-autonomous mode using information about other vehicles determined using methods from some embodiments. In some embodiments, the control unit 140 may be configured to operate as an autonomous driving system (ADS). In some embodiments, the control unit 140 may be configured to operate as an automated driver assistance system (ADAS).
[0042] Figure 1C This is a component block diagram of system 150, which illustrates components and supporting systems applicable to implementing some implementation schemes. (See reference) Figure 1A , Figure 1B and Figure 1C The vehicle 100 may include a control unit 140, which may include various circuits and devices for controlling the operation of the vehicle 100. Figure 1C In the illustrated example, control unit 140 includes processor 164, memory 166, input module 168, output module 170, and radio module 172. Control unit 140 may be coupled to and configured to control driving control component 154, navigation component 156, and one or more sensors 158 of vehicle 100. Radio module 172 may be configured to communicate with base station 180 via wireless communication link 182 (e.g., 5G, etc.), which provides connectivity to a server 184 of a third party (such as a law enforcement agency like a highway maintenance agency) via network 186 (e.g., the Internet).
[0043] Figure 2 Examples are shown of subsystems, computing elements, computing devices, or units that can be used within a vehicle management system 200, specifically within a vehicle 100. (See reference) Figures 1A to 2 In some implementations, various computing elements, computing devices, or units within the vehicle management system 200 may be implemented within a system of interconnected computing devices (i.e., subsystems) that communicate data and commands to each other (e.g., by...). Figure 2 (As indicated by the arrow in the diagram). In other embodiments, various computing elements, computing devices, or units within the vehicle management system 200 may be implemented within a single computing device, such as individual threads, processes, algorithms, or computing elements. Therefore, Figure 2Each exemplified subsystem / computing element is also generally referred to herein as a "module," which may be implemented in one or more processing systems constituting the vehicle management system 200. However, the use of the term "module" in describing various implementations is not intended to imply or require the corresponding functionality to be implemented within a single autonomous (or semi-autonomous) vehicle management system computing device, in multiple computing systems, or in a combination of dedicated hardware modules, software-implemented modules, and dedicated processing systems within a distributed vehicle computing system, although each is a potential specific implementation. Rather, the use of the term "module" is intended to encompass subsystems with independent processing systems, computing elements (e.g., threads, algorithms, subroutines, etc.) running in one or more computing devices and processing systems, and combinations of subsystems and computing elements.
[0044] In various implementations, the vehicle management system 200 may include a radar sensing module 202, a camera sensing module 204, a positioning engine module 206, a map fusion and arbitration module 208, a route planning module 210, a sensor fusion and road world model (RWM) management module 212, a motion planning and control module 214, and a behavior planning and prediction module 216. Modules 202 to 216 are merely examples of some modules in one example configuration of the vehicle management system 200. In other configurations consistent with some implementations, additional modules may be included, such as additional modules for other sensing sensors (e.g., a LiDAR sensing module, etc.), additional modules for planning and / or control, additional modules for modeling, etc., and / or some modules of vehicle 202 to 216 may be excluded from the vehicle management system 200. Each of modules 202 to 216 may exchange data, calculation results, and commands with each other. Examples of some interactions between modules 202 to 216 are provided by Figure 2 The arrows in the diagram illustrate this. Furthermore, the vehicle management system 200 can receive and process data from sensors (e.g., radar, lidar, cameras, inertial measurement units (IMUs), etc.), navigation systems (e.g., Global Navigation Satellite System (GNSS) receivers, IMUs, etc.), vehicle networks (e.g., Controller Area Network (CAN) buses), and databases in memory (e.g., digital map data). The vehicle management system 200 can output vehicle control commands or signals to the Adaptive Drive-by-Wire (ADS) system / control unit 220, which is a system, subsystem, or computing device that directly interfaces with the vehicle's steering, throttle, and braking controls. Figure 2 The configurations of the vehicle management system 200 and ADS system / control unit 220 illustrated herein are merely example configurations, and other configurations of the vehicle management system and other vehicle components may be used in some implementations. As an example, Figure 2The configurations of the vehicle management system 200 and ADS system / control unit 220 illustrated herein can be used in vehicles configured for autonomous or semi-autonomous operation, while different configurations can be used in non-autonomous vehicles.
[0045] Camera perception module 204 may receive data from one or more cameras (such as cameras (e.g., 122, 136)) and process the data to identify and determine the location of other vehicles and objects (e.g., passengers, etc.) near and / or inside vehicle 100. Camera perception module 204 may include using neural network processing and artificial intelligence methods to identify objects and vehicles and pass such information to sensor fusion and RWM trained model 212 and / or other modules, such as augmented reality projection system / control unit 221.
[0046] The radar perception module 202 may receive data from one or more detection and ranging sensors (such as radar (e.g., 132) and / or lidar (e.g., 138)) and process the data to identify and determine the location of other vehicles and objects in the vicinity of vehicle 100. The radar perception module 202 may include the use of neural network processing and artificial intelligence methods to identify objects and vehicles and pass such information to a sensor fusion and RWM-trained model 212.
[0047] The positioning engine module 206 can receive data from various sensors and process that data to determine the location of the vehicle 100. These various sensors may include, but are not limited to, GNSS sensors, IMUs, and / or other sensors connected via a CAN bus. The positioning engine module 206 may also utilize input from one or more cameras (such as cameras (e.g., 122, 136)) and / or any other available sensors (such as radar, lidar, etc.).
[0048] The map fusion and arbitration module 208 can access data within a high-definition (HD) map database and receive output from the positioning engine module 206, processing the data to further determine the location of vehicle 100 within the map, such as its position within a traffic lane, its position within a street map, etc. The HD map database can be stored in memory (e.g., memory 166). For example, the map fusion and arbitration module 208 can convert latitude and longitude information from GNSS data into a position within a ground road map contained in the HD map database. GNSS positioning locking includes errors, so the map fusion and arbitration module 208 can be used to determine the best guessed position of the vehicle within the road based on arbitration between GNSS coordinates and HD map data. For example, while GNSS coordinates might place the vehicle near the middle of a two-lane road in the HD map, the map fusion and arbitration module 208 can determine, based on the direction of travel, that the vehicle is most likely aligned with the lane in the same direction of travel. The map fusion and arbitration module 208 can then pass the map-based location information to a sensor fusion and RWM-trained model 212.
[0049] Route planning module 210 can use an HD map and input from the operator or dispatcher to plan a route to a specific destination for vehicle 100. Route planning module 210 can pass map-based location information to sensor fusion and RWM-trained model 212. However, the use of prior maps is not required by other modules (such as sensor fusion and RWM-trained model 212). For example, other processing systems can operate and / or control the vehicle based solely on perception data without providing a map, thereby constructing concepts of lanes, boundaries, and local maps as perception data is received.
[0050] The sensor fusion and RWM trained model 212 can receive data and outputs generated by the radar sensing module 202, camera sensing module 204, map fusion and arbitration module 208, and route planning module 210, and use some or all of these inputs to estimate or refine the position and state of the vehicle 100 relative to the road, other vehicles on the road, and other objects near and / or inside the vehicle 100. For example, the sensor fusion and RWM trained model 212 can combine image data from the camera sensing module 204 with arbitration map position information from the map fusion and arbitration module 208 to refine the determined location of the vehicle within a traffic lane. As another example, the sensor fusion and RWM trained model 212 can combine object recognition and image data from the camera sensing module 204 with object detection and ranging data from the radar sensing module 202 to determine and refine the relative positions of other vehicles and objects near the vehicle. As another example, the sensor fusion and RWM trained model 212 can receive information about the location and direction of travel of other vehicles from vehicle-to-vehicle (V2V) communication (such as via a CAN bus) and combine this information with information from the radar perception module 202 and the camera perception module 204 to refine the position and motion of other vehicles. The sensor fusion and RWM trained model 212 can output refined position and status information of vehicle 100, as well as refined position and status information of other vehicles or objects near vehicle 100 or objects within vehicle 100, to the motion planning and control module 214, the behavior planning and prediction module 216, and / or the augmented reality projection system / control unit 221. As another example, the sensor fusion and RWM trained model 212 can apply facial recognition technology to images to identify specific facial patterns inside and / or outside the vehicle.
[0051] As a further example, the sensor fusion and RWM-trained model 212 can use dynamic traffic control commands to guide vehicle 100 to change speed, lane, direction of travel, or other navigation elements, and combine this information with other received information to determine refined location and status information. The sensor fusion and RWM-trained model 212 can output the refined location and status information of vehicle 100, as well as the refined location and status information of other vehicles and objects near vehicle 100 or objects within vehicle 100, via wireless communication (such as via C-V2X connection, other wireless connections, etc.) to motion planning and control module 214, behavior planning and prediction module 216, augmented reality projection system / control unit 221, and / or devices remote from vehicle 100, such as data servers, other vehicles, etc.
[0052] As a further example, the sensor fusion and RWM trained model 212 can monitor perception data from various sensors (such as perception data from radar perception module 202, camera perception module 204, other perception modules, etc.) and / or data from one or more sensors themselves to analyze the status in the vehicle's sensor data. The sensor fusion and RWM trained model 212 can be configured to detect sensor data status, such as sensor measurements being at, above, or below thresholds, the occurrence of certain types of sensor measurements (e.g., seat positioning movement, seat height change, etc.), and can output the sensor data as part of the refined position and status information of the vehicle 100, provided via wireless communication (such as via C-V2X connection, other wireless connections, etc.) to the behavior planning and prediction module 216, the augmented reality projection system / control unit 221, and / or devices remote from the vehicle 100 (such as data servers, other vehicles, etc.).
[0053] Detailed location and status information may include vehicle descriptors associated with the vehicle and its owner and / or operator, such as: vehicle specifications (e.g., size, weight, color, type of onboard sensors, etc.); vehicle location, speed, acceleration, direction of travel, attitude, orientation, destination, fuel / power level, and other status information; vehicle emergency status (e.g., the vehicle is an emergency vehicle or a private individual is in an emergency); vehicle restrictions (e.g., weight / width load, turning restrictions, high-occupancy vehicle (HOV) authorization, etc.); vehicle capabilities (e.g., all-wheel drive, four-wheel drive, snow tires, chains, supported connection types, onboard sensor operating status, onboard sensor resolution level, etc.); equipment issues (e.g., low tire pressure, weak brakes, sensor malfunction, etc.); owner / operator travel preferences (e.g., preferred lanes, roads, routes and / or destinations, preference for avoiding tolls or highways, preference for the fastest route, etc.); permission to provide sensor data to a data broker server (e.g., 184); and / or owner / operator identification information.
[0054] The behavior planning and prediction module 216 of the autonomous transportation system 200 can use refined position and state information of the vehicle 100, as well as position and state information of other vehicles and objects output from the sensor fusion and RWM-trained model 212, to predict the future behavior of other vehicles and / or objects. For example, the behavior planning and prediction module 216 can use such information to predict the future relative position of these other vehicles based on its own vehicle positioning and speed, as well as the positioning and speed of other vehicles in the vicinity. Such predictions can take into account information from HD maps and route planning to anticipate changes in the relative vehicle position as the primary vehicle and other vehicles travel along a road. The behavior planning and prediction module 216 can output the behavior and position predictions of other vehicles and objects to the motion planning and control module 214. Additionally, the behavior planning and prediction module 216 can use object behavior combined with position predictions to plan and generate control signals for controlling the motion of the vehicle 100. For example, based on route planning information, detailed location data in road information, and the relative positions and movements of other vehicles, the behavior planning and prediction module 216 can determine that vehicle 100 needs to change lanes and accelerate, such as to maintain or achieve a minimum distance from other vehicles and / or prepare for a turn or exit. Therefore, the behavior planning and prediction module 216 can calculate or otherwise determine changes in wheel steering angle and throttle setting, which, along with various such parameters necessary to achieve such lane changes and accelerations, will be commanded to the motion planning and control module 214 and the ADS system / control unit 220. One such parameter could be a calculated steering wheel command angle.
[0055] The motion planning and control module 214 can receive data and information output from the sensor fusion and RWM-trained model 212, as well as behavior and position predictions of other vehicles and objects from the behavior planning and prediction module 216. It uses this information to plan and generate control signals for controlling the motion of the vehicle 100, and to verify that such control signals meet the safety requirements of the vehicle 100. For example, based on route planning information, detailed location data in road information, and the relative positions and movements of other vehicles, the motion planning and control module 214 can verify various control commands or instructions and transmit them to the ADS system / control unit 220.
[0056] The ADS system / control unit 220 can receive commands or instructions from the motion planning and control module 214 and convert such information into mechanical control signals for controlling the wheel angles, braking, and throttle of the vehicle 100. For example, the ADS system / control unit 220 can respond to a calculated steering wheel command angle by transmitting a corresponding control signal to the steering wheel controller.
[0057] The ADS system / control unit 220 can receive data and information outputs from the motion planning and control module 214 and / or other modules in the vehicle management system 200, and determine whether an event is occurring that needs to be notified to the decision-maker in the vehicle 100 based on the received data and information outputs.
[0058] In some embodiments, the vehicle management system 200 may include functionality to perform safety checks or monitoring of various commands, plans, or other decisions made by various modules that may affect the safety of vehicles and occupants. Such safety check or monitoring functionality may be implemented within a dedicated module or distributed among various modules and included as part of that functionality. In some embodiments, various safety parameters may be stored in memory, and the safety check or monitoring functionality may compare determined values (e.g., relative distance to nearby vehicles, distance to the road centerline, etc.) with corresponding safety parameters and issue warnings or commands if the safety parameters are violated or will be violated. For example, the safety or supervisory function in the behavior planning and prediction module 216 (or a separate module) can determine the current or future separation distance between another vehicle (such as the model 212 refined by sensor fusion and RWM training) and the vehicle (e.g., based on the world model refined by the model 212 trained by sensor fusion and RWM), compare the separation distance with a safe separation distance parameter stored in memory, and issue instructions to the motion planning and control layer 214 to accelerate, decelerate, or turn if the current or predicted separation distance violates the safe separation distance parameter.
[0059] Figure 3 This is a block diagram illustrating example components of a System-on-Chip (SOC) 300 for use in a processing system (e.g., a V2X processing system) according to various implementation schemes. (See also...) Figures 1A to 3 The processing device SOC 300 may include several heterogeneous processors, such as a digital signal processor (DSP) 303, a modem processor 304, an image and object recognition processor 306, a mobile display processor 307, an application processor 308, and a resource and power management (RPM) processor 317. The processing device SOC 300 may also include one or more coprocessors 310 (e.g., vector coprocessors) connected to one or more of the heterogeneous processors 303, 304, 306, 307, 308, and 317.
[0060] Each of these processors may include one or more cores and an independent / internal clock. Each processor / core may perform operations independently of the other processors / cores. For example, the processing device SOC 300 may include a processor running a first type of operating system (e.g., FreeBSD, LINUX, OS X, etc.) and a processor running a second type of operating system (e.g., Microsoft Windows). In some embodiments, the application processor 308 may be the main processor, central processing unit (CPU), microprocessor unit (MPU), arithmetic logic unit (ALU), etc. of the SOC 300. The graphics processor 306 may be a graphics processing unit (GPU).
[0061] The processing device SOC 300 may include analog circuitry and custom circuitry 314 for managing sensor data, analog-to-digital conversion, wireless data transmission, and performing other specialized operations, such as processing encoded audio and video signals for rendering in a web browser. The processing device SOC 300 may also include system components and resources 316, such as voltage regulators, oscillators, phase-locked loops, peripheral bridges, data controllers, memory controllers, system controllers, access ports, timers, and other similar components for supporting processors and software clients (e.g., web browsers) running on computing devices.
[0062] The processing device SOC 300 may also include a dedicated circuit (CAM) 305 for camera actuation and management. This CAM includes, provides, controls, and / or manages the operation of one or more cameras (e.g., a main camera, webcam, 3D camera, etc.), video display data from camera firmware, image processing, video preprocessing, video front-end (VFE), embedded JPEG, high-definition video codecs, etc. The CAM 305 may be a separate processing unit and / or include a separate or internal clock.
[0063] In some embodiments, the image and object recognition processor 306 may be configured with processor-executable instructions and / or dedicated hardware configured to perform image processing and object recognition analysis as described in various embodiments. For example, the image and object recognition processor 306 may be configured to process images received from a camera via CAM 305 to identify and / or mark other vehicles. In some embodiments, the processor 306 may be configured to process radar or lidar data.
[0064] System components and resources 316, analog and custom circuitry 314, and / or CAM 305 may include circuitry for interfacing with peripheral devices such as cameras, radar, lidar, electronic displays, wireless communication devices, external memory chips, etc. Processors 303, 304, 306, 307, and 308 may be interconnected via interconnect / bus module 324 to one or more memory elements 312, system components and resources 316, analog and custom circuitry 314, CAM 305, and RPM processor 317. This interconnect / bus module may include reconfigurable logic gate arrays and / or implement bus architectures (e.g., CoreConnect, AMBA, etc.). Communication may be provided by advanced interconnects such as high-performance on-chip networks (NoC).
[0065] The processing device SOC 300 may also include input / output modules (not shown) for communicating with external resources such as clock 318 and voltage regulator 320. External resources (e.g., clock 318, voltage regulator 320) may be shared by two or more internal SOC processors / cores (e.g., DSP 303, modem processor 304, graphics processor 306, application processor 308, etc.).
[0066] In some implementations, the processing device SOC 300 may be included in a control unit (e.g., 140) for use in a vehicle (e.g., 100). The control unit may include communication links for communicating with a telephone network (e.g., 180), the Internet, and / or a web server (e.g., 184), as described.
[0067] The processing device SOC 300 may also include additional hardware and / or software components suitable for collecting sensor data from sensors, including motion sensors (e.g., accelerometers and gyroscopes of an IMU), user interface elements (e.g., input buttons, touchscreen displays, etc.), microphone arrays, sensors for monitoring physical conditions (e.g., position, orientation, motion, orientation, vibration, pressure, etc.), cameras, compasses, GPS receivers, and communication circuitry (e.g., Bluetooth). ® (such as WLAN, Wi-Fi, etc.) and other well-known components of modern electronic devices.
[0068] Figure 4 This is a block diagram illustrating various operations performed on images as part of an autonomous driving system, and the processing involved in implementing these operations in various implementation schemes. (Reference) Figures 1A to 4Image frames 402 from multiple device cameras can be received by an image processing system such as a camera perception module 204. This image processing system may include multiple modules, processing systems, and trained machine model / AI modules configured to perform various operations necessary to obtain information from the images to support vehicle navigation and safe operation. While not implying inclusion, Figure 4 Examples of the processes involved in supporting autonomous vehicle operation, identifying visual attacks, and taking mitigation actions are illustrated under various implementation schemes.
[0069] Image frame 402 can be processed by object detection module 404, which performs operations associated with detecting objects within the image frame based on various image processing techniques. As discussed, autonomous vehicle image processing involves multiple detection methods and analysis modules that focus on different aspects of using image streams to provide the information needed for safe navigation of autonomous driving systems. The processing of image frames in object detection module 404 can involve multiple different detectors and modules that process the image in different ways to identify objects, define boundary blocks covering the objects, and identify the location of detected objects within the frame coordinates. The outputs of various detection methods can be combined in an ensemble detection, which can be a list, table, or data structure of detections performed by the individual detectors processing the image frame. Therefore, the ensemble detection in object detection module 404 can aggregate the outputs of various detection mechanisms and modules for object classification and tracking and vehicle control decisions.
[0070] As also discussed, image processing supporting autonomous driving systems involves other image processing tasks 406. As an example of other tasks, image frames can be analyzed to determine road features and the 3D depth of detected objects. Other processing tasks 406 may include panoptic segmentation, a computer vision task that includes both instance segmentation and semantic segmentation. Instance segmentation involves identifying and classifying objects of multiple categories observed within an image frame. Semantic segmentation is the task of associating individual pixels or groups of pixels in a digital image with categories or classification labels such as “tree,” “traffic sign,” “pedestrian,” “road,” “building,” “car,” “sky,” etc. By addressing instance segmentation and semantic segmentation together, panoptic segmentation enables autonomous driving systems to understand a given scene in greater detail.
[0071] The outputs of object detection method 404 and other tasks 406 can be used for object classification 410. As described, this may involve classifying features and objects detected in an image frame using classifications important to the decision-making process of an autonomous driving system (e.g., road features, traffic signs, pedestrians, other vehicles, etc.). As illustrated, the methods described herein can be used to examine identified features, such as traffic signs 408 within segments or bounding boxes in an image frame, to assign classifications to individual objects and obtain information about objects or features (e.g., a speed limit of 50 km / h based on identified traffic sign 408). Also as part of object classification 410, the techniques described herein can be used to examine the image frame for projection attacks.
[0072] In operation 414, the output of other task 406 can also be analyzed for the correlation of features and elements from one frame to the next. As further described herein, such correlations can be part of identifying inconsistencies in an image that may indicate camera or visual attacks. Furthermore, such correlations can be used in operation 416 to determine the plausibility of various tasks.
[0073] The output of object classification 410 can be used to track various features and objects from one frame to the next 412. As described above, feature and object tracking is important for identifying the trajectory of features / objects relative to the vehicle for navigation and collision avoidance. Therefore, tracking operation 412 can provide safe multi-object tracking 420 to support vehicle control functions 422 of an autonomous driving system. Additionally, in some implementation methods, feature / object tracking can be used to detect inconsistencies that may indicate or suggest a visual attack. Therefore, the operation of tracking 412 can be used to make a safety decision 418, which can be used to trigger or define a safety response 424, such as one or more responses to mitigate the risk or impact of a visual attack.
[0074] Figure 5This is a functional block diagram illustrating operation 500 in various implementations. As described, vehicle camera 502 may provide multiple images (such as image frame streams from one or more cameras) to various image processing modules, processing systems, and AI modules that identify, identify, and measure various objects within image frames. Such measurements may include identifying features and objects and assigning identifiers (IDs) to identified features, objects (or bounding boxes surrounding such elements) within image frames. As an example, object measurement may include identifying traffic signs within an image, such as enclosing the sign within a bounding box with identified coordinates (e.g., pixel coordinates) within the image, and labeling detected traffic signs (e.g., labels including the classification of the detected traffic signs). Object measurement may also include assigning labels to patches detected in the image (e.g., features that look like traffic signs, pedestrians, or cars but are actually two-dimensional images, etc.). Measurement may also include assigning misclassification labels, lighting check labels, light values, semantic consistency labels, depth plausibility labels, contextual consistency labels, and object label consistency labels, all of which can then be used for consistency checks to identify inconsistencies within and across image frames that may indicate or suggest a visual attack.
[0075] These measurements 504 can be associated with a trajectory spanning multiple image frames in box 506. That is, individual features / objects identified in box 504 and their associated labels can be traced across a series of image frames to identify the marked features / objects relative to the vehicle's trajectory. In box 508, this information can be used to update the trajectory and perform multiple analyses to detect visual attacks or inconsistencies that suggest or indicate visual attacks. Such tests or detectors may include a time consistency check 510, a consistency counter check 512, and a past history check 514. The results of these analyses can then be processed together in box 516 to make a security decision. The output of the security decision can be an appropriate security response in box 518. For example, a decision can be made to determine whether an attack exists in block 520. If a decision is made that an attack exists, the attack can be reported to the agency in box 522, and actions can be taken to protect the vehicle, such as removing suspicious information from data used when operating the vehicle (524). If a decision is made that no attack exists, the security response system can remain idle in box 526.
[0076] Figure 6 This is a block diagram illustrating the operations and data structures involved in an execution time consistency check according to some implementation schemes. This time consistency check is used to identify potential visual attacks based on the characteristic of inconsistency between objects from one frame to the next. Reference Figures 1A to 6 Temporal consistency checks can be configured to detect temporal inconsistencies in values and / or classifications between two consecutive or sequential frames based on objects, features, or trajectories.
[0077] refer to Figure 6 Measurements of features and objects (e.g., measurement 504) can be considered together in a data structure for a given image frame (e.g., frame (t+1)), which includes measurements of bounding boxes, classifications, and attack classifications for each identifier's features (e.g., #1 to #N). These measurements 602 in each frame can be processed together with values from the previous frame 606 (e.g., frame t) by an association algorithm 604 to achieve tracking of a given feature or object from one frame to the next. This can be stored in a trajectory segment pool data structure 608, which includes bounding boxes, features, packet classifications, feature categorization, and tracking of leading class consistency from frame to frame for multiple trajectories. This process of associating bounding boxes, feature classifications, etc., frame by frame and constructing the trajectory segment pool data structure across multiple trajectories enables the resolution of inconsistencies from one frame to another across various trajectory identifiers through an extended sequence of image frames.
[0078] A temporal consistency check can be performed to look for changes between two consecutive frames that are inconsistent with reality for a given feature or object. To be sensitive to various attack methods and / or image processing that may attempt to fool cameras and other imaging devices, the detection method (or detector) can evaluate features / objects based on multiple measurements or factors across multiple image frames. Therefore, instead of only considering the labels assigned to features / objects from one frame to the next, the detection method can also evaluate other measurements or values individually or collectively. The detection method can determine whether a feature changes from one frame to another, such as in terms of size, shape, location, or classification, and whether the change is inconsistent with reality (e.g., moving too quickly, changing shape in an unnatural way, etc.).
[0079] Temporal consistency checks can be performed based on two or more values (e.g., Value A, Value B) of a feature / object between two image frames, such as whether the difference between two values exceeds a threshold of natural or expected change, and thus indicates an attack. In equation form, this can be expressed as: |Value A - Value B| > Threshold, where Value A and Value B are the values of a feature from image (t) and the value of a feature from trajectory (t-1). Values Value A, Value B, and Threshold can be floating-point values, vectors, or matrices. The function can return a binary output, such as inconsistency if the difference exceeds the threshold, or consistency if the difference is ≤ the threshold. Temporal consistency features can track multiple characteristics that identify an object. This can be a list of functions output after each successful association loop. For example, temporal consistency features can include inconsistency values across multiple time periods or image frames (e.g., consistent (t-2), inconsistent (t-1), inconsistent (t)).
[0080] As an example, a temporal consistency check can determine whether a bounding box has moved or shifted from one frame to the next by an amount exceeding a threshold, such as movement that indicates unnaturalness or exceeds the capacity of the category assigned to the object or box. For example, a traffic sign will not actually move, and therefore its frame-by-frame movement is limited by relative movement that can be expected based on the vehicle's own speed. As another example, a pedestrian can move frame-by-frame, based on relative movement relative to a vehicle plus the pedestrian's own walking or running speed (which will be defined by a threshold). If an object or the bounding box around the object moves frame-by-frame at a rate indicating an unrealistic speed, this is an indication of inconsistency, and this inconsistency could be evidence of a visual attack. Other dynamic features that can be compared to a threshold include changes in color (e.g., color shifts occurring at a rate greater than expected for natural objects), changes in lighting, and changes in depth within the field of view.
[0081] Other types of changes can also indicate visual attacks, such as outlier changes in non-safe features, or changes in values that are expected to be static, such as changes in traffic sign detection labels, traffic sign classification labels, or semantic labels from one frame to the next.
[0082] Temporal consistency classification can include two categories of checks and features. The first category tracks changes in outliers of non-safe features, such as changes in the label or classification assigned to a given feature or object (e.g., a traffic sign). This first category can distinguish between static features that are not expected to move from one frame to the next and dynamic features that are expected to move, and therefore can be compared to a movement or velocity threshold, among other considerations. The second category relates to changes in attack classification, which can monitor the output from an attack classifier within the image processing system. For example, temporal consistency classification could include tracking inconsistencies in output from a projector detector and inconsistencies in output from a patch detector.
[0083] The threshold applied in temporal consistency checks is important because small changes, especially in frame-to-frame positioning, are expected. Using threshold testing allows temporal consistency check detectors to avoid false detections based on normal movement or changes of specific objects or features.
[0084] Figure 7This is a block diagram illustrating the operations and data structures involved in performing inconsistency counter checks according to some implementation schemes. Different kinds of attacks and different causes of inconsistencies from one image frame to the next may involve a form of noise that can be explained by tracking or counting the number of inconsistencies within a given trajectory. For example, patching or projector attacks may be consistent noise because patches remain on traffic signs throughout the detection process, and the effects of patches are not always stable, which can lead to misclassification. On the other hand, solar glare or natural objects moving through the frame (such as leaves falling in front of vehicles) may cause misclassification of objects by temporarily fixing or interfering with the image of the object. By tracking or counting the number of inconsistencies within a given trajectory, non-attack inconsistencies may have a smaller count across multiple image frames, while visual attack cases are more likely to persist throughout the observation period.
[0085] refer to Figures 1A to 7 The consistency counter check may involve using a counting function 704 to compare various values of boxes, features, classifications, etc., within the trajectory segment pool 702 of the first frame (e.g., frame t) with corresponding values in subsequent frames (e.g., frame (t+1)). As illustrated, each of the frame trajectory segment pool data structures 702, 706 may include a data field for counting at corresponding times. Across multiple frames, if the counter for a given inconsistency exceeds a threshold, this can indicate a visual attack. The counting check can be applied to a reference... Figure 6 Each feature and object identified and tracked in the temporal consistency check described is described. If the same inconsistency occurs in the trajectory segment pool of a subsequent frame, the counter for that feature in the subsequent trajectory segment pool is incremented.
[0086] A suitable threshold for identifying various types of potential visual attacks could be determined by recording different types of image frame-to-frame inconsistencies observed during a series of non-adversarial road tests (i.e., road tests where visual attacks are known to be absent). Therefore, this number could be based on the maximum number of inconsistencies seen during the longest lifetime of the object-related light trail under normal driving conditions when no threat exists, and the threshold number could vary based on the classification of the object or feature. For example, a larger threshold might be appropriate for images of other vehicles (which are moving and present a collision hazard) compared to traffic signs (which are stationary and located outside the road).
[0087] Figure 8This is a block diagram illustrating the operations and data structures involved in performing a past history check according to some implementation schemes. Another check can be performed to detect unexpected changes between objects or features that have been observed in the past (e.g., when a vehicle has traveled the same route in previous days) and the observed objects or features. For example, frequent (e.g., daily) travel between home and work or other common destinations makes it possible to build a database of frequently observed objects or features that are always tracked (such as traffic signs, traffic lights, etc.). If the same object is observed on many such trips, it can be assumed that these objects are permanent features and should be observed every time. However, if the common features of the objects are suddenly lost, this inconsistency between the current trajectory and past trajectories for normal, fixed objects (e.g., traffic signs, traffic lights, etc.) may indicate a visual attack that needs to be evaluated. When determining whether there are changes between past and current trajectories, differences in distance or viewing angle can be considered.
[0088] like Figure 8 As illustrated, past history checks can be supported by adding past historical values or labels to features and objects stored in frame trajectory segment pool data structures 802, 808. For example, a past historical value at time t in trajectory segment pool 802 of frame t can be associated with past historical trajectory segment pool 804 via association function 806, and in response to confirming the association of a specific object or feature, the past historical value or label can be updated in the trajectory segment pool of the next frame (e.g., frame (t+1)). As illustrated, such association and updating can be performed on a per-object / feature and per-trajectory basis.
[0089] Figures 9 to 11 This is an illustration of an alternative decision-making algorithm, based on several implementation schemes, for identifying potential attacks on vehicle cameras using multiple detection methods. (Reference) Figures 1A to 11 Inconsistencies identified in each of the temporal inconsistency counters and past history checks described above can be used for decision tests or algorithms executed in the processing system. Decision tests or algorithms may depend on desired or appropriate sensitivity levels and the tolerance or acceptability of false positive decisions. In the description of these figures, the detector may be a detection criterion (e.g., a function) that depends on the counter value and its threshold. In each decision test or algorithm, some detectors may be deactivated or ignored depending on the operating environment, such as during severe weather (e.g., rain, fog, or snow), when images may be distorted or affected by natural phenomena.
[0090] like Figure 9As illustrated, a simple decision test 900 could be: if any of the multiple detection methods or detectors 902 indicates an actual or possible attack (e.g., outputs indicating an actual or possible attack), then a visual attack is identified or suspected. As illustrated, the outputs or conclusions of the multiple detectors 902 are accumulated in an OR function 904 such that any positive detector output (e.g., a detector detecting image inconsistency) leads to a decision that a visual attack exists. Therefore, a single triggered detector (i.e., a single detector detecting image inconsistency) is sufficient to make a decision that an attack exists or is likely to exist.
[0091] like Figure 10 As illustrated, in some decision tests or algorithms 1000, the OR function may be applied to some detectors 1002, but not all detectors 902, such as those that are more reliable (e.g., have a lower false positive rate) or associated with more dangerous attacks, while the outputs of other detectors may be evaluated based on voting or the AND function 1006.
[0092] like Figure 11 As illustrated, in some implementations, a decision test or algorithmic identification of the presence of a visual attack can be determined based on the votes of all active detectors 902 (such as using a majority aggregator 1108). If a majority (or unanimous) vote is achieved (i.e., most detectors detect image inconsistencies) (e.g., in determination box 1110), a visual attack can be detected; otherwise, it can be determined that no attack exists. Figure 11 It is also illustrated that the output of each detector can be weighted or adjusted by weighting factors 1102, 1104, 1106 associated with or suitable for each type of detector 902 or detection method. In this embodiment, the majority aggregator 1108 can aggregate weighted scores, and the determination of the majority in the determination box 1110 can be whether the sum exceeds a threshold (i.e., the weighted sum of detectors that detect image inconsistencies).
[0093] Figure 12 This is a process flowchart of example method 1200 according to various implementation schemes, a method for verifying object detection and image inconsistencies, performed by a processing system on a device (e.g., a vehicle), for detecting and reacting to potential attacks on the device's camera system. Reference Figures 1A to 12The operation of method 1200 may be performed by a processing system (e.g., 102, 120, 240) comprising one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130) and / or hardware elements, any one or a combination thereof, which may be configured to perform any operation of method 12. Furthermore, one or more processors within the processing system may be configured with software or firmware to perform various operations of the method. To encompass any of the processors, hardware elements, and software elements that may be involved in performing method 1200, the element performing the method operation is referred to as the "processing system". Additionally, components for performing the functions of method 1200 may include the processing system (e.g., 102, 120, 240), which includes one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130), memory 112, radio module 118, and / or vehicle sensors (e.g., 122, 136).
[0094] In block 1202, the processing system may perform operations including receiving multiple images (such as, but not limited to, camera image frame streams) from one or more cameras of a device (e.g., a vehicle). For example, multiple image frames may be received from one or more forward-facing cameras used for observing the road ahead for navigation and collision avoidance.
[0095] In box 1204, the processing system can perform operations including performing image processing to identify and measure objects required for navigation and device operation. As described, multiple camera images can be processed by multiple different processing systems, including trained neural network processing systems, to extract information necessary for safe navigation of the device. As described, these operations can include identifying and measuring objects in the environment, including identifying objects as both objects and object types, identifying traffic signs and classifying or categorizing traffic signs, and detecting the contours and depth of roads. Image processing operations can also include identifying objects that need to be avoided, which can include classifying the identified objects. Further image processing that can be performed as part of the operations in box 1204 can include assigning lighting check labels, determining lighting or lighting values, performing semantic consistency checks on identified objects, performing depth plausibility analysis and labeling such objects, performing contextual consistency analysis and providing labels accordingly, and checking the consistency of object labels. Such operations can also include associating various object measurements with the current vehicle trajectory and object trajectories (i.e., the correlation between measurements and trajectories).
[0096] In box 1206, the processing system can perform operations including multiple different processes performed on multiple images to detect different types of image inconsistencies. Performing more than one different process for identifying image inconsistencies provides better detection sensitivity and reduces false alarm events. (See reference...) Figure 13The described processes may include temporal consistency checks on detected objects and their classifications, inconsistency counter checks to identify when the number of image inconsistencies exceeds a threshold indicating a potential attack on a vehicle camera, and past history checks to identify when commonly observed objects suddenly move, change, or disappear.
[0097] In box 1208, the processing system may perform operations including identifying visual attacks using the results of multiple different processes on multiple images. In box 1208, the results of the various checks performed in box 1206 can be used in a decision algorithm to identify whether an attack on a vehicle camera is occurring or is likely to occur. As described, such a decision algorithm can be as simple as identifying a visual attack if any one of the different inconsistency checking processes indicates the likelihood of contact. More complex algorithms may include assigning weights to each of the various inconsistency detection methods (referred to herein as "detectors") and accumulating the results in a voting or thresholding algorithm to determine whether a visual attack is more likely.
[0098] In determination box 1210, the processing system may perform operations including determining whether a visual attack has been identified or is likely to be identified based on the analysis performed in boxes 1206 and 1208.
[0099] In response to determining that a visual attack has been identified or may have occurred (i.e., determination box 1210 = "Yes"), the processing system may perform actions including performing one or more mitigation actions in box 1212. In some embodiments, mitigation actions may include ignoring, eliminating, or removing malicious tracks (or potentially malicious tracks or objects) from the tracking database by the vehicle ADS, or disabling or otherwise ignoring malicious features associated with objects identified in an image. In some embodiments, mitigation actions may include reporting the detected attack to a remote system, such as law enforcement or highway maintenance organizations, enabling the cessation or removal of the threat or cause of the malicious attack. In some embodiments, mitigation actions may include outputting an indication of a visual attack, such as a warning or notification to the operator.
[0100] The operations in method 1200 can be performed continuously. Therefore, in response to determining that an attack has not yet been identified in one or more cameras (i.e., determining box 1210 = "No") and after execution 1212, the processing system can repeat method 1200 by receiving multiple images again from one or more device cameras in box 1202 and performing the methods as described.
[0101] Figure 13This is a process flowchart of example operation 1206 according to various implementation schemes. This example operation can be performed as part of method 1200 for performing multiple different processes on multiple images to detect different types of image inconsistencies. Reference Figures 1A to 13 Operation 1206 may be performed by a processing system (e.g., 102, 120, 240) comprising one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130) and / or hardware elements, any one or a combination thereof being configured to perform any operation within the operation. Furthermore, one or more processors within the processing system may be configured with software or firmware to perform various operations. To encompass any of the processors, hardware elements, and software elements that may be involved in performing the illustrated operations, the element performing the method operation is referred to as the "processing system". Additionally, components for the functionality of performing the illustrated operations may include the processing system (e.g., 102, 120, 240), which includes one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130), memory 112, and / or a vehicle camera (e.g., 122, 136).
[0102] After image processing of multiple images has been performed in block 1204 of method 1200, the processing system can perform operations including performing a temporal consistency check on multiple images spanning a time period in block 1302. As described, the operation in block 1302 may include image frame-to-frame comparisons configured to detect temporal inconsistencies in values and / or classifications between two consecutive or sequential frames based on objects, features, or trajectories. Measurements performed on each image frame may be associated with objects, labels, and / or values in previous image frames, enabling the tracking of features, objects, labels, classifications, and values from one frame to the next. The process in block 1302 enables the identification of inconsistencies from one frame to another across various trajectories through an extended sequence of image frames. In some embodiments, the process in block 1302 may include tracking inconsistency outputs from a projector detector and inconsistency outputs from a patch detector.
[0103] In block 1304, the processing system may perform operations including performing an inconsistency counter check on multiple images, which determines whether the number of inconsistencies in the image frames meets a threshold. In some embodiments, the process in block 1304 may include counting image inconsistency events and determining whether the count for a given inconsistency exceeds a threshold that may indicate a visual attack. This counting check performed in block 1304 may be applied to each feature and object identified and tracked in the temporal consistency check in block 1302.
[0104] In block 1306, the processing system may perform operations including performing a past history check on multiple images, which compares previously identified objects in previously processed images with identified objects in currently acquired images to identify changes in at least one of the following: objects, object location, or object classification. In some embodiments, the process in block 1306 may include determining whether there is an unintended change between objects or features that have been observed in the past (e.g., when a vehicle has traveled the same route in previous days) and the observed objects or features. Unintended changes may be the movement or disappearance of objects with permanent classifications (such as traffic signs or road features), or the appearance of objects with permanent classifications that were not previously present.
[0105] The following paragraphs describe specific implementation examples. While some of the following specific implementation examples are described as example systems and methods, further example implementations may include: example operations discussed in the following paragraphs that can be implemented by various computing devices; example methods discussed in the following paragraphs implemented by a device computing device including a processing system comprising one or more processors configured with processor-executable instructions to perform the operations of the methods of the following specific implementation examples; example methods discussed in the following paragraphs implemented by a device computing device including components for performing the methods of the following specific implementation examples; and example methods discussed in the following paragraphs that can be implemented as a non-transitory processor-readable storage medium storing processor-executable instructions configured to cause the processing system of the device computing device to perform the operations of the methods of the following specific implementation examples.
[0106] Example 1. A method for detecting a visual attack performed by a processing system on a device, the method comprising: receiving a plurality of images from one or more cameras of the device; performing a plurality of different processes on the plurality of images to detect different types of image inconsistencies; using the results of the plurality of different processes on the plurality of images to identify a visual attack; and performing one or more mitigation actions in response to identifying the visual attack.
[0107] Example 2. The method according to Example 1, wherein performing the plurality of different processes on the plurality of images to detect different types of image inconsistencies includes performing a temporal consistency check on images spanning a time period.
[0108] Example 3. The method according to any one of Examples 1 to 2, wherein performing the plurality of different processes on the plurality of images to detect different types of image inconsistencies includes performing an inconsistency counter check on the plurality of images, the inconsistency counter check determining whether the number of inconsistencies in the images meets a threshold.
[0109] Example 4. The method according to any one of Examples 1 to 3, wherein performing the plurality of different processes on the plurality of images to detect different types of image inconsistencies includes performing a past history check on the plurality of images, the past history check comparing previously identified objects in previously processed images with identified objects in currently acquired images to identify changes in at least one of the objects, the location of the objects, or the classification of the objects.
[0110] Example 5. The method according to any one of Examples 1 to 4, wherein the result of using the plurality of different processes on the plurality of images includes one or more of the following: if any of the different types of image inconsistencies are detected, a visual attack is identified; if the number of detected image inconsistencies of the different types exceeds a threshold, a visual attack is identified; if a majority of the detectors detect image inconsistencies, a visual attack is identified; or if a weighted majority of the detectors detect image inconsistencies, a visual attack is identified, wherein the weights applied to each of the detectors are predetermined.
[0111] Example 6. The method according to any one of Examples 1 to 5, wherein the mitigation action performed in response to the identification of a visual attack includes one or more of the following: removing a malicious trajectory from a tracking database, outputting an indication of the visual attack, or disabling malicious features associated with an object identified in the camera image.
[0112] Example 7. The method according to any one of Examples 1 to 6, wherein performing mitigation actions in response to identifying a visual attack includes reporting the detected attack to a remote system.
[0113] Example 8. The method according to any one of Examples 1 to 7, wherein the device is a vehicle.
[0114] As used in this application, the terms "component," "module," "system," etc., are intended to include computer-related entities such as, but not limited to, hardware, firmware, combinations of hardware and software, software, or software being executed, configured to perform specific operations or functions. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a wireless device and the wireless device itself can be referred to as a component. One or more components may reside within a process and / or a thread of execution, and components may reside on a processor or core and / or be distributed across two or more processors or cores. Furthermore, these components may execute on various non-transitory computer-readable media on which various instructions and / or data structures are stored. Components may communicate via local and / or remote processes, function or procedure calls, electronic signals, data packets, memory read / write, and other known network, computer, processor, and / or process-related communication methods.
[0115] Several different cellular and mobile communication services and standards are available and envisioned for the future, all of which are feasible and benefit from various implementation schemes for reporting the detection of visual attacks on devices. These services and standards include, for example, the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE) systems, 3rd Generation Wireless (3G), 4th Generation Wireless (4G), 5th Generation Wireless (5G), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), 3GSM, General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA) systems (e.g., cdmaOne, CDMA1020TM), Enhanced Data Rate Evolution of GSM (EDGE), Advanced Mobile Phone Systems (AMPS), Digital AMPS (IS-136 / TDMA), Evolved Data Optimization (EV-DO), Digital Enhanced Cordless Telecommunications (DECT), Global Interoperability for Microwave Access (WiMAX), Wireless Local Area Networks (WLAN), Wi-Fi Protected Access I and II (WPA, WPA2), and Integrated Digital Enhanced Network (iDEN). Each of these technologies relates to the transmission and reception of, for example, voice, data, signaling, and / or content messages. It should be understood that any references to terms and / or technical details relating to individual telecommunications standards or technologies are for illustrative purposes only and are not intended to limit the scope of the claims to a particular communication system or technology, unless specifically stated in the language of the claims.
[0116] The various embodiments illustrated and described are provided merely as examples illustrating the various features of the claims. However, the features shown and described with respect to any given embodiment are not necessarily limited to the associated embodiment and may be used or combined with other embodiments shown and described. Furthermore, the claims are not intended to be limited to any one of the exemplary embodiments. For example, one or more of methods and operations 400a-400c may replace or be combined with one or more operations of methods and operations 400a-400c.
[0117] The foregoing method descriptions and process flowcharts are provided as illustrative examples only and are not intended to require or imply that the operations of the various embodiments must be performed in the given order. As those skilled in the art will appreciate, the operations in the foregoing embodiments can be performed in any order. Words such as “afterward,” “then,” “next,” etc., are not intended to limit the order of operations; these words are used to guide the reader through the description of the method. Furthermore, any reference to singular claim elements (e.g., references using the articles “a,” “an,” or “the”) should not be construed as limiting that element to the singular. Additionally, references to the term “and / or” should be understood to include both conjunctions and antonyms. For example, “A and / or B” means “A and B” as well as “A or B.”
[0118] The various exemplary logic blocks, modules, components, circuits, and algorithmic operations described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and operations have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the claims.
[0119] Hardware for implementing the various exemplary logic units, logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or executed using a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any combination thereof designed to perform the functions described herein. While the general-purpose processor may be a microprocessor, in alternative embodiments, the processing system may perform operations including as any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of receiver intelligent objects, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.
[0120] In one or more embodiments, the described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The operation of the methods or algorithms disclosed herein may be implemented in a processor-executable software module or processor-executable instructions, which may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium that can be accessed by a computer or processor. By way of example and without limitation, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EEPROM, FLASH memory, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage intelligent objects, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. The above combinations may also be included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operation of a method or algorithm may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable and / or computer-readable storage medium, which may be incorporated into a computer program product.
[0121] The above description of the disclosed embodiments is provided to enable any person skilled in the art to implement or use the claims. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the claims. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but should be granted the broadest scope consistent with the following claims and the principles and novel features disclosed herein.
Claims
1. A method for detecting visual attacks performed by a processing system on a device, the method comprising: Receive multiple images from one or more cameras of the device; Perform multiple different processes on the multiple images to detect different types of image inconsistencies; Visual attacks are identified using the results of the various processes applied to the multiple images. as well as One or more mitigation actions are performed in response to the identification of the visual attack.
2. The method of claim 1, wherein performing the plurality of different processes on the plurality of images to detect different types of image inconsistencies includes performing a temporal consistency check on images spanning a time period.
3. The method of claim 1, wherein performing the plurality of different processes on the plurality of images to detect different types of image inconsistencies includes performing an inconsistency counter check on the plurality of images, the inconsistency counter check determining whether the number of inconsistencies in the images meets a threshold.
4. The method of claim 1, wherein performing the plurality of different processes on the plurality of images to detect different types of image inconsistencies includes performing a past history check on the plurality of images, the past history check comparing previously identified objects in previously processed images with identified objects in currently acquired images to identify changes in at least one of the objects, the location of the objects, or the classification of the objects.
5. The method of claim 1, wherein the results of the plurality of different processes applied to the plurality of images include one or more of the following: If any of the different types of image inconsistencies is detected, a visual attack is identified; If the number of detected image inconsistencies of the different types exceeds a threshold, a visual attack is identified; If most detectors detect image inconsistencies, a visual attack is identified; or A visual attack is identified if a weighted majority of the detectors in the detectors detect image inconsistencies, wherein the weights applied to each of the detectors are predetermined.
6. The method of claim 1, wherein the mitigation action performed in response to the identification of a visual attack includes one or more of the following: removing a malicious trajectory from a tracking database, outputting an indication of the visual attack, or disabling malicious features associated with an object identified in the camera image.
7. The method of claim 1, wherein performing mitigation actions in response to identifying a visual attack includes reporting the detected attack to a remote system.
8. The method according to claim 1, wherein the device is a means of transportation.
9. An apparatus comprising: One or more memory units; One or more cameras; as well as A processing system coupled to the one or more memories and the one or more cameras, and including one or more processors configured to: Receive multiple images from one or more cameras of the device; Perform multiple different processes on the image to detect different types of image inconsistencies; Visual attacks are identified using the results of the various processes applied to the multiple images. as well as In response to the identification of a visual attack, one or more mitigation actions are performed.
10. The apparatus of claim 9, wherein the one or more processors are further configured to perform a temporal consistency check on camera images spanning a time period.
11. The apparatus of claim 9, wherein the one or more processors are further configured to perform an inconsistency counter check on a camera image, the inconsistency counter check determining whether the number of inconsistencies in the camera image meets a threshold.
12. The apparatus of claim 9, wherein the one or more processors are further configured to perform a past history check on the camera images, the past history check comparing previously identified objects in previously processed camera images with identified objects in currently acquired camera images to identify changes in at least one of the objects, the location of the objects, or the classification of the objects.
13. The apparatus of claim 9, wherein the one or more processors are further configured to use the results of the plurality of different processes on the plurality of images to: If any one of the different detectors detects an image inconsistency, a visual attack is identified; If the number of detected image inconsistencies of the different types exceeds a threshold, a visual attack is identified; If inconsistencies are detected in most of the aforementioned different types of images, a visual attack is identified; or A visual attack is identified if a weighted majority of the different detectors detects image inconsistencies, wherein the weights applied to each of the different detectors are predetermined.
14. The apparatus of claim 9, wherein the one or more processors are further configured to perform mitigation actions in response to the identification of a visual attack, the mitigation actions including removing malicious tracks from a tracking database or disabling one or more malicious features associated with an object identified in the camera image.
15. The apparatus of claim 9, wherein the one or more processors are further configured to perform mitigation actions in response to the identification of a visual attack, the mitigation actions including reporting the detected attack to a remote system.
16. The apparatus of claim 9, wherein the apparatus is a means of transportation.
17. A non-transitory processor-readable medium storing processor-executable instructions configured to cause a device's processing system to: Receive multiple images from one or more cameras of the device; Perform multiple different processes on the multiple images to detect different types of image inconsistencies; Visual attacks are identified using the results of the various processes applied to the multiple images. as well as In response to the identification of a visual attack, one or more mitigation actions are performed.
18. The non-transitory processor-readable medium of claim 17, wherein the stored processor-executable instructions are further configured to cause the processing system of the device to perform a temporal consistency check on camera images spanning a time period.
19. The non-transitory processor-readable medium of claim 17, wherein the stored processor-executable instructions are further configured to cause the processing system of the device to perform an inconsistency counter check on a camera image, the inconsistency counter check determining whether the number of inconsistencies in the camera image meets a threshold.
20. The non-transitory processor-readable medium of claim 17, wherein the stored processor-executable instructions are further configured to cause the processing system of the device to perform a past history check on a camera image, the past history check comparing previously identified objects in a previously processed image with identified objects in a currently acquired image to identify changes in at least one of the objects, the location of the objects, or the classification of the objects.
21. The non-transitory processor-readable medium of claim 17, wherein the stored processor-executable instructions are further configured to cause the processing system of the device to use the results of the plurality of different processes on the camera image to: If any of the different types of image inconsistencies is detected, a visual attack is identified; If the number of detected image inconsistencies of the different types exceeds a threshold, a visual attack is identified; If most of the detectors detect image inconsistencies, a visual attack is identified; or A visual attack is identified if a weighted majority of the detectors in the detectors detect image inconsistencies, wherein the weights applied to each detector are predetermined.
22. The non-transitory processor-readable medium of claim 17, wherein the stored processor-executable instructions are further configured to cause the processing system of the device to perform mitigation actions in response to recognizing a visual attack, the mitigation actions including one or more of the following: removing a malicious trajectory from a tracking database, outputting an indication of the visual attack, or disabling malicious features associated with an object identified in the camera image.
23. The non-transitory processor-readable medium of claim 17, wherein the stored processor-executable instructions are further configured to cause the processing system of the device to perform mitigation actions in response to recognizing a visual attack, the mitigation actions including reporting the detected attack to a remote system.