Reasonableness and consistency checker for vehicle device cameras
By using multiple image processing models and consistency checks to identify visual attacks in autonomous vehicles, the problem of image processing inconsistency in camera systems is solved, improving the system's security and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-11-20
- Publication Date
- 2026-06-26
AI Technical Summary
Camera systems in autonomous vehicles are vulnerable to malicious visual attacks, which can lead to inconsistent image processing and compromise safe operation.
Multiple trained image processing models are used to process camera images. Through semantic segmentation, depth estimation, object detection and object classification, consistency checks are performed to identify inconsistencies in the image processing output, and mitigation actions are performed when an attack is detected.
Effectively identify and mitigate visual attacks, improve the safety and reliability of autonomous driving systems, and ensure the accuracy and consistency of image processing.
Smart Images

Figure CN122295703A_ABST
Abstract
Description
Background Technology
[0001] With the emergence of autonomous and semi-autonomous vehicles, robotic vehicles, and other types of mobile devices using Advanced Driver Assistance Systems (ADAS) and Automatic Automated Driving Systems (ADS), devices with such systems are becoming vulnerable to new forms of malicious behavior and threats; namely, spoofing or otherwise attacking the camera systems at the heart of autonomous vehicles' navigation and object avoidance. While such attacks may be rare now, they are expected to become a significant problem in the future as devices with autonomous driving systems proliferate. Summary of the Invention
[0002] Various aspects include methods that can be implemented on the processing system of the device, and systems for implementing methods for assessing the plausibility and / or consistency of cameras and advanced driver assistance system (ADAS) cameras used in inspection (ADS) to identify potential malicious attacks. These aspects may include: using multiple trained image processing models to process images received from the device's cameras to obtain multiple image processing outputs; performing multiple consistency checks on the multiple image processing outputs, wherein a consistency check in the multiple consistency checks compares each of the multiple different outputs to detect inconsistencies; detecting attacks on the cameras based on the inconsistencies; and performing mitigation actions in response to the identification of an attack.
[0003] In some aspects, using multiple trained image processing models to process images received from the camera of the device to obtain multiple image processing outputs may include: performing semantic segmentation processing on the image using a trained semantic segmentation model to associate masks of pixel groups in the image with classification labels; performing depth estimation processing on the image using a trained depth estimation model to identify distances to objects in the image; performing object detection processing on the image using a trained object detection model to identify objects in the image and define bounding boxes around the identified objects; and performing object classification processing on the image using a trained object classification model to classify the objects in the image.
[0004] In some aspects, performing multiple consistency checks on multiple image processing outputs may include: performing a semantic consistency check that compares the classification labels associated with the mask from the semantic segmentation process with the bounding boxes of object detections from the object detection process in the image to identify inconsistencies between the mask classification and the detected objects; and providing an indication of detected classification inconsistencies in response to the inconsistencies between the mask classification and the detected objects in the image.
[0005] Some aspects may also include: in response to the consistency between the classification label associated with the mask from the semantic segmentation process and the bounding box of the object detection from the object detection process, performing a positional consistency check that compares the position of the classification mask from the semantic segmentation process in the image with the position of the bounding box of the object detection from the object detection process in the image to identify inconsistencies between the position of the classification mask and the position of the detected object bounding box, and providing an indication of the detected classification inconsistency if the position of the classification mask in the image is inconsistent with the position of the detected object bounding box.
[0006] In some aspects, performing multiple consistency checks on multiple image processing outputs may include: performing a depth reasonableness check that compares the depth estimate of a detected object from an object detection process with the depth estimates of individual pixels or groups of pixels from a depth estimation process to identify a distribution in the depth estimates across pixels of the detected object that is inconsistent with a depth distribution associated with a classification of a mask covering the detected object from a semantic classification process; and providing an indication of detected depth inconsistencies in cases where the distribution in the depth estimates across pixels of the detected object differs from the depth distribution associated with the classification of the mask.
[0007] In some aspects, performing multiple consistency checks on multiple image processing outputs may include: performing a contextual consistency check that compares depth estimates of bounding boxes covering detected objects from object detection processing with depth estimates of masks covering detected objects from semantic segmentation processing to determine whether the distribution of the mask depth estimates differs from the depth estimates of the bounding boxes; and providing an indication of detected contextual inconsistencies if the distribution of the mask depth estimates is the same as or similar to the distribution of the bounding box depth estimates.
[0008] In some aspects, performing multiple consistency checks on multiple image processing outputs may include: performing a label consistency check that compares detected objects from object detection processing with the labels of the detected objects from object classification processing to determine whether the object classification labels are consistent with the detected objects; and providing an indication of the detected label inconsistency if the object classification labels are inconsistent with the detected objects.
[0009] In some aspects, performing mitigation actions in response to identifying an attack may include adding indications of inconsistencies from each of a plurality of consistency checks to information about each detected object, which is provided to the autonomous driving system for tracking the detected object. In some aspects, performing mitigation actions in response to identifying an attack may include reporting the detected attack to a remote system.
[0010] Another aspect includes devices such as vehicles, which include memory and a processor configured to perform operations of any of the methods outlined above. Another aspect may include devices such as vehicles having various components for performing functions corresponding to any of the methods outlined above. Another aspect may include a non-transitory processor-readable storage medium having processor-executable instructions stored thereon, the processor-executable instructions being configured to cause one or more processors of the device processing system to perform various operations corresponding to any of the methods outlined above. Attached Figure Description
[0011] The accompanying drawings, incorporated herein and forming part of this specification, illustrate exemplary embodiments of the claims and, together with the general description given above and the detailed description given below, serve to interpret the features of the claims.
[0012] Figures 1A to 1C This is a component block diagram illustrating a typical system of automated devices in the form of vehicles suitable for implementing various schemes.
[0013] Figure 2 It is a functional block diagram showing the functional elements or modules of an autonomous driving system suitable for implementing various implementation schemes.
[0014] Figure 3 It is a component block diagram applicable to processing systems that implement various implementation schemes.
[0015] Figure 4A and Figure 4B This is a block diagram illustrating the various operations performed on multiple images as part of a typical autonomous driving system.
[0016] Figure 5A and Figure 5B This is a processing block diagram illustrating various operations performed on multiple images according to various embodiments, which can be performed as part of an autonomous driving system, including operations for identifying inconsistencies in the image processing results, which can indicate visual attacks on the device's cameras.
[0017] Figure 6 This is a process flowchart of an example method according to various implementation schemes for detecting visual attacks performed by a processing system on a device (e.g., a vehicle) for detecting and responding to potential attacks on the device's camera system.
[0018] Figure 7It is a process flowchart according to some embodiments of an image processing method that can be performed on images from a camera of a device to support ADS or ADAS, the output of which can be processed to identify inconsistencies that may indicate a visual attack or potential visual attack.
[0019] Figures 8A to 8D This is a flowchart illustrating a method for identifying inconsistencies in the processing of images from a camera used to identify visual attacks or potential visual attacks, according to some implementation schemes. Detailed Implementation
[0020] Various embodiments will be described in detail with reference to the accompanying drawings. Where possible, the same reference numerals will be used throughout the drawings to refer to the same or similar parts. References to specific examples and embodiments are for illustrative purposes and are not intended to limit the scope of the claims.
[0021] Various implementations include methods and vehicle processing systems for processing individual images to identify and respond to attacks (referred to herein as "visual attacks") on a device's (e.g., a vehicle's) camera. These implementations address potential risks to devices (e.g., vehicles) that may be caused by malicious visual attacks and unintentional actions that make images acquired by the camera appear to include false objects or obstacles to be avoided, forged traffic signs, images that may interfere with depth and distance determination, and similarly misleading images that may interfere with the safe autonomous driving operation of the device. Various implementations provide methods for identifying actual or potential visual attacks based on inconsistencies in individual images, including semantic classification inconsistencies, semantic classification location inconsistencies, depth plausibility inconsistencies, contextual inconsistencies, and label inconsistencies. When a visual attack or potential attack is identified, some implementations include a processing system performing one or more mitigation actions to resolve or adapt to a visual attack on a camera in ADS or ADAS operation, and / or reporting the detected attack to an external third party, such as law enforcement or highway maintenance agencies, to stop the attack.
[0022] Various implementation schemes can improve the operational safety of automated and semi-automated devices (e.g., vehicles) by providing effective methods and systems for detecting malicious attacks on camera systems and taking mitigation actions such as reducing the risk to vehicles, providing output instructions, and / or reporting attacks to appropriate authorities.
[0023] The terms “onboard” or “within a vehicle” are used interchangeably herein to refer to equipment or components contained within, attached to, and / or carried by a device (e.g., a vehicle or equipment providing the functionality of a vehicle). Onboard equipment typically includes a processing system, which may include one or more processors, a System-on-a-Chip (SoC), and / or a System-on-a-Package (SIP), any of which may include one or more components, systems, units, and / or modules that implement functionality (collectively referred to herein as “processing system” for simplicity). Onboard equipment and aspects of functionality may be implemented in hardware components, software components, or a combination of hardware and software components.
[0024] The term "System-on-a-Chip" (SOC) is used herein to refer to a single integrated circuit (IC) chip containing multiple resources and / or processors integrated on a single substrate. A single SOC may contain circuitry for digital, analog, mixed-signal, and radio frequency functions. A single SOC may also include any number of general-purpose and / or special-purpose processors (digital signal processors, modem processors, video processors, etc.), memory blocks (e.g., ROM, RAM, flash memory, etc.), and resources (e.g., timers, voltage regulators, oscillators, etc.). A SOC may also include software for controlling the integrated resources and processors, as well as software for controlling peripheral devices.
[0025] The term "System-in-Package" (SIP) may be used herein to refer to a single module or package containing multiple resources, computing units, cores and / or processors on two or more IC chips, a substrate, or a System-on-a-Chip (SoC). For example, a SIP may include a single substrate on which multiple IC chips or semiconductor dies are stacked in a vertical configuration. Similarly, a SIP may include one or more multi-chip modules (MCMs) on which multiple ICs or semiconductor dies are packaged into a single substrate. A SIP may also include multiple independent SoCs coupled together and packaged in close proximity via high-speed communication circuitry, such as on a single motherboard or in a single wireless device. The proximity of the SoCs facilitates high-speed communication and the sharing of memory and resources.
[0026] The term "device" is used herein to refer to any of a variety of devices, systems, and equipment that may use camera vision systems and are therefore potentially vulnerable to visual attacks. Some non-limiting examples of devices to which various implementations may be applied include autonomous and semi-autonomous vehicles, mobile robots, mobile machinery, autonomous and semi-autonomous farm equipment, autonomous and semi-autonomous construction and paving equipment, autonomous and semi-autonomous military equipment, etc.
[0027] As used herein, the term "processing system" refers to one or more processors, including multi-core processors, that are organized and configured to perform various computational functions. Various implementation methods may be implemented in one or more of a plurality of processors within any of the various vehicle computers and processing systems described herein.
[0028] As used herein, the term "semantic segmentation" encompasses image processing, such as associating individual pixels or groups of pixels in a digital image with classification labels, such as "tree," "traffic sign," "pedestrian," "road," "building," "car," "sky," etc., via a trained model. The coordinates of a group of pixels can be in the form of a "mask" associated with a classification label within the image, where the mask is defined by coordinates within the image (e.g., pixel coordinates) or by coordinates and regions within the image.
[0029] Camera systems and image processing play a crucial role in current and future automated and semi-automated devices, such as ADS or ADAS systems implemented in automated and semi-automated vehicles, mobile robots, mobile machinery, and automated and semi-automated farm equipment. In such devices, multiple cameras provide images of the road and surrounding landscape, thus providing data available for navigation (e.g., road following), object recognition, collision avoidance, and hazard detection. The processing of image data in modern ADS or ADAS systems has advanced far beyond basic object recognition and tracking, including understanding information posted on street signs, understanding road conditions, and navigating complex road situations (e.g., turning lanes, avoiding pedestrians and cyclists, maneuvering around traffic cones, etc.).
[0030] Processing camera data fields involves several tasks (sometimes referred to as "visual tasks") that are crucial for the safe operation of automated devices such as vehicles. Visual tasks, typically performed by camera systems to support ADS and ADAS operations, include semantic segmentation, depth estimation, object detection, and object classification. These image processing operations are critical to supporting basic navigation ADS / ADAS operations, including road following with depth estimation for path planning, object detection in three dimensions (3D), object identification or classification, traffic sign recognition (including temporary traffic signs and signs reflected in map data), and panoramic segmentation.
[0031] In modern ADS and ADAS systems, camera images can be processed by multiple different analysis engines, sometimes referred to as a "visual pipeline." To identify and understand the scene around a device (e.g., a vehicle), the multiple different analysis engines in the visual pipeline are typically neural network-type artificial intelligence / machine learning (AI / ML) modules. These models are trained to perform different analytical tasks on image data and output specific types of information. For example, such trained AI / ML analysis modules in a visual pipeline may include models trained to perform semantic segmentation analysis on individual images, models trained to perform depth estimation of pixels, pixel groups, and regions / bounding boxes for objects within an image, models trained to perform object detection (i.e., detecting objects within an image), and models trained to perform object classification (i.e., determining and assigning categories to detected objects). Such trained AI / ML analysis modules can analyze image frames and image sequences to identify and interpret objects in real time. The information outputs of these image processing training models can be combined to generate informational data structures that use camera images (e.g., in a tracked object data structure) to identify and track objects. These camera images can be used by the device's ADS or ADAS processor to support navigation, collision avoidance, and following traffic processes (e.g., traffic signs or signals).
[0032] A key operation achieved by processing image data in the vision pipeline is object detection and classification (i.e., identifying and understanding the meaning or implications of objects). In addition to object detection, the position of detected objects relative to the device in three dimensions (3D) is also important for navigation and collision avoidance. Examples of objects that ADS and ADAS operations need to identify, classify, and in some cases interpret or understand include traffic signs, pedestrians, other vehicles, road obstacles, road boundaries and lane markings, and road features that differ from information included in detailed map data and observed during previous driving experiences.
[0033] Traffic signs are objects that need to be identified, classified, and processed to understand the text displayed in applications of autonomous vehicles (e.g., speed limits). This processing is necessary so that the guidance and regulations identified by the signs can be incorporated into the decisions of the autonomous driving system. Typically, traffic signs have recognizable shapes depending on the type of information displayed (e.g., stop, yield, speed limit, etc.). However, sometimes the information displayed differs from the meaning or classification corresponding to the shape, such as text in different languages or shapes that are not actually traffic signs (e.g., advertisements, T-shirt designs, protest signs, etc.). Furthermore, traffic signs may identify requirements or regulations (e.g., speed limits or traffic controls) that are inconsistent with information appearing in map data that ADS or ADAS may rely on.
[0034] Pedestrians and other vehicles are important targets for detection, classification, and close tracking to avoid collisions and properly plan vehicle paths. Classifying pedestrians and other vehicles can help predict their future location or trajectory, which is crucial for future planning performed by autonomous driving systems.
[0035] In addition to identifying, classifying, and obtaining information about detected objects, image data can be processed in a way that allows the position of these objects to be tracked frame by frame, making it possible to determine the trajectory of the object relative to the device (or the device relative to the object) to support navigation and collision avoidance functions.
[0036] Visual attacks, and the obfuscated or conflicting images that could mislead the image analysis process of autonomous driving systems, can originate from many different sources and involve a variety of different types of attacks. Visual attacks can target semantic segmentation operations, depth estimation, and / or object detection and recognition functions that are critical image processing capabilities of ADS or ADAS systems. Visual attacks can include projector attacks and patching attacks.
[0037] In projector vision attacks, images are projected onto a vehicle camera by a projector with the aim of creating false or misleading image data to interfere with Adaptive Dashboards (ADS) or Advanced Driver Assistance Systems (ADAS). For example, a projector can be used to project an image onto a road so that, when viewed in the camera's two-dimensional visual plane, the image appears three-dimensional and resembles an object to be avoided. An example of this type of attack would be a projection of a picture or shape resembling a pedestrian (or other object) onto the road, which, when viewed from the perspective of a vehicle camera, appears to be a pedestrian in the road. Another example is a projector that projects images onto structures along the road, such as projecting an image of a stop sign onto a building wall that would otherwise be blank. Yet another example is a projector that directly targets a device camera, injecting an image (e.g., an incorrect traffic sign) into the image.
[0038] Examples of patched visual attacks include images of objects that can be identified (such as traffic signs) but are fake, inappropriate, or placed where such objects shouldn't be. For example, a T-shirt with an image of a stop sign could interfere with an autonomous driving system, making it difficult to determine whether a vehicle should stop or ignore the sign, especially if the person wearing the shirt is walking or running rather than at or near an intersection. As another example, images of the rear of a vehicle or interfering shapes could interfere with image processing modules that estimate the depth and 3D localization of objects.
[0039] While some methods have been proposed to address image distortion and interference, a comprehensive multi-factor approach has not been identified. Therefore, camera-based ADS or ADAS operations remain vulnerable to many visual attacks.
[0040] Various implementations offer an integrated security solution to address threats posed by attacks on device cameras that support autonomous driving and steering systems, based on analysis of individual images from device cameras. These implementations include the use of various types of consistency checks (sometimes called detectors) that identify inconsistencies in the outputs of different image processing processes as part of ADS / ADAS image analysis and object tracking processes. As used herein, the term "image processing" refers to computational and neural network processing performed by a device (such as a vehicle ADS or ADAS system) on device camera images to produce data (generally referred to herein as image processing "outputs") that provides information in the format required for the device system's object detection, collision avoidance, navigation, and other functions. Examples of image processing covered in this term may include various types of processes that output different types of information, such as depth estimation of individual pixels and groups of pixels, object recognition bounding box coordinates, object recognition labels, etc. Consistency checkers compare two or more outputs of an image processing module or vision pipeline to identify differences in the outputs that reveal inconsistent analytical results or conclusions. Each of the consistency checkers or detectors compares the outputs of selected different camera vision pipelines to identify / recognize inconsistencies in the respective outputs. In doing so, the consistency checker system is able to identify visual attacks in a single image. Some example consistency checkers include depth plausibility checks, semantic consistency checks, positional inconsistency checks, contextual consistency checks, and label consistency checks; however, other implementations may use more or fewer consistency checkers, such as comparing the shape of detected objects with object classification and / or semantic segmentation mask labels.
[0041] In depth reasonableness checks, depth estimates from individual pixels or groups of pixels derived from depth estimation processes performed on the semantic segmentation mask and the identified objects are compared to determine whether the distribution of depth estimates across pixels of detected objects is consistent with or inconsistent with the depth distribution across the semantic segmentation mask. By estimating the depth to individual pixels or groups of pixels, the distribution of depth estimates for detected objects in a digital image can be obtained. For a single entity object (e.g., a vehicle, a pedestrian, etc.), the distribution of pixel depth estimates across the object should be narrow (i.e., a small fraction or percentage of depth estimates). Conversely, objects that are not entities (e.g., projections on a road, banners with holes in the middle, objects that appear to be vehicles with gaps passing through them, etc.) can exhibit a broad distribution of pixel depth estimates (i.e., some pixel depth estimates differ from the average depth estimate of the remaining pixels covering the detected object by more than a threshold fraction or percentage). By analyzing the pixel depth estimates of detected objects to identify when objects exhibit a distribution of depth estimates that exceeds a threshold difference, fraction, or percentage (i.e., depth estimate inconsistency), objects with illogical depth distributions can be identified. This may indicate that the detected object is not what it appears to be (e.g., a projection of a real object, a banner or sign showing an object that is not the actual object, etc.), and thus indicate a visual attack.
[0042] In semantic consistency checking, the output of the semantic segmentation process of an image is compared with the bounding boxes surrounding the detected objects from the object detection process to determine whether the labels assigned to the semantic segmentation masks are consistent with or inconsistent with the detected object bounding boxes. For example, the semantic segmentation process or visual pipeline can label each mask with category labels (e.g., "tree", "traffic sign", "pedestrian", "road", "building", "car", "sky", etc.), and the object detection / visual pipeline and / or object classification process can use a neural network AI / ML model to identify objects, which has been trained on a broad training dataset including images of objects that have already been assigned ground truth labels. In semantic consistency checking, mask labels from semantic segmentation that differ from or do not cover the labels assigned in the object detection / object classification process will be identified as inconsistent.
[0043] In location inconsistency checking, if semantic consistency checking finds that the mask label matches the detected object bounding box, location inconsistency checking can be performed if the semantic segmentation mask is located in a similar position within the image or overlaps within the bounding box of the detected object within a threshold amount. The mask and bounding box can have different sizes, so the region overlap ratio can be less than one. However, if the mask and bounding box appear in approximately the same position in the image, the region overlap ratio can be equal to or greater than a threshold overlap value, which is set to identify cases where the overlap between the mask and bounding box is insufficient to indicate that they belong to the same object. If the overlap ratio is less than the threshold overlap value, this indicates that the semantic segmentation mask is focused on something different from the detected object, and therefore there is a semantic location inconsistency that could indicate an actual or potential visual attack.
[0044] In context consistency checks, depth estimates of detected objects and the rest of the environment in the scene can be examined for inconsistencies that indicate erroneous images or deceptive objects. In some implementations, the inspector or detector can compare the estimated depth values of pixels covering the detected object or the mask encompassing it with the estimated depth values of pixels covering the overlapping mask. In some implementations, the inspector or detector can compare the distribution of estimated pixel depth values across the bounding box of the detected object or the object with the distribution of estimated pixel depth values across the overlapping mask, thereby comparing the difference with a threshold indicating an actual or potential visual attack or an inconsistency that allows for action.
[0045] In label consistency checks, detected objects from the object detection process are compared with the labels obtained from the object classification process to determine whether the object classification labels match the detected objects. Label inconsistencies can be identified when the labels assigned to the same object or mask by the two labeling processes (semantic segmentation and object detection / classification) do not match or fall into different unique categories (e.g., "tree" vs. "car" or "traffic sign" vs. "pedestrian").
[0046] In some implementations, the outputs of some or all of the different consistency checks may be numeric values, such as "1" or "0," to indicate whether an inconsistency was detected in an image. For example, a "0" may be output to indicate a genuinely detected object within the image, and a "1" may be output to indicate a non-genuinely detected object, malicious image data, a visual attack, or other indication of untrusted image data. In some implementations, the outputs of some or all of the different consistency checks may include further information about the detected inconsistencies, such as identifiers of the detected objects associated with the inconsistency, the pixel coordinates of each detected inconsistency within the image, the number of inconsistencies detected in a given image, and other types of information for identifying and tracking multiple inconsistencies detected in a given image.
[0047] The output of the inconsistency checks can then be used to determine whether a visual attack is occurring or may be occurring. In some implementations, the results of all inconsistency checks may be considered when determining whether a visual attack is occurring or may be occurring. In some implementations, the results of individual inconsistency checks can be used to determine whether different types of visual attacks are occurring or may be occurring.
[0048] Some implementations include performing one or more mitigation actions in response to determining that a visual attack is occurring or may be occurring. In some implementations, mitigation actions may involve appending information about the conclusions from various inconsistency checks to data fields provided to the ADS or ADAS object tracking information, enabling the system to determine how to react to the detected object. For example, information about an object tracked by the ADS or ADAS may include information about whether any of the multiple inconsistency checks indicates an attack or unreliable information, which can help the ADS / ADAS determine how to navigate relative to such an object. In some implementations, indications of detected inconsistencies in the image processing results may be reported to the operator. In some implementations, information indicating a visual attack determined based on one or more identified inconsistency results may be communicated to remote services, such as highway management, law enforcement, etc.
[0049] Various implementation schemes can be implemented in various devices, Figure 1A and Figure 1B The example provided is a non-limiting illustration of a means of transport, form 100. (See reference.) Figure 1A and Figure 1BThe vehicle 100 may include a control unit 140 and a plurality of sensors 102 to 138, including a satellite geolocation system receiver 108, occupancy sensors 112, 116, 118, 126, 128, tire pressure sensors 114, 120, cameras 122, 136, microphones 124, 134, an impact sensor 130, radar 132, and lidar 138. These sensors 102 to 138, located in or on the vehicle, can be used for various purposes, such as automatic and semi-automatic navigation and control, collision avoidance, location determination, etc., and to provide sensor data about objects and people in or on the vehicle 100. Sensors 102 to 138 may include one or more of a wide variety of sensors capable of detecting various information available for navigation, collision avoidance, and automatic and semi-automatic navigation and control. Each of the sensors 102 to 138 may communicate wirelessly with the control unit 140 and with each other. Specifically, the sensors may include one or more cameras 122, 136, or other optical or photoelectric sensors. Cameras 122, 136, or other optical or photoelectric sensors may include outward-facing sensors for imaging objects outside the vehicle 100 and / or in-vehicle sensors for imaging objects (including passengers) inside the vehicle 100. In some embodiments, the number of cameras may be less than two or more than two. For example, more than two cameras may be present, such as two front-facing cameras with different fields of view (FOV), four side-facing cameras, and two rear-facing cameras. Sensors may also include other types of object detection and ranging sensors, such as radar 132, lidar 138, IR sensors, and ultrasonic sensors. The sensors may also include tire pressure sensors 114, 120, humidity sensors, temperature sensors, satellite geolocation sensors 108, accelerometers, vibration sensors, gyroscopes, gravimeters, impact sensors 130, force gauges, pressure gauges, strain sensors, fluid sensors, chemical sensors, gas content analyzers, hazardous substance sensors, microphones 124, 134 (inside or outside the vehicle 100), occupancy sensors 112, 116, 118, 126, 128, proximity sensors, and other sensors.
[0050] The vehicle control unit 140 may be configured with processor-executable instructions to perform operations in some embodiments using information received from various sensors, particularly cameras 122 and 136. In some embodiments, the control unit 140 may supplement the processing of multiple images with distance and relative positioning (e.g., relative azimuth) obtainable from radar 132 and / or lidar 138 sensors. The control unit 140 may also be configured to control the steering, braking, and speed of the vehicle 100 when operating in automatic or semi-automatic mode using information about other vehicles determined using methods from some embodiments. In some embodiments, the control unit 140 may be configured to operate as an automated driving system (ADS). In some embodiments, the control unit 140 may be configured to operate as an automated driver assistance system (ADAS).
[0051] Figure 1C This is a component block diagram of system 150, illustrating components and supporting systems applicable to implementing some implementation schemes. (See reference) Figure 1A , Figure 1B and Figure 1C The vehicle 100 may include a control unit 140, which may include various circuits and devices for controlling the operation of the vehicle 100. Figure 1C In the illustrated example, control unit 140 includes processor 164, memory 166, input module 168, output module 170, and radio module 172. Control unit 140 may be coupled to and configured to control driving control component 154, navigation component 156, and one or more sensors 158 of vehicle 100. Radio module 172 may be configured to communicate with base station 180 via wireless communication link 182 (e.g., 5G, etc.), which provides connectivity to a server 184 of a third party (such as a law enforcement agency like a highway maintenance agency) via network 186 (e.g., the Internet).
[0052] Figure 2 Examples are shown of subsystems, computing elements, computing devices, or units within a device management system 200 that can be used within a vehicle 100. (See reference) Figures 1A to 2 In some implementations, various computing elements, computing devices, or units within the device management system 200 may be implemented within a system of interconnected computing devices (i.e., subsystems) that communicate data and commands to each other (e.g., by...). Figure 2 (As indicated by the arrow in the diagram). In other embodiments, various computing elements, computing devices, or units within the vehicle management system 200 may be implemented within a single computing device, such as individual threads, processes, algorithms, or computing elements. Therefore, Figure 2Each exemplified subsystem / computing element is also generally referred to herein as a "module," which may be implemented in one or more processing systems constituting the device management system 200. However, the use of the term "module" in describing various embodiments is not intended to imply or require the implementation of corresponding functionality within a single computing device or processing system of the ADS or ADAS device management system, in multiple computing systems or processing systems, or in a combination of dedicated hardware modules, software-implemented modules, and dedicated processing systems in a distributed device computing system, although each is a potential specific implementation. Rather, the use of the term "module" is intended to encompass subsystems with independent processing systems, computing elements (e.g., threads, algorithms, subroutines, etc.) operating in one or more computing devices and processing systems, and combinations of subsystems and computing elements.
[0053] In various implementations, the device management system 200 may include a radar sensing module 202, a camera sensing module 204, a positioning engine module 206, a map fusion and arbitration module 208, a route planning module 210, a sensor fusion and road world model (RWM) management module 212, a motion planning and control module 214, and a behavior planning and prediction module 216.
[0054] Modules 202 to 216 are merely examples of some modules in one example configuration of the device management system 200. In other configurations consistent with some implementations, additional modules may be included, such as additional modules for other sensing sensors (e.g., LIDAR sensing modules, etc.), additional modules for planning and / or control, additional modules for modeling, etc., and / or some modules in Modules 202 to 216 may be excluded from the device management system 200.
[0055] Each of modules 202 through 216 can exchange data, computation results, and commands with each other. Examples of some interactions between modules 202 through 216 are provided below. Figure 2 The arrows in the diagram illustrate this. Furthermore, the device management system 200 can receive and process data from sensors (e.g., radar, lidar, cameras, inertial measurement units (IMUs), etc.), navigation systems (e.g., Global Navigation Satellite System (GNSS) receivers, IMUs, etc.), vehicle networks (e.g., Controller Area Network (CAN) buses), and databases in memory (e.g., digital map data). The device management system 200 can output vehicle control commands or signals to the ADS or ADAS system / control unit 220, which is a system, subsystem, or computing device that directly interfaces with the vehicle's steering, throttle, and braking controls.
[0056] Figure 2The configurations of the device management system 200 and ADS / ADAS system / control unit 220 illustrated herein are merely example configurations, and other configurations of the vehicle management system and other vehicle components may be used in some implementations. As an example, Figure 2 The device management system 200 and ADS / ADAS system / control unit 220 illustrated herein can be configured for use in devices (e.g., vehicles) that are configured for automatic or semi-automatic operation, while different configurations can be used in non-automatic devices.
[0057] Camera perception module 204 may receive data from one or more cameras (such as cameras (e.g., 122, 136)) and process the data to identify and determine the location of other vehicles and objects (e.g., passengers, etc.) near and / or inside vehicle 100. Camera perception module 204 may include a trained neural network processing module that implements artificial intelligence methods to process image data to achieve object and vehicle identification, localization, and classification, and to pass such information to sensor fusion and RWM trained model 212 and / or other modules of the ADS / ADAS system.
[0058] The radar perception module 202 may receive data from one or more detection and ranging sensors (such as radar (e.g., 132) and / or lidar (e.g., 138)) and process the data to identify and determine the location of other vehicles and objects in the vicinity of vehicle 100. The radar perception module 202 may include the use of neural network processing and artificial intelligence methods to identify objects and vehicles, and pass such information to the sensor fusion and RWM trained model 212 of the ADS / ADAS system.
[0059] The positioning engine module 206 can receive data from various sensors and process that data to determine the location of the vehicle 100. These various sensors may include, but are not limited to, GNSS sensors, IMUs, and / or other sensors connected via a CAN bus. The positioning engine module 206 may also utilize input from one or more cameras (such as cameras (e.g., 122, 136)) and / or any other available sensors (such as radar, lidar, etc.).
[0060] Map fusion and arbitration module 208 can access data within a high-definition (HD) map database and receive output from positioning engine module 206, processing the data to further determine the location of vehicle 100 within the map, such as its position within a traffic lane, its position within a street map, etc. The HD map database can be stored in memory (e.g., memory 166). For example, map fusion and arbitration module 208 can convert latitude and longitude information from GNSS data into a position within a ground road map contained in the HD map database. GNSS positioning locking includes errors, so map fusion and arbitration module 208 can be used to determine the best guessed position of the vehicle within the road based on arbitration between GNSS coordinates and HD map data. For example, while GNSS coordinates might place the vehicle near the middle of a two-lane road in the HD map, map fusion and arbitration module 208 can determine the vehicle's most likely alignment with the lane of travel that aligns with its direction of travel. Map fusion and arbitration module 208 can then pass the map-based location information to sensor fusion and RWM trained model 212.
[0061] Route planning module 210 can use an HD map and input from the operator or dispatcher to plan a route to a specific destination for vehicle 100. Route planning module 210 can pass map-based location information to sensor fusion and RWM trained model 212. However, the use of prior maps is not required by other modules (such as sensor fusion and RWM trained model 212). For example, other processing systems can operate and / or control the vehicle based solely on perception data without providing a map, thereby constructing concepts of lanes, boundaries, and local maps as perception data is received.
[0062] The sensor fusion and RWM trained model 212 can receive data and outputs generated by the radar sensing module 202, camera sensing module 204, map fusion and arbitration module 208, and route planning module 210, and use some or all of these inputs to estimate or refine the position and state of the vehicle 100 relative to the road, other vehicles on the road, and other objects near and / or inside the vehicle 100. For example, the sensor fusion and RWM trained model 212 can combine image data from the camera sensing module 204 with arbitration map position information from the map fusion and arbitration module 208 to refine the determined location of the vehicle within a traffic lane. As another example, the sensor fusion and RWM trained model 212 can combine object recognition and image data from the camera sensing module 204 with object detection and ranging data from the radar sensing module 202 to determine and refine the relative positions of other vehicles and objects near the vehicle. As another example, sensor fusion and RWM trained model 212 can receive information about the location and direction of travel of other vehicles from vehicle-to-vehicle (V2V) communication (such as via CAN bus) and combine this information with information from radar sensing module 202 and camera sensing module 204 to refine the position and movement of other vehicles.
[0063] The sensor fusion and RWM trained model 212 can output refined position and status information of the vehicle 100, as well as refined position and status information of other vehicles and objects near the vehicle 100 or objects inside the vehicle 100, to the motion planning and control module 214 and / or the behavior planning and prediction module 216. As another example, the sensor fusion and RWM trained model 212 can apply facial recognition technology to images to identify specific facial patterns inside and / or outside the vehicle.
[0064] As a further example, the sensor fusion and RWM trained model 212 can use dynamic traffic control commands to guide vehicle 100 to change speed, lane, direction of travel, or other navigation elements, and combine this information with other received information to determine refined location and status information. The sensor fusion and RWM trained model 212 can output the refined location and status information of vehicle 100, as well as the refined location and status information of other vehicles and objects near vehicle 100 or objects inside vehicle 100, via wireless communication (such as via C-V2X connection, other wireless connections, etc.) to motion planning and control module 214, behavior planning and prediction module 216, and / or devices remote from vehicle 100, such as data servers, other vehicles, etc.
[0065] As a further example, the sensor fusion and RWM trained model 212 can monitor perception data from various sensors (such as perception data from radar perception module 202, camera perception module 204, other perception modules, etc.) and / or data from one or more sensors themselves to analyze the conditions in the vehicle's sensor data. The sensor fusion and RWM trained model 212 can be configured to detect conditions in the sensor data, such as sensor measurements being at, above, or below thresholds, the occurrence of certain types of sensor measurements (e.g., seat positioning movement, seat height change, etc.), and can output the sensor data as part of the refined position and status information of the vehicle 100, provided via wireless communication (such as via C-V2X connection, other wireless connections, etc.) to the behavior planning and prediction module 216 and / or devices remote from the vehicle 100 (such as data servers, other vehicles, etc.).
[0066] Detailed location and status information may include vehicle descriptors associated with the vehicle and its owner and / or operator, such as: vehicle specifications (e.g., size, weight, color, onboard sensor type, etc.); vehicle location, speed, acceleration, direction of travel, attitude, orientation, destination, fuel / power level, and other status information; vehicle emergency status (e.g., the vehicle is an emergency vehicle or a private individual is in an emergency); vehicle restrictions (e.g., weight / width load, turning restrictions, high-occupancy vehicle (HOV) authorization, etc.); vehicle capabilities (e.g., all-wheel drive, four-wheel drive, snow tires, chains, supported connection types, onboard sensor operating status, onboard sensor resolution level, etc.); equipment issues (e.g., low tire pressure, weak brakes, sensor malfunction, etc.); owner / operator travel preferences (e.g., preferred lane, road, route, and / or destination, preference for avoiding tolls or highways, preference for the fastest route, etc.); permission to provide sensor data to a data broker server (e.g., 184); and / or owner / operator identification information.
[0067] The behavior planning and prediction module 216 of the device management system 200 can use refined location and status information of the vehicle 100, as well as location and status information of other vehicles and objects output from sensor fusion and RWM trained by model 212, to predict the future behavior of other vehicles and / or objects. For example, the behavior planning and prediction module 216 can use such information to predict the future relative positions of these other vehicles based on its own vehicle positioning and speed, as well as the positioning and speed of other vehicles in the vicinity. Such predictions can take into account information from HD maps and route planning to anticipate changes in the relative vehicle positions as the primary vehicle and other vehicles travel along the road.
[0068] The behavior planning and prediction module 216 can output the behavior of other vehicles and objects, as well as position predictions, to the motion planning and control module 214. Additionally, the behavior planning and prediction module 216 can use object behavior combined with position predictions to plan and generate control signals for controlling the motion of vehicle 100. For example, based on route planning information, refined position in road information, and the relative positions and movements of other vehicles, the behavior planning and prediction module 216 can determine that vehicle 100 needs to change lanes and accelerate, such as to maintain or achieve a minimum distance from other vehicles and / or prepare for a turn or exit. Therefore, the behavior planning and prediction module 216 can calculate or otherwise determine the steering angle of the wheels and changes in throttle settings, which, along with various such parameters necessary to achieve such lane changes and accelerations, will be commanded to the motion planning and control module 214 and the ADS system / control unit 220. One such parameter could be a calculated steering wheel command angle.
[0069] The motion planning and control module 214 can receive data and information output from sensor fusion and the RWM trained model 212, as well as behavior and position predictions of other vehicles and objects from the behavior planning and prediction module 216. It uses this information to plan and generate control signals for controlling the motion of the vehicle 100, and to verify that such control signals meet the safety requirements of the vehicle 100. For example, based on route planning information, detailed location data in road information, and the relative positions and movements of other vehicles, the motion planning and control module 214 can verify various control commands or instructions and transmit them to the ADS system / control unit 220.
[0070] The ADS system / control unit 220 can receive commands or instructions from the motion planning and control module 214 and convert such information into mechanical control signals for controlling the wheel angles, braking, and throttle of the vehicle 100. For example, the ADS system / control unit 220 can respond to a calculated steering wheel command angle by transmitting a corresponding control signal to the steering wheel controller.
[0071] The ADS system / control unit 220 can receive data and information outputs from the motion planning and control module 214 and / or other modules in the device management system 200, and determine whether an event is occurring that needs to be notified to the decision-maker in the vehicle 100 based on the received data and information outputs.
[0072] Figure 3 This is a block diagram illustrating examples of components of a system-on-a-chip (SOC) 300 for use in a processing system (e.g., a V2X processing system) according to various embodiments, used when performing operations in a device. (See also...) Figures 1A to 3The processing device SOC 300 may include several heterogeneous processors, such as a digital signal processor (DSP) 303, a modem processor 304, an image and object recognition processor 306, a mobile display processor 307, an application processor 308, and a resource and power management (RPM) processor 317. The processing device SOC 300 may also include one or more coprocessors 310 (e.g., vector coprocessors) connected to one or more of the heterogeneous processors 303, 304, 306, 307, 308, and 317.
[0073] Each of these processors may include one or more cores and an independent / internal clock. Each processor / core may perform operations independently of the other processors / cores. For example, the processing device SOC 300 may include a processor running a first type of operating system (e.g., FreeBSD, LINUX, OS X, etc.) and a processor running a second type of operating system (e.g., Microsoft Windows). In some embodiments, the application processor 308 may be the main processor, central processing unit (CPU), microprocessor unit (MPU), arithmetic logic unit (ALU), etc. of the SOC 300. The graphics processor 306 may be a graphics processing unit (GPU).
[0074] The processing device SOC 300 may include analog circuitry and custom circuitry 314 for managing sensor data, analog-to-digital conversion, wireless data transmission, and performing other specialized operations, such as processing encoded audio and video signals for rendering in a web browser. The processing device SOC 300 may also include system components and resources 316, such as voltage regulators, oscillators, phase-locked loops, peripheral bridges, data controllers, memory controllers, system controllers, access ports, timers, and other similar components for supporting processors and software clients (e.g., web browsers) running on computing devices.
[0075] The processing device SOC 300 may also include a dedicated circuit (CAM) 305 for camera actuation and management. This CAM includes, provides, controls, and / or manages the operation of one or more cameras (e.g., a main camera, webcam, 3D camera, etc.), video display data from camera firmware, image processing, video preprocessing, video front-end (VFE), embedded JPEG, high-definition video codecs, etc. The CAM 305 may be a separate processing unit and / or include a separate or internal clock.
[0076] In some embodiments, the image and object recognition processor 306 may be configured with processor-executable instructions and / or dedicated hardware configured to perform image processing and object recognition analysis as described in various embodiments. For example, the image and object recognition processor 306 may be configured to process images received from a camera via CAM 305 to identify and / or mark other vehicles. In some embodiments, the processor 306 may be configured to process radar or lidar data.
[0077] System components and resources 316, analog and custom circuitry 314, and / or CAM 305 may include circuitry for interfacing with peripheral devices such as cameras, radar, lidar, electronic displays, wireless communication devices, external memory chips, etc. Processors 303, 304, 306, 307, and 308 may be interconnected via interconnect / bus module 324 to one or more memory elements 312, system components and resources 316, analog and custom circuitry 314, CAM 305, and RPM processor 317. This interconnect / bus module may include reconfigurable logic gate arrays and / or implement bus architectures (e.g., CoreConnect, AMBA, etc.). Communication may be provided by advanced interconnects such as high-performance on-chip networks (NoC).
[0078] The processing device SOC 300 may also include input / output modules (not shown) for communicating with external resources such as clock 318 and voltage regulator 320. External resources (e.g., clock 318, voltage regulator 320) may be shared by two or more internal SOC processors / cores (e.g., DSP 303, modem processor 304, graphics processor 306, application processor 308, etc.).
[0079] In some implementations, the processing device SOC 300 may be included in a control unit (e.g., 140) for use in a vehicle (e.g., 100). The control unit may include communication links for communicating with a telephone network (e.g., 180), the Internet, and / or a web server (e.g., 184), as described.
[0080] The processing device SOC 300 may also include additional hardware and / or software components suitable for collecting sensor data from sensors, including motion sensors (e.g., accelerometers and gyroscopes of an IMU), user interface elements (e.g., input buttons, touchscreen displays, etc.), microphone arrays, sensors for monitoring physical conditions (e.g., position, orientation, motion, orientation, vibration, pressure, etc.), cameras, compasses, satellite navigation system receivers, and communication circuitry (e.g., Bluetooth). ®(such as WLAN, Wi-Fi, etc.) and other well-known components of modern electronic devices.
[0081] Figure 4A This is a processing block diagram 400 illustrating various operations performed on camera images from a device camera as part of routine ADS or ADAS processing. (Reference) Figures 1A to 4A Image frames 402 from multiple device cameras can be received by an image processing system such as a camera perception module 204. This image processing system may include multiple modules, processing systems, and trained machine model / AI modules configured to perform various operations necessary to obtain information from the images to support vehicle navigation and safe operation. While not implying inclusion, Figure 4A Examples of some of the processes involved in supporting the operation of automated devices are shown.
[0082] Image frame 402 can be processed by object detection module 404, which performs operations associated with detecting objects within the image frame based on various image processing techniques. As discussed, automated vehicle image processing involves multiple detection methods and analysis modules that focus on different aspects of the image to provide the information needed for safe navigation in an ADS or ADAS system. The processing of the image frame in object detection module 404 can involve multiple different detectors and modules that process the image in different ways to identify objects, define boundary blocks covering the objects, and identify the location of detected objects within the frame coordinates. The outputs of various detection methods can be combined in an ensemble detection, which can be a list, table, or data structure of detections performed by the individual detectors processing the image frame. Therefore, the ensemble detection in object detection module 404 can aggregate the outputs of various detection mechanisms and modules for object classification and tracking, and vehicle control decisions.
[0083] As discussed, image processing supporting autonomous driving systems involves other image processing tasks 406. As an example of other tasks, image frames can be analyzed to determine road features and the 3D depth of detected objects. Other processing tasks 406 may include panoptic segmentation, a computer vision task that includes both instance segmentation and semantic segmentation. Instance segmentation involves identifying and classifying objects of multiple categories observed within an image frame. By addressing instance segmentation and semantic segmentation problems together, panoptic segmentation enables ADS or ADAS systems to understand a given scene in greater detail.
[0084] The outputs of object detection method 404 and other tasks 406 can be used for object classification 410. As described, this may involve classifying features and objects detected in an image frame using classifications important to the decision-making process of an autonomous driving system (e.g., road features, traffic signs, pedestrians, other vehicles, etc.). As illustrated, the methods described herein can be used to examine identified features, such as traffic signs 408 within segments or bounding boxes in an image frame, to assign classifications to individual objects and obtain information about the objects or features (e.g., a speed limit of 50 km / h based on the identified traffic sign 408).
[0085] The output of object classification 410 can be used to track various features and objects from one frame to the next 412. As described above, feature and object tracking is important for identifying features / objects relative to the device's trajectory in order to navigate and avoid collisions.
[0086] Figure 4B This is an example of the components and data flow diagram 420 illustrating the processing of camera images used to generate data for object tracking in conventional ADS and ADAS systems. (Reference) Figures 1A to 4B Image data from each camera 422a to 422n of the device can be provided to and processed by multiple neural network AI modules, which are trained to perform specific types of image processing, including semantic segmentation, depth estimation, object detection, and object classification.
[0087] Image data from one or more cameras 422a to 422n can be processed by a semantic segmentation module 424, which can be an AI / ML network trained to receive image data as input and produce outputs that associate groups of pixels or masks in the image with classification labels. Semantic segmentation is a computational process of dividing a digital image into multiple segments, masks, or “superpixels,” where each segment is identified or corresponds to a predefined category or class. The goal of semantic segmentation is to assign labels to each pixel or group of pixels in the image (e.g., pixels spanning a mask) such that pixels with the same label share certain characteristics. Non-limiting examples of classification labels include “trees,” “traffic signs,” “pedestrians,” “roads,” “buildings,” “cars,” “sky,” etc. The coordinates of the labeled masks within the digital image can be defined by coordinates within the image (e.g., pixel coordinates) or by the coordinates and regions of each mask within the image.
[0088] The AI / ML semantic segmentation module 424 can employ an encoder-decoder architecture, where the encoder performs feature extraction and the decoder performs pixel-by-pixel classification. The encoder may include a series of convolutional layers followed by pooling layers, thereby increasing depth while reducing spatial dimensionality. The decoder reverses this process through a series of upsampling and deconvolutional layers, thereby restoring spatial dimensionality while applying the learned features to each pixel for segmentation. Using such a process, the semantic segmentation module 424 in devices such as transportation vehicles can achieve real-time detection of pedestrians, road signs, and other vehicles.
[0089] Image data from one or more cameras 422a to 422n can be processed by a depth estimation module 426, which is trained to receive image data as input and produce an output that estimates the distance from the camera or device to the object associated with each pixel or group of pixels. The depth estimation module 426 can use various methods to estimate the distance or depth of each pixel. Non-limiting examples of such methods include using a model with a dense vision transformer trained on the data set to achieve monocular depth estimation of individual pixels and groups of pixels, as described in R. Ranftl et al., “Computer Vision and Pattern Recognition (cs.CV)”, arXiv:2103.13413 [cs.CV]. Another non-limiting example of this approach uses a hierarchical transformer encoder to capture and transmit the global context of the image, and a lightweight decoder to generate the estimated depth map while taking local connectivity into account, as described in D. Kim et al., “Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth,” arXiv:2201.07436v3 [cs.CV]. Additionally, parallax-based stereo depth estimation methods can also be used to estimate the depth of objects associated with pixels in two (or more) images separated by a known distance, such as two images taken approximately simultaneously by two spaced-apart cameras, or two images taken by a single camera at different instances on a moving device.
[0090] Image data from one or more cameras 422a to 422n can be processed by an object detection module 428, which can be an AI / ML network trained to receive image data as input and produce output identifying individual objects within the image, including pixel coordinates defining bounding boxes around each detected object. As an example, the object detection module 428 may include neural network layers configured and trained to divide a digital image into regions or grids, pass pixel data within each region or grid through a convolutional network to extract features, and then process the extracted features through layers trained to classify objects and define bounding box coordinates. Known methods for training the neural network for the object detection module can use a large training dataset of images (e.g., images collected by cameras on vehicles traveling along many driving routes) that include a variety of objects that may be encountered, annotated with ground truth information including appropriate labels for each object in the images, wherein the appropriate labels are manually assigned to each object in each training image.
[0091] Image data from one or more cameras 422a to 422n can be processed by an object classification module 430, which can be an AI / ML network trained to receive image data as input and produce output classifying objects in the images. Object classification involves classifying detected objects into predefined classes or labels, which can be performed after object detection and is crucial for decision-making, path planning, and event prediction within an automated navigation framework. Known methods for training object classification modules for ADS or ADAS applications can utilize a broad training database of images containing various objects with ground-based information about suitable classifications for each object.
[0092] As illustrated, the outputs of image processing modules 424 to 430 can be combined to generate a data structure 432 that, for each object identified in the image, includes an object tracking number or identifier, a bounding box (i.e., the pixel coordinates of the box defining the object), and the object's classification. This data structure can then be used for object tracking 434 to support ADS or ADAS navigation, path planning, and collision avoidance processing.
[0093] Although reference Figure 4A and Figure 4BThe described processing can provide sufficient information about the scene surrounding the device to enable automated manipulation; however, the results may be susceptible to visual attacks that could deceive or confuse one or more of the image processing modules 424 to 430. To overcome this vulnerability, various implementations include a consistency check configured to identify inconsistencies in the output of the image processing modules 424 to 430, which can be used to identify actual or potential visual attacks.
[0094] Figure 5A This is a processing block diagram 500 illustrating various operations performed on camera images from a device camera as part of routine ADS or ADAS processing. (Reference) Figures 1A to 5A Image frames 402 from multiple device cameras can be received by an image processing system such as a camera perception module 204. This image processing system may include multiple modules, processing systems, and trained machine model / AI modules configured to perform various operations necessary to obtain information from the images to support vehicle navigation and safe operation. While not implying inclusion, Figure 5A Examples are given of some of the processes involved in supporting the operation of automated devices and identifying visual attacks and taking mitigation actions, according to various implementation schemes.
[0095] Image frame 402 can be processed by object detection module 404, which performs operations associated with detecting objects within the image frame based on various image processing techniques. As discussed, image processing for autonomous vehicles involves multiple detection methods and analysis modules that focus on different aspects of using image streams to provide the information needed for safe navigation in autonomous driving systems. The processing of image frames in object detection module 404 can involve multiple different detectors and modules that process the image in different ways to identify objects, define boundary blocks covering the objects, and identify the location of detected objects within the frame coordinates. The outputs of various detection methods can be combined in an ensemble detection, which can be a list, table, or data structure of detections performed by the individual detectors processing the image frame. Therefore, the ensemble detection in object detection module 404 can aggregate the outputs of various detection mechanisms and modules for object classification and tracking, and vehicle control decisions.
[0096] As discussed, image processing supporting autonomous driving systems involves other image processing tasks 406. As an example of other tasks, image frames can be analyzed to determine road features and the 3D depth of detected objects. Other processing tasks 406 may include panoptic segmentation, a computer vision task that includes both instance segmentation and semantic segmentation. Instance segmentation involves identifying and classifying objects of multiple categories observed within an image frame. By addressing instance segmentation and semantic segmentation problems together, panoptic segmentation enables autonomous driving systems to understand a given scene in greater detail.
[0097] The outputs of object detection method 404 and other tasks 406 can be used for object classification 410. As described, this may involve classifying features and objects detected in an image frame using classifications important to the decision-making process of an autonomous driving system (e.g., road features, traffic signs, pedestrians, other vehicles, etc.). As illustrated, the methods described herein can be used to examine identified features, such as traffic signs 408 within segments or bounding boxes in an image frame, to assign classifications to individual objects and obtain information about the objects or features (e.g., a speed limit of 50 km / h based on the identified traffic sign 408). Also as part of object classification 410, the techniques described herein can be used to examine the image frame for projection attacks.
[0098] In operation 502, the outputs of set object detection 404 and other processing tasks 406 can also be correlated, allowing the outputs of selected processing tasks to be compared in task consistency check 504. As further described herein, task consistency check 504 can be configured to identify inconsistencies in the outputs of two or more different image processing methods performed on an image, which could indicate camera or visual attacks. Consistency checker 504 may also be referred to as or used as a sensor, detector, or configured to identify inconsistencies between the outputs of two or more different types of image processing involved in ADS and ADAS systems that rely on cameras for navigation and object avoidance.
[0099] The output of object classification 410 can be combined with indications of inconsistencies identified by consistency checker 504 to include indications of inconsistencies in object tracking data within multi-object tracking operation 506. As described above, feature and object tracking is crucial for identifying the trajectory of features / objects relative to the vehicle for navigation and collision avoidance. Using the output of consistency checker 404, multiple tracking operation 506 provides safe multi-object tracking 508 to support vehicle control functions 220 of the autonomous driving system. Additionally, feature / object tracking can be used in safety decision module 510, which is configured to detect inconsistencies that may indicate or suggest a visual attack. Such safety decisions can be used to report conclusions to remote services 512.
[0100] Figure 5B This is an example of the components and data flow diagram 520 illustrating the process of generating data for object tracking, including camera image processing and consistency checks, according to various implementation schemes. (Reference) Figures 1A to 5B Image data from each camera 422a to 422n of the device can be provided to and processed by multiple neural network AI modules, which are trained to perform specific types of image processing, including semantic segmentation, depth estimation, object detection, and object classification.
[0101] For reference Figure 4B As described, image data from one or more cameras 422a to 422n can be processed by multiple image processing modules 424 to 430. As described, image processing modules 424 to 430 can be AI / ML modules, which include: a semantic segmentation module 424 trained to associate groups of pixels or masks in the image with classification labels; a depth estimation module 426 that estimates the depth of each pixel or group of pixels; an object detection module 428 that identifies individual objects within bounding boxes in the image; and an object classification module 430 that classifies the objects in the image.
[0102] In various implementations, the outputs of image processing modules 424 to 430 are examined for inconsistencies between the outputs of different modules, which may indicate or prove a visual attack. As illustrated, the outputs of selected processing modules may be associated with a specific consistency checker 504. For example, the outputs of semantic segmentation module 424 and object detection module 428 may be provided to a semantic consistency checker 522, the outputs of semantic segmentation module 424, depth estimation module 426, and object detection module 428 may be provided to a depth plausibility checker 524, the outputs of semantic segmentation module 424, depth estimation module 426, and object detection module 428 may be provided to a contextual consistency checker 526, and the outputs of object detection module 428 and object classification module may be provided to a label consistency checker 524.
[0103] As described, the semantic consistency checker 522 compares the output of the semantic segmentation processing of the image with the bounding boxes surrounding the detected objects from the object detection processing to determine whether the label assigned to the semantic segmentation mask matches or does not match the detected object bounding boxes. In some embodiments, mask labels from semantic segmentation that differ from or do not cover the labels assigned in the object detection / object classification processing can be identified as inconsistencies. In some embodiments, in the case of label matching, the positions of the corresponding segmentation mask and the detected object bounding boxes in the image can be compared, and inconsistencies are identified if the overlap between the mask and bounding box positions is not within a threshold percentage. In the case of identifying any inconsistency, an appropriate indication of the inconsistency (e.g., a "1" in the inconsistency label or its location) can be output for object tracking.
[0104] As described, the depth reasonableness checker 524 can compare the distribution of depth estimates from individual pixels or pixel masks from semantic segmentation with the depth estimates of pixels across detected objects to determine whether the two depth distributions are consistent or inconsistent. In some implementations, depth inconsistencies can be identified if the distribution of depth estimates across pixels across a segmentation mask differs from the distribution of depth estimates across pixels within the mask by more than a threshold amount, and an appropriate indication of the inconsistency (e.g., a "1" in the inconsistency label or its location) can be output for object tracking.
[0105] As described, inconsistencies indicating spoofed images or deceptive objects can be examined by the context consistency checker 526, the depth estimate of the detected object, and the depth estimate of the rest of the environment in the scene. In some embodiments, the checker or detector can compare the estimated depth value of the pixels of the detected object or the mask covering the object with the estimated depth value of the pixels of the overlapping mask. In some embodiments, the checker or detector can compare the distribution of the estimated pixel depth values across the bounding box of the detected object or the object with the distribution of the estimated pixel depth values across the overlapping mask, thereby comparing the difference with a threshold indicating an actual or potential visual attack or an inconsistency that allows for action. Upon identification of an inconsistency, an appropriate indication of the inconsistency (e.g., a "1" in the inconsistency label or its location) can be output for object tracking.
[0106] As described, the label consistency checker 528 compares the labels assigned to detected objects from the object detection process with the labels of the same objects or regions (within a mask) obtained from the object classification process to determine whether the object classification labels are consistent with the detected objects. Label inconsistencies can be identified in cases where the labels assigned to the same object or mask by the two labeling processes (semantic segmentation and object detection / classification) do not match or fall into different unique categories, and appropriate indications of the inconsistencies (e.g., a "1" or location of the inconsistent label) can be output for object tracking.
[0107] The outputs of consistency checkers 522 to 528, in the form of indications of attacks (or potential attacks) or real data (e.g., in a single-bit flag), can be combined with or appended to the outputs of image processing modules 424 to 430 to generate data structure 530. This data structure includes an object tracking number or identifier for each object identified in the image, a bounding box (i.e., pixel coordinates defining the box covering the object), the object's classification, and indications of different consistency or inconsistency results from consistency checkers 522 to 528. As an illustrative example, object #1 includes an indication (e.g., 1 or 0) that indicates the semantic consistency check identified an inconsistency that might indicate an attack, while other consistency checkers did not find the inconsistency. Data structure 530 can then be used for object tracking 532 to support ADS or ADAS navigation, path planning, and collision avoidance processing, with the improvement that the object data includes information related to indications of potential attacks identified by inconsistency checkers 522 to 528.
[0108] Figure 6 This is a process flowchart of example method 600 according to various implementations, a method for detecting visual attacks performed by a processing system on a device (e.g., a vehicle), the device being used to detect and react to potential attacks on the device's camera system. Reference Figures 1A to 6The operation of method 600 may be performed by a processing system (e.g., 102, 120, 240) comprising one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130) and / or hardware elements, any one or a combination thereof, which may be configured to perform any operation of method 600. Furthermore, one or more processors within the processing system may be configured with software or firmware to perform various operations of the method. To encompass any of the processors, hardware elements, and software elements that may be involved in performing method 600, the element performing the method operation is referred to as the "processing system". Additionally, components for performing the functions of method 600 may include the processing system (e.g., 102, 120, 240), which includes one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130), memory 112, radio module 118, and one or more cameras (e.g., 122, 136).
[0109] In box 602, the processing system may perform operations including receiving images (such as, but not limited to, images from a camera image frame stream) from one or more cameras of a device (e.g., a vehicle). For example, images may be received from a forward-facing camera used by an ADS or ADAS for observing the road ahead for navigation and collision avoidance purposes.
[0110] In block 604, the processing system may perform operations including processing images received from the device's camera to obtain multiple image processing outputs. In some embodiments, image processing may be performed by multiple neural network processors that have been trained using machine learning methods (referred to herein as "trained image processing models") to receive images as input and generate outputs of the type that provide the device system (e.g., an ADS or ADAS system) with the processed information required by the device. In some embodiments, the operations performed in block 604 may include processing images received from the device's camera using multiple different trained image processing models to obtain multiple different image processing outputs. As described, camera images may be processed by multiple different processing systems, including trained neural network processing systems, to extract information necessary for the safe navigation of the device. (See reference...) Figure 7 In more detail, these operations may include semantic segmentation, depth estimation, object detection, and / or object classification.
[0111] In block 606, the processing system may perform operations including: performing multiple consistency checks on multiple image processing outputs, wherein each of the multiple consistency checks compares each of the multiple outputs to detect inconsistencies. In some embodiments, the operations performed in block 606 may include: performing multiple consistency checks on multiple distinct image processing outputs, wherein each of the multiple consistency checks compares two or more selected outputs of the multiple distinct outputs to detect inconsistencies. (See reference...) Figures 8A to 8D In a more detailed description, the multiple consistency checks may include: a semantic consistency check, which compares the classification label associated with the mask from the semantic segmentation process with the bounding boxes of object detections from the object detection process in the image; a positional consistency check, which compares the position of the classification mask from the semantic segmentation process within the image with the position of the bounding boxes of object detections from the object detection process within the image; a depth reasonableness check, which compares the depth estimate of the detected object from the object detection process with the depth estimates of individual pixels or groups of pixels from the depth estimation process; and a contextual consistency check, which compares the depth estimate of the bounding box covering the detected object from the object detection process with the depth estimate of the mask covering the detected object from the semantic segmentation process.
[0112] In block 608, the processing system may perform operations including: using detected inconsistencies to identify an attack on the device's cameras. In some embodiments, the processing system may identify an attack on one or more cameras of the device in response to detecting one or a threshold number of inconsistencies in an image. In some embodiments, the results of various consistency checks performed in block 606 may be used in a decision algorithm to identify whether an attack on a vehicle camera is occurring or is likely to occur. Such a decision algorithm may be as simple as identifying a visual attack, provided that any of the different inconsistency checking processes indicates the likelihood of contact. More complex algorithms may include assigning weights to each of the various inconsistency checks and accumulating the results in a voting or thresholding algorithm to determine whether a visual attack is more likely.
[0113] In determination box 610, the processing system can detect attacks based on inconsistencies in image processing performed as in box 606.
[0114] In response to the detection of a visual attack (or determination that a visual attack may exist) (i.e., determination box 610 = "Yes"), the processing system may perform a mitigation action in box 612. In some embodiments, the mitigation action may include adding an indication of inconsistency from each of the plurality of consistency checks to information about each detected object, which is provided to the autonomous driving system for tracking the detected object. Adding an indication of inconsistency to the object tracking information enables a device (e.g., a vehicle's) ADS or ADAS to identify and compensate for visual attacks, such as ignoring or underemphasizing information from the camera being attacked. In some embodiments, the mitigation action may include reporting the detected attack to a remote system, such as law enforcement or highway maintenance organizations, so that the threat or cause of the malicious attack can be stopped or removed. In some embodiments, the mitigation action may include outputting an indication of the visual attack, such as a warning or notification to the operator. In some embodiments, the processing system may perform more than one mitigation action.
[0115] The operation of method 600 can be performed continuously. Therefore, in response to no attack being detected (i.e., determining box 610 = "No") and / or after taking mitigation actions in box 612, the processing system can repeat method 600 by receiving another image from the device camera again in box 602 and performing the methods as described.
[0116] Figure 7 This is a flowchart of a process, according to some embodiments, to perform image processing methods on images from a camera of a device to support Adaptive Digital Subsystems (ADS) or Advanced Driver Assistance Systems (ADAS), where the output of the ADS or ADAS can be processed to identify inconsistencies that may indicate visual attacks or potential visual attacks. Specifically, Figure 7 Examples are illustrated of operations that can be performed in block 604 of method 600 when processing images received from the camera of the device, according to various embodiments. Reference Figures 1A to 7 Operation 604 may be performed by a processing system (e.g., 102, 120, 240) comprising one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130) and / or hardware elements, any one or a combination thereof being configured to perform any operation of the operation. Furthermore, one or more processors within the processing system may be configured with software or firmware to perform various operations. To encompass any of the processors, hardware elements, and software elements that may be involved in performing the illustrated operations, the element performing the method operation is referred to as the "processing system". Additionally, components for the functionality of performing the illustrated operations may include the processing system (e.g., 102, 120, 240), which includes one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130), memory 112, and / or a vehicle camera (e.g., 122, 136).
[0117] After receiving an image from the device's camera (e.g., an image frame from an image stream from the camera), the processing system may perform operations including: performing semantic segmentation processing on the image using a trained semantic segmentation model in box 702 to associate masks of pixel groups in the image with classification labels. The semantic segmentation processing may include processing performed by an AI / ML network trained to receive image data as input and produce output that associates pixel groups or masks in the image with classification labels. Semantic segmentation may include dividing the image into multiple masks, where each mask is assigned a predefined category or class.
[0118] In box 704, the processing system can perform operations including: performing depth estimation processing on an image using a trained AI / ML depth estimation model to identify distances to pixels covering detected objects in the image. The depth estimation performed in box 704 can generate a map of pixel depth estimates across some or all of the images. As described above, the depth estimation processing can use an AI / ML depth estimation model based on monocular depth estimation, or a hierarchical transformer encoder to capture and transmit the global context of the image, and a lightweight decoder to generate the estimated depth map. Pixel depth estimation can also, or alternatively, use stereo depth estimation methods based on spatial and / or temporal disparity.
[0119] In box 706, the processing system may perform operations including: performing object detection processing on an image using an AI / ML network object detection model trained to identify objects in the image and define bounding boxes around the identified objects. In some embodiments, the object detection processing may include processing by a neural network layer configured and trained to divide the digital image into regions or grids, pass pixel data within each region or grid through a convolutional network to extract features, and then process the extracted features through layers trained to classify objects and define bounding box coordinates. The output of box 706 may be multiple bounding boxes surrounding the detected objects within each image.
[0120] In box 708, the processing system may perform operations including: performing object classification processing on an image using an AI / ML network object classification model trained to classify objects in the image. In some implementations, the object classification processing may include classifying detected objects into predefined classes or labels.
[0121] Figures 8A to 8D This is a flowchart illustrating a method for identifying inconsistencies in the processing of images from a camera of a device used to identify visual attacks or potential visual attacks, according to some implementation schemes. Specifically, Figures 8A to 8D Example methods 800a to 800d are illustrated, which can be executed in block 606 of method 600 to identify as referenced. Figure 7 The inconsistencies between the results of the image processing operations in box 604 of method 600, as described in boxes 702 to 708, are presented. Figures 8A to 8D The order in which methods 800a to 800d are described is arbitrary, and the processing system can execute methods 800a to 800d in any order, and in some embodiments, fewer methods may be executed than all of them. Reference Figures 1A to 8D The operations in methods 800a to 800d may be performed by a processing system (e.g., 102, 120, 240) comprising one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130) and / or hardware elements, any one or a combination thereof, which may be configured to perform any operation of the operations. Furthermore, one or more processors within the processing system may be configured with software or firmware to perform various operations. To encompass any of the processors, hardware elements, and software elements that may be involved in performing the illustrated operations, the element performing the method operations is referred to as the “processing system.” Additionally, components for the functionality of performing the illustrated operations may include the processing system (e.g., 102, 120, 240), which includes one or more processors (e.g., 110, 123, 124, 126, 127, 128, 130), memory 112, and / or a vehicle camera (e.g., 122, 136).
[0122] refer to Figure 8A In block 802 of method 800a, the processing system may perform an operation including a semantic consistency check that compares the classification labels associated with the mask from the semantic segmentation process with bounding boxes of objects detected in the image from the object detection process to identify inconsistencies between the mask classification and the detected objects. As described herein, the semantic consistency check may include the processing system comparing the output of the semantic segmentation process of the image with bounding boxes around objects detected in the object detection process to determine whether the labels assigned to the semantic segmentation mask are consistent with or inconsistent with the detected object bounding boxes.
[0123] In box 804, the processing system can determine whether any classification inconsistencies in the image are identified in the semantic segmentation processing and object detection processing of the image.
[0124] In response to determining that one or more classification inconsistencies are identified in an image (i.e., determination box 804 = "Yes"), the processing system may perform operations including: providing an indication of the detected classification inconsistency in box 806 in response to the mask classification being inconsistent with detected objects in the image. In some embodiments, this indication may be information provided to a decision-making process configured to determine whether or potentially detecting a visual attack on the camera is detected based on one or more identified inconsistencies. In some embodiments, this indication may be information that can be included with or attached to object tracking information as described herein. In some embodiments, this indication may be information that can be included in a report of an image attack, or used to generate a report of such an image attack for submission to a remote server as described herein. In some embodiments, this indication may be another signal, information, or response that enables the device ADS or ADAS to respond to or adapt to the identified inconsistency.
[0125] In response to determining that no classification inconsistency in image processing was identified (i.e., determining box 804 = "No"), the processing system may perform the following operations: performing a positional consistency check in box 808, which compares the position of the classification mask from the semantic segmentation process within the image with the position of the bounding box of the object detection from the object detection process within the image to identify inconsistencies between the position of the classification mask and the position of the detected object bounding box.
[0126] In box 810, the processing system may perform operations including providing an indication of a detected classification inconsistency when the location of the classification mask within the image is inconsistent with the location of the detected object bounding box. As described, this indication may be information provided to the decision-making process, information that may be included in or attached to object tracking information, information that may be included in or used to generate information to a remote server, and / or enable the device ADS or ADAS to respond to or adapt to another signal, information, or response to the identified inconsistency.
[0127] Subsequently, the processing system may perform the operations of block 606 of method 600 as described, and / or other operations for checking inconsistencies in image processing, such as performing method 800b. Figure 8B ), 800c ( Figure 8C ) and / or 800d ( Figure 8D The operations in ).
[0128] refer to Figure 8BIn block 812 of method 800b, the processing system may perform an operation including a depth plausibility check, which compares depth estimates of detected objects from object detection processing with depth estimates of individual pixels or groups of pixels from depth estimation processing to identify a distribution in the depth estimates across pixels of detected objects that is inconsistent with a depth distribution associated with a classification of a mask covering the detected objects from semantic classification processing. As described herein, the depth plausibility check may include: identifying inconsistencies between depth or distance estimates of pixels or groups of pixels within a classification mask and / or detected objects and depth or distance estimates of the classification mask and / or detected objects as a whole within the image.
[0129] In block 814, the processing system may perform operations including: providing an indication of a reasonableness check for the detected depth when the depth or distance estimate of a pixel or group of pixels within a classification mask and / or a detected object is inconsistent with the depth or distance estimate of the classification mask and / or the detected object as a whole within the image. As described, this indication may be information provided to a decision-making process, information that may be included in or attached to object tracking information, information that may be included in or used to generate information to a remote server, and / or enable the device ADS or ADAS to respond to or adapt to another signal, information, or response to the identified inconsistency.
[0130] Subsequently, the processing system may perform the operations of block 606 of method 600 as described, and / or other operations for checking inconsistencies in image processing, such as performing method 800a ( Figure 8A ), 800c ( Figure 8C ) and / or 800d ( Figure 8D The operations in ).
[0131] refer to Figure 8C In block 822 of method 800c, the processing system may perform an operation including a context consistency check that compares depth estimates of bounding boxes covering detected objects from object detection processing with depth estimates of masks covering detected objects from semantic segmentation processing to determine whether the distribution of mask depth estimates differs from the depth estimates of bounding boxes. As described herein, the context consistency check may include identifying inconsistencies between the distribution of depth estimates of classification masks and the distribution of depth estimates of bounding boxes of detected objects.
[0132] In block 824, the processing system may perform operations including providing an indication of detected contextual inconsistency when the distribution of the mask's depth estimate is the same as or similar to the distribution of the bounding box's depth estimate. As described, this indication may be information provided to the decision-making process, information that may be included in or attached to object tracking information, information that may be included in or used to generate information to a remote server, and / or enable the device ADS or ADAS to respond to or adapt to another signal, information, or response to the identified inconsistency.
[0133] Subsequently, the processing system may perform the operations of block 606 of method 600 as described, and / or other operations for checking inconsistencies in image processing, such as performing method 800a ( Figure 8A ), 800b ( Figure 8B ) and / or 800d ( Figure 8D The operations in ).
[0134] refer to Figure 8D In block 832 of method 800d, the processing system may perform an operation including a label consistency check, which compares detected objects from the object detection process with the labels of the detected objects from the object classification process to determine whether the object classification label is consistent with the detected object. As described herein, the label consistency check may include the processing system determining whether labels assigned to the same object or mask by the two labeling processes (semantic segmentation and object detection / classification) do not match or belong to different unique categories (e.g., "tree" vs. "car" or "traffic sign" vs. "pedestrian").
[0135] In block 834, the processing system may perform operations including providing an indication of a detected label inconsistency when the object classification label is inconsistent with the detected object. As described, this indication may be information provided to the decision-making process, information that may be included in or attached to object tracking information, information that may be included in or used to generate information to a remote server, and / or enable the device ADS or ADAS to respond to or adapt to another signal, information, or response to the identified inconsistency.
[0136] Subsequently, the processing system may perform the operations of block 606 of method 600 as described, and / or other operations for checking inconsistencies in image processing, such as performing method 800a ( Figure 8A ), 800b ( Figure 8B ) and / or 800c ( Figure 8C The operations in ).
[0137] Specific implementation embodiments are described in the following paragraphs. While some of the following specific implementation examples are described in the form of example systems and methods, further example implementations may include: example operations discussed in the following paragraphs that can be implemented by various computing devices; example methods implemented by an apparatus (e.g., a vehicle) discussed in the following paragraphs, the apparatus including a processing system comprising one or more processors configured to perform operations of the methods of the following specific implementation examples using processor-executable instructions; example methods implemented by an apparatus discussed in the following paragraphs, the apparatus including components for performing the functions of the methods of the following specific implementation examples; and example methods discussed in the following paragraphs that can be implemented as a non-transitory processor-readable storage medium storing processor-executable instructions configured to cause the processing system of the apparatus to perform operations of the methods of the following specific implementation examples.
[0138] Example 1. A method for detecting 1. A method for detecting a visual attack performed by a processing system on a device, the method comprising: processing images received from a camera of the device using a plurality of trained image processing models to obtain a plurality of image processing outputs; performing a plurality of consistency checks on the plurality of image processing outputs, wherein the consistency check in the plurality of consistency checks compares each of the plurality of image processing outputs to detect inconsistencies; detecting an attack on the camera based on the inconsistencies; and performing mitigation actions in response to identifying the attack.
[0139] Example 2. The method according to Example 1, wherein processing the image received from the camera of the device using multiple trained image processing models to obtain multiple image processing outputs includes: performing semantic segmentation processing on the image using a trained semantic segmentation model to associate a mask of a group of pixels in the image with a classification label; performing depth estimation processing on the image using a trained depth estimation model to identify the distance to an object in the image; performing object detection processing on the image using a trained object detection model to identify objects in the image and define bounding boxes around the identified objects; and performing object classification processing on the image using a trained object classification model to classify the objects in the image.
[0140] Example 3. According to the method of Example 2, wherein performing the plurality of consistency checks on the plurality of image processing outputs includes: performing a semantic consistency check, the semantic consistency check comparing a classification label associated with a mask from a semantic segmentation process with a bounding box of an object detected from an object detection process in the image to identify inconsistencies between the mask classification and the detected objects; and providing an indication of detected classification inconsistencies in response to the inconsistency between the mask classification and the detected objects in the image.
[0141] Example 4. According to the method of Example 3, the method further includes: in response to a classification label associated with a mask from semantic segmentation processing being consistent with a bounding box of object detection from object detection processing, performing a position consistency check, the position consistency check comparing the position of the classification mask from semantic segmentation processing in the image with the position of the bounding box of object detection from object detection processing in the image to identify inconsistencies between the position of the classification mask and the position of the detected object bounding box; and providing an indication of detected classification inconsistency if the position of the classification mask in the image is inconsistent with the position of the detected object bounding box.
[0142] Example 5. The method according to any one of Examples 2 to 4, wherein performing the plurality of consistency checks on the plurality of image processing outputs includes: performing a depth reasonableness check, the depth reasonableness check comparing the depth estimate of a detected object from an object detection process with the depth estimates of individual pixels or groups of pixels from a depth estimation process to identify a distribution in the depth estimates across pixels of the detected object that is inconsistent with a depth distribution associated with a classification of a mask covering the detected object from a semantic classification process; and providing an indication of detected depth inconsistency in the case that the distribution in the depth estimates across pixels of the detected object differs from the depth distribution associated with the classification of the mask.
[0143] Example 6. The method according to any one of Examples 2 to 5, wherein performing the plurality of consistency checks on the plurality of image processing outputs includes: performing a context consistency check, the context consistency check comparing a depth estimate of a bounding box covering a detected object from an object detection process with a depth estimate of a mask covering the detected object from a semantic segmentation process to determine whether the distribution of the depth estimate of the mask is different from the depth estimate of the bounding box; and providing an indication of detected context inconsistency if the distribution of the depth estimate of the mask is the same as or similar to the distribution of the depth estimate of the bounding box.
[0144] Example 7. The method according to any one of Examples 2 to 6, wherein performing the plurality of consistency checks on the plurality of image processing outputs includes: performing a label consistency check, wherein the label consistency check compares a detected object from an object detection process with a label of the detected object from an object classification process to determine whether the object classification label is consistent with the detected object; and, in the case that the object classification label is inconsistent with the detected object, providing an indication of the detected label inconsistency.
[0145] Example 8. The method according to any one of Examples 2 to 7, wherein performing mitigation actions in response to identifying the attack comprises: adding an indication of inconsistency from each of the plurality of consistency checks to information about each detected object, the information being provided to the autonomous driving system for tracking the detected object.
[0146] Example 9. The method according to any one of Examples 2 to 8, wherein performing mitigation actions in response to identifying the attack includes reporting the detected attack to a remote system.
[0147] As used in this application, the terms "component," "module," "system," etc., are intended to include computer-related entities such as, but not limited to, hardware, firmware, combinations of hardware and software, software, or software being executed, configured to perform specific operations or functions. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a wireless device and the wireless device itself can be referred to as a component. One or more components may reside within a process and / or execution thread, and components may reside on a processor or core and / or be distributed across two or more processors or cores. Furthermore, these components may execute on various non-transitory computer-readable media on which various instructions and / or data structures are stored. Components may communicate via local and / or remote processes, function or procedure calls, electronic signals, data packets, memory read / write, and other known network, computer, processor, and / or process-related communication methods.
[0148] Several different cellular and mobile communication services and standards are available and envisioned for the future, all of which are feasible and benefit from various implementation schemes for reporting the detection of visual attacks on devices. These services and standards include, for example, the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE) systems, 3rd Generation Wireless (3G), 4th Generation Wireless (4G), 5th Generation Wireless (5G), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), 3GSM, General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA) systems (e.g., cdmaOne, CDMA1020TM), Enhanced Data Rate Evolution of GSM (EDGE), Advanced Mobile Phone Systems (AMPS), Digital AMPS (IS-136 / TDMA), Evolved Data Optimization (EV-DO), Digital Enhanced Cordless Telecommunications (DECT), Global Interoperability for Microwave Access (WiMAX), Wireless Local Area Networks (WLAN), Wi-Fi Protected Access I and II (WPA, WPA2), and Integrated Digital Enhanced Network (iDEN). Each of these technologies relates to the transmission and reception of, for example, voice, data, signaling, and / or content messages. It should be understood that any references to terms and / or technical details relating to the various telecommunications standards or technologies are for illustrative purposes only and are not intended to limit the scope of the claims to a particular communication system or technology, unless specifically stated in the language of the claims.
[0149] The various embodiments illustrated and described are provided merely as examples illustrating the various features of the claims. However, the features shown and described with respect to any given embodiment are not necessarily limited to the associated embodiment and may be used or combined with other embodiments shown and described. Furthermore, the claims are not intended to be limited to any single example embodiment.
[0150] The foregoing method descriptions and process flowcharts are provided as illustrative examples only and are not intended to require or imply that the operations of the various embodiments must be performed in the given order. As those skilled in the art will recognize, the operations in the foregoing embodiments can be performed in any order. Words such as “afterward,” “then,” “next,” etc., are not intended to limit the order of operations; these words are used to guide the reader through the description of the method. Furthermore, any reference to singular claim elements (e.g., references using the articles “a,” “an,” or “the”) should not be construed as limiting that element to the singular. Additionally, references to the term “and / or” should be understood to include both conjunctions and antonyms. For example, “A and / or B” means “A and B” as well as “A or B.”
[0151] The various exemplary logic blocks, modules, components, circuits, and algorithmic operations described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and operations have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the claims.
[0152] Hardware for implementing the various exemplary logic units, logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or performed by a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any combination thereof designed to perform the functions described herein. While the general-purpose processor may be a microprocessor, in alternative embodiments, the processing system may use any conventional processor, controller, microcontroller, or state machine to perform operations. The processor may also be implemented as a combination of receiver intelligent objects, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.
[0153] In one or more embodiments, the described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The operation of the methods or algorithms disclosed herein may be implemented in a processor-executable software module or processor-executable instructions, which may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium that can be accessed by a computer or processor. By way of example and without limitation, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EEPROM, FLASH memory, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage intelligent objects, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. The above combinations are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operation of a method or algorithm may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable and / or computer-readable storage medium, which may be incorporated into a computer program product.
[0154] The above description of the disclosed embodiments is provided to enable any person skilled in the art to implement or use the claims. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the claims. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but should be granted the broadest scope consistent with the following claims and the principles and novel features disclosed herein.
Claims
1. A method for detecting visual attacks performed by a processing system on a device, the method comprising: Multiple trained image processing models are used to process images received from the camera of the device to obtain multiple image processing outputs; Multiple consistency checks are performed on the plurality of image processing outputs, wherein the consistency check in the plurality of consistency checks compares each of the plurality of image processing outputs to detect inconsistencies; Attacks on the camera are detected based on the aforementioned inconsistencies; as well as Mitigation actions are performed in response to the identification of the attack.
2. The method of claim 1, wherein using a plurality of trained image processing models to process the images received from the camera of the device to obtain a plurality of image processing outputs comprises: The image is processed using a trained semantic segmentation model to associate a mask of a group of pixels in the image with a classification label. The trained depth estimation model is used to perform depth estimation processing on the image to identify the distance to objects in the image; The image is processed using a trained object detection model to identify objects in the image and define bounding boxes around the identified objects. as well as The image is processed using a trained object classification model to classify the objects in the image.
3. The method of claim 2, wherein performing the plurality of consistency checks on the plurality of image processing outputs comprises: A semantic consistency check is performed, which compares the classification labels associated with the mask from the semantic segmentation process with the bounding boxes of object detection from the object detection process in the image to identify inconsistencies between the mask classification and the detected objects. as well as In response to a mismatch between the mask classification and the detected objects in the image, an indication of the detected classification inconsistency is provided.
4. The method according to claim 3, further comprising: In response to the consistency between the classification label associated with the mask from the semantic segmentation process and the bounding box of the object detection from the object detection process, a position consistency check is performed. The position consistency check compares the position of the classification mask from the semantic segmentation process in the image with the position of the bounding box of the object detection from the object detection process in the image to identify inconsistencies between the position of the classification mask and the position of the detected object bounding box. as well as If the location of the classification mask within the image is inconsistent with the location of the detected object bounding box, an indication of the detected classification inconsistency is provided.
5. The method of claim 2, wherein performing the plurality of consistency checks on the plurality of image processing outputs comprises: A depth reasonableness check is performed, which compares the depth estimate of the detected object from the object detection process with the depth estimate of each pixel or group of pixels from the depth estimation process to identify the distribution in the depth estimate across the pixels of the detected object that is inconsistent with the depth distribution associated with the classification of the mask covering the detected object from the semantic classification process. as well as In cases where the distribution in the depth estimate across pixels of a detected object differs from the depth distribution associated with the mask classification, it provides an indication of inconsistency in the detected depth.
6. The method of claim 2, wherein performing the plurality of consistency checks on the plurality of image processing outputs comprises: A context consistency check is performed, which compares the depth estimates of the bounding boxes covering the detected objects from the object detection process with the depth estimates of the masks covering the detected objects from the semantic segmentation process to determine whether the distribution of the depth estimates of the masks is different from the depth estimates of the bounding boxes. as well as If the distribution of the depth estimate of the mask is the same as or similar to the distribution of the depth estimate of the bounding box, an indication of detected contextual inconsistency is provided.
7. The method of claim 2, wherein performing the plurality of consistency checks on the plurality of image processing outputs comprises: Perform a label consistency check, which compares the detected objects from the object detection process with the labels of the detected objects from the object classification process to determine whether the object classification labels are consistent with the detected objects. as well as If the object classification label is inconsistent with the detected object, an indication of the detected label inconsistency is provided.
8. The method of claim 1, wherein performing mitigation actions in response to identifying the attack comprises: Indications of inconsistency from each of the plurality of consistency checks are added to information about each detected object, which is provided to the autonomous driving system for tracking the detected objects.
9. The method of claim 1, wherein performing mitigation actions in response to identifying the attack includes reporting the detected attack to a remote system.
10. An apparatus comprising: A processing system, the processing system including one or more processors configured to perform the following operations: Multiple trained image processing models are used to process images received from the camera of the device to obtain multiple image processing outputs; Multiple consistency checks are performed on the plurality of image processing outputs, wherein the consistency check in the plurality of consistency checks compares each of the plurality of image processing outputs to detect inconsistencies; Attacks on the camera are detected based on the aforementioned inconsistencies; as well as Mitigation actions are performed in response to the identification of the attack.
11. The apparatus of claim 10, wherein, in order to process the image received from the camera of the apparatus, the one or more processors are further configured to: The image is processed using a trained semantic segmentation model to associate a mask of a group of pixels in the image with a classification label. The trained depth estimation model is used to perform depth estimation processing on the image to identify the distance to objects in the image; The image is processed using a trained object detection model to identify objects in the image and define bounding boxes around the identified objects. as well as The image is processed using a trained object classification model to classify the objects in the image.
12. The apparatus of claim 11, wherein the one or more processors are further configured to perform the plurality of consistency checks on the plurality of image processing outputs, and the one or more processors are further configured to: A semantic consistency check is performed, which compares the classification labels associated with the mask from the semantic segmentation process with the bounding boxes of object detections from the object detection process in the image to identify inconsistencies between the mask classification and the detected objects; and In response to a mismatch between the mask classification and the detected objects in the image, an indication of the detected classification inconsistency is provided.
13. The apparatus of claim 12, wherein in response to a classification label associated with a mask from the semantic segmentation process coinciding with a bounding box of an object detection from the object detection process, the one or more processors are further configured to: A positional consistency check is performed, which compares the position of the classification mask from the semantic segmentation process within the image with the position of the bounding box of the object detection process within the image to identify inconsistencies between the position of the classification mask and the position of the detected object bounding box; and If the location of the classification mask within the image is inconsistent with the location of the detected object bounding box, an indication of the detected classification inconsistency is provided.
14. The apparatus of claim 11, wherein, in order to perform the plurality of consistency checks on the plurality of image processing outputs, the one or more processors are further configured to: A depth reasonableness check is performed, which compares the depth estimate of the detected object from the object detection process with the depth estimates of individual pixels or groups of pixels from the depth estimation process to identify a distribution in the depth estimates across pixels of the detected object that is inconsistent with the depth distribution associated with the classification of the mask covering the detected object from the semantic classification process; and In cases where the distribution in the depth estimate across pixels of a detected object differs from the depth distribution associated with the mask classification, it provides an indication of inconsistency in the detected depth.
15. The apparatus of claim 11, wherein, in order to perform the plurality of consistency checks on the plurality of image processing outputs, the one or more processors are further configured to: Perform a context consistency check, which compares the depth estimates of bounding boxes covering detected objects from the object detection process with the depth estimates of masks covering the detected objects from the semantic segmentation process to determine whether the distribution of the mask depth estimates differs from the depth estimates of the bounding boxes; and If the distribution of the depth estimate of the mask is the same as or similar to the distribution of the depth estimate of the bounding box, an indication of detected contextual inconsistency is provided.
16. The apparatus of claim 11, wherein, in order to perform the plurality of consistency checks on the plurality of image processing outputs, the one or more processors are further configured to: Perform a label consistency check, which compares the detected objects from the object detection process with the labels of the detected objects from the object classification process to determine whether the object classification labels are consistent with the detected objects; and If the object classification label is inconsistent with the detected object, an indication of the detected label inconsistency is provided.
17. The apparatus of claim 10, wherein the one or more processors are further configured to perform mitigation actions in response to identifying the attack, the mitigation actions adding an indication of inconsistency from each of the plurality of consistency checks to information about each detected object, the information being provided to the autonomous driving system for tracking the detected objects.
18. The apparatus of claim 10, wherein the one or more processors are further configured to perform mitigation actions in response to identifying the attack, the mitigation actions reporting the detected attack to a remote system.
19. A non-transitory processor-readable medium storing processor-executable instructions configured to cause a processing system of a device to perform operations, the operations including: Multiple trained image processing models are used to process images received from the camera of the device to obtain multiple image processing outputs; Multiple consistency checks are performed on the plurality of image processing outputs, wherein the consistency check in the plurality of consistency checks compares each of the plurality of image processing outputs to detect inconsistencies; Attacks on the camera are detected based on the aforementioned inconsistencies; as well as Mitigation actions are performed in response to the identification of the attack.
20. The non-transitory processor-readable medium of claim 19, wherein the processor-executable instructions are further configured to cause the processing system to perform operations such that processing the images received from the camera of the device using a plurality of trained image processing models to obtain a plurality of image processing outputs includes: The image is processed using a trained semantic segmentation model to associate a mask of a group of pixels in the image with a classification label. The trained depth estimation model is used to perform depth estimation processing on the image to identify the distance to objects in the image; The image is processed using a trained object detection model to identify objects in the image and define bounding boxes around the identified objects. as well as The image is processed using a trained object classification model to classify the objects in the image.