Defence of multi-modal fusion model against single-source attack
By introducing an outlier removal network and a robust feature fusion layer into the multimodal learning model, the problem of insufficient robustness of the multimodal model under single-source adversarial perturbation is solved, and significant performance improvement is achieved in multiple tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2022-06-16
- Publication Date
- 2026-07-24
AI Technical Summary
Existing multimodal learning models are not robust enough to single-source adversarial perturbations and are easily attacked by a single modality, leading to model failure. They cannot effectively utilize information from multiple undisturbed modalities to make correct predictions.
A combined strategy of outlier removal network and robust feature fusion layer is adopted. The model is trained by outlier removal learning and adaptive gating strategy to detect inconsistencies between modes and allow only undisturbed modal information to pass through during the fusion process, thereby improving the single-source adversarial robustness of the model.
It significantly improves the robustness of multimodal models in action recognition, object detection and sentiment analysis tasks, with gains of 7.8-25.2%, 19.7-48.2% and 1.6-6.7%, respectively, while maintaining performance on clean data.
Smart Images

Figure CN115482442B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to robust multimodal machine learning systems. More specifically, this application relates to improving the robustness of multimodal machine learning systems by training and using an outlier removal network (odd-one-out network) with robust fusion layers. Background Technology
[0002] In the real world, information can be captured and represented by different modalities. For example, a set of pixels in an image can be associated with labels and textual interpretations; sound can be associated with vibrations caused by speed, operating conditions, or environmental conditions; and ultrasound can be associated with distance, size, and density. Different modalities can be characterized by very different statistical properties. For example, images are often represented as pixel intensities or outputs of a feature extractor, while sound can be a time series, and ultrasound can produce point clouds. Due to the different statistical properties of different information resources, discovering relationships between different modalities is crucial. Multimodal learning is a good model for representing joint representations of different modalities. Multimodal learning models are also able to fill in missing modalities while taking into account the observed modalities. Summary of the Invention
[0003] A multimodal sensing system includes a controller. The controller can be configured to receive a first signal from a first sensor, a second signal from a second sensor, and a third signal from a third sensor; extract a first feature vector from the first signal, a second feature vector from the second signal, and a third feature vector from the third signal; determine an odd-one-out vector from the first, second, and third feature vectors via an outlier removal network of a machine learning network based on inconsistent modality prediction; fuse the first, second, and third feature vectors and the odd-one-out vector into a fused feature vector; and output the fused feature vector.
[0004] A multimodal sensing method includes receiving a first signal from a first sensor, a second signal from a second sensor, and a third signal from a third sensor; extracting a first feature vector from the first signal, extracting a second feature vector from the second signal, and extracting a third feature vector from the third signal; determining an outlier vector from the first, second, and third feature vectors via an outlier removal network of a machine learning network based on inconsistent modal prediction; fusing the first, second, and third feature vectors and the outlier vector into a fused feature vector; and outputting the fused feature vector.
[0005] A multimodal perception system for autonomous vehicles includes a first sensor and a controller. The first sensor is one of a video, RADAR, LIDAR, or ultrasonic sensor. The controller can be configured to receive a first signal from the first sensor, a second signal from a second sensor, and a third signal from a third sensor; extract a first feature vector from the first signal, a second feature vector from the second signal, and a third feature vector from the third signal; determine an outlier vectors from the first, second, and third feature vectors via an outlier removal network of a machine learning network based on inconsistent modal predictions; fuse the first, second, and third feature vectors and the outlier vectors into a fused feature vector; output the fused feature vector; and control the autonomous vehicle based on the fused feature vector. Attached Figure Description
[0006] Figure 1 This is a block diagram of a system used to train a neural network.
[0007] Figure 2 It is a graphical representation of an exemplary single-source adversarial perturbation on a multimodal model with both vulnerable and robust outputs.
[0008] Figure 3 This is a block diagram of a data annotation system that utilizes machine learning models.
[0009] Figure 4 It is a graphical representation of a multimodal hybrid network.
[0010] Figure 5 This is a block diagram of an electronic computing system.
[0011] Figure 6 It is a graphical representation of a multimodal fusion network with an outlier removal network.
[0012] Figure 7 It is a graphical representation of the outlier removal network.
[0013] Figure 8 It is a graphical representation of a robust feature fusion layer with inputs that have outliers removed.
[0014] Figure 9 This is a flowchart of a robust training strategy for feature fusion and outlier removal networks.
[0015] Figure 10A This is a graphical representation of an example action recognition result.
[0016] Figure 10B This is a graphical representation of an exemplary two-dimensional object detection result.
[0017] Figure 10CThis is a graphical representation of an example sentiment analysis result.
[0018] Figure 11 This is a schematic diagram of a control system configured to control a vehicle.
[0019] Figure 12 This is a schematic diagram of a control system configured to control manufacturing machines.
[0020] Figure 13 This is a schematic diagram of a control system configured to control power tools.
[0021] Figure 14 This is a schematic diagram of a control system configured to control an automated personal assistant.
[0022] Figure 15 This is a schematic diagram of a control system configured as a control and monitoring system.
[0023] Figure 16 This is a schematic diagram of a control system configured to control a medical imaging system. Detailed Implementation
[0024] Detailed embodiments of the invention are disclosed herein as needed; however, it should be understood that the disclosed embodiments are merely examples of the invention, which may be implemented in various and alternative forms. The drawings are not necessarily to scale; some features may be exaggerated or minimized to show detail of specific components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but only as a representative basis for teaching those skilled in the art to use the invention in various ways.
[0025] The term "substantially" may be used herein to describe the disclosed or claimed embodiments. The term "substantially" may modify values or relative characteristics disclosed or claimed in this disclosure. In such cases, "substantially" may mean that the modified value or relative characteristic is within ±0%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, or 10% of that value or relative characteristic.
[0026] The term sensor refers to a device that detects or measures physical properties and records, indicates, or otherwise responds to those properties. Sensors include optical, light, imaging, or photonic sensors (e.g., charge-coupled devices (CCDs), CMOS active pixel sensors (APS), infrared (IR) sensors, CMOS sensors), acoustic, sound, or vibration sensors (e.g., microphones, seismic detectors, hydrophones), automotive sensors (e.g., wheel speed, parking, radar, oxygen, blind spot, torque), chemical sensors (e.g., ion-sensitive field-effect transistors (ISFETs), oxygen, carbon dioxide, chemiluminescent resistors, holographic sensors), current, potential, magnetic, or radio frequency sensors (e.g., Hall effect, magnetometers, magnetoresistive, Faraday cups, galvanometers), environmental, weather, humidity, or moisture sensors (e.g., weather radar, solar meters), and flow or fluid velocity sensors. Examples include mass airflow sensors, anemometers, ionizing radiation or subatomic particle sensors (e.g., ionization chambers, Geiger counters, neutron detectors), navigation sensors (e.g., Global Positioning System (GPS) sensors, magnetohydrodynamic (MHD) sensors), position, angle, displacement, distance, velocity or acceleration sensors (e.g., LiDAR, accelerometers, ultra-wideband radar, piezoelectric sensors), force, density or level sensors (e.g., strain gauges, nuclear density meters), thermal sensors, heat sensors or temperature sensors (e.g., infrared thermometers, pyrometers, thermocouples, thermistors, microwave radiometers), or other devices, modules, machines or subsystems whose purpose is to detect or measure physical properties and record, indicate or otherwise respond to them.
[0027] In addition to achieving high performance in many vision tasks, multimodal models are expected to be robust to single-source failures due to the availability of redundant information between modalities. This disclosure provides a solution for the robustness of multimodal neural networks against worst-case (i.e., adversarial) perturbations on a single modality. This disclosure will illustrate that standard multimodal fusion models are vulnerable to single-source adversarial attacks; for example, an attack on any single modality can overcome correct information from multiple undisturbed modalities and cause the model to fail. This unexpected weakness exists in a variety of multimodal tasks and requires solutions. This disclosure proposes an adversarially robust fusion strategy that trains the model to compare information from all input sources, detect inconsistencies in perturbed modalities compared to other modalities, and only allow information from undisturbed modalities to pass through. This method significantly improves the single-source robustness of existing methods, achieving gains of 7.8–25.2% for action recognition, 19.7–48.2% for object detection, and 1.6–6.7% for sentiment analysis, without compromising performance on undisturbed (i.e., clean) data based on experimental results.
[0028] Figure 1 A system 100 for training a neural network is shown. System 100 may include an input interface for accessing training data 192 of the neural network. For example, as... Figure 1 As shown, the input interface can be comprised of a data storage interface 180 that allows access to training data 192 from data storage device 190. For example, data storage interface 180 can be a memory interface or permanent storage interface, such as a hard disk or SSD interface, but it can also be a personal, local area, or wide area network interface, such as a Bluetooth, Zigbee, or Wi-Fi interface, or an Ethernet or fiber optic interface. Data storage device 190 can be an internal data storage device of system 100, such as a hard disk drive or SSD, but it can also be an external data storage device, such as a network-accessible data storage device.
[0029] In some embodiments, data storage device 190 may further include a data representation 194 of an untrained variant of the neural network, which may be accessed by system 100 from data storage device 190. However, it will be understood that the training data 192 and data representation 194 of the untrained neural network may also be accessed from different data storage devices, for example, via different subsystems of data storage interface 180. Each subsystem may be of the type described above for data storage interface 180. In other embodiments, the data representation 194 of the untrained neural network may be generated internally by system 100 based on the design parameters of the neural network, and therefore may not be explicitly stored on data storage device 190. System 100 may also include a processor subsystem 160, which may be configured to provide an iterative function as an alternative to a stack of layers of the neural network to be trained during operation of system 100. In one embodiment, the corresponding layers of the replaced layer stack may have mutually shared weights and may receive the output of the previous layer or, with respect to the first layer of the layer stack, receive an initial activation and a portion of the input of the layer stack as input. The system may also include multiple layers. Processor subsystem 160 can also be configured to iteratively train a neural network using training data 192. Here, the training iterations of processor subsystem 160 may include a forward propagation portion and a backward propagation portion. Processor subsystem 160 can be configured to perform the forward propagation portion by determining an equilibrium point of the iterative function that converges to a fixed point and by providing the equilibrium point as an alternative to the output of the layer stack in the neural network, wherein determining the equilibrium point includes using a numerical root-finding algorithm to find the root solution of the iterative function minus its input, and other operations that define the forward propagation portion that can be performed. System 100 may also include an output interface for outputting a data representation 196 of the trained neural network, which may also be referred to as trained model data 196. For example, as well as Figure 1As shown, the output interface can be comprised of a data storage interface 180, which in these embodiments is an input / output (“IO”) interface through which trained model data 196 can be stored in the data storage device 190. For example, the data representation 194 defining an “untrained” neural network can be at least partially replaced by the data representation 196 of a trained neural network during or after training, because the parameters of the neural network, such as the network weights, hyperparameters, and other types of parameters, can be adapted to reflect training on the training data 192. This also... Figure 1 The same data record on data storage device 190 is referred to by reference numerals 194 and 196. In other embodiments, data representation 196 may be stored separately from data representation 194 defining an "untrained" neural network. In some embodiments, the output interface may be separate from data storage interface 180, but it can typically be of the type described above for data storage interface 180.
[0030] Figure 2 This is a graphical representation 200 of an exemplary single-source adversarial perturbation on a multimodal model with vulnerable and robust output. A scene 202 of a truck traveling along a road is analyzed through different modalities 204. In this example, the different modalities include a video camera 204a, a LiDAR sensor 204b, and a microphone 204c. Data from the different modalities is processed by a processor or controller in a multimodal model 206, and a prediction of scene 208 is output, which can be used to control systems such as robotic systems, autonomous vehicles, industrial systems, or other electrical / electromechanical systems. If an adversarial perturbation occurs to one of the modalities (e.g., video camera 204a), the prediction 206 of the scene may be an inaccurate prediction 206a. However, a robust prediction 206b of the truck can be produced using the robust multimodal model 206, even with the presence of adversarial perturbations to video camera 204a. This disclosure will propose a system and method for generating robust predictions in the face of adversarial perturbations to modalities, as well as a system and method for training a robust multimodal model.
[0031] Figure 3A data annotation system 300 is depicted that implements a system for annotating data. The data annotation system 300 may include at least one computing system 302. The computing system 302 may include at least one processor 304 operatively connected to a memory unit 308. The processor 304 may include one or more integrated circuits that implement the functions of a central processing unit (CPU) 306. The CPU 306 may be a commercially available processing unit that implements an instruction set, such as one of the x86, ARM, Power, or MIPS instruction set families. During operation, the CPU 306 may execute stored program instructions retrieved from the memory unit 308. The stored program instructions may include software that controls the operation of the CPU 306 to perform the operations described herein. In some examples, the processor 304 may be a system-on-a-chip (SoC) that integrates the functions of the CPU 306, memory unit 308, network interface, and input / output interface into a single integrated device. The computing system 302 may implement an operating system for managing various aspects of the operation.
[0032] Memory cell 308 may include volatile and non-volatile memory for storing instructions and data. Non-volatile memory may include solid-state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 302 is deactivated or loses power. Volatile memory may include static and dynamic random access memory (RAM) for storing program instructions and data. For example, memory cell 308 may store a machine learning model 310 or algorithm, a training dataset 312 for the machine learning model 310, and an original source dataset 315. Model 310 may include as described in this disclosure and Figure 7 The anomaly removal network is illustrated in the figure. Furthermore, the training dataset 312 may include, as described in this disclosure, and... Figure 4 , Figure 6 , Figure 7 and Figure 8 The illustration shows the features and feature extractor. Furthermore, the original source 315 may include features derived from those described in this disclosure. Figure 4 and Figure 6 The diagram shows the data for multiple input modalities.
[0033] The computing system 302 may include a network interface device 322 configured to provide communication with external systems and devices. For example, the network interface device 322 may include wired and / or wireless Ethernet interfaces defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 322 may include a cellular communication interface for communicating with cellular networks (e.g., 3G, 4G, 5G). The network interface device 322 may be further configured to provide a communication interface to an external network 324 or the cloud.
[0034] External network 324 may be referred to as the World Wide Web or the Internet. External network 324 can establish standard communication protocols between computing devices. External network 324 can allow information and data to be easily exchanged between computing devices and the network. One or more servers 330 can communicate with external network 324.
[0035] The computing system 302 may include an input / output (I / O) interface 320, which may be configured to provide digital and / or analog inputs and outputs. The I / O interface 320 may include an additional serial interface (e.g., a Universal Serial Bus (USB) interface) for communicating with external devices.
[0036] The computing system 302 may include a human-machine interface (HMI) device 318, which may include any device that enables the system 300 to receive control input. Examples of input devices may include HMI inputs such as a keyboard, mouse, touchscreen, voice input device, and other similar devices. The computing system 302 may include a display device 332. The computing system 302 may include hardware and software for outputting graphical and textual information to the display device 332. The display device 332 may include an electronic display screen, projector, printer, or other suitable device for displaying information to a user or operator. The computing system 302 may also be configured to allow interaction with remote HMIs and remote display devices via a network interface device 322.
[0037] System 300 can be implemented using one or more computing systems. While this example depicts a single computing system 302 implementing all the described features, the various features and functions can be decoupled and implemented by multiple computing units that communicate with each other. The specific system architecture chosen can depend on a variety of factors.
[0038] System 300 can implement machine learning algorithm 310 configured to analyze raw source dataset 315. Raw source dataset 315 may include raw or unprocessed sensor data, which may represent the input dataset for the machine learning system. Raw source dataset 315 may include video, video clips, images, text-based information, and raw or partially processed sensor data (e.g., radar maps of objects). In some examples, machine learning algorithm 310 may be a neural network algorithm designed to perform a predetermined function. For example, a neural network algorithm may be configured in automotive applications to identify pedestrians in video images.
[0039] Computer system 300 may store a training dataset 312 for machine learning algorithm 310. Training dataset 312 may represent a set of previously constructed data used to train machine learning algorithm 310. Training dataset 312 may be used by machine learning algorithm 310 to learn weight factors associated with the neural network algorithm. Training dataset 312 may include source datasets containing the corresponding outputs or results that machine learning algorithm 310 attempts to replicate via the learning process. In this example, training dataset 312 may include source videos with and without pedestrians, along with corresponding presence and location information. The source videos may include various scenes in which pedestrians are identified.
[0040] Machine learning algorithm 310 can operate in learning mode using training dataset 312 as input. Machine learning algorithm 310 can be executed in multiple iterations using data from training dataset 312. For each iteration, machine learning algorithm 310 can update its internal weight factors based on the obtained results. For example, machine learning algorithm 310 can compare its output results (e.g., annotations) with those included in training dataset 312. Since training dataset 312 includes expected results, machine learning algorithm 310 can determine when performance is acceptable. After machine learning algorithm 310 reaches a predetermined performance level (e.g., 100% consistency with results associated with training dataset 312), machine learning algorithm 310 can be executed using data not in training dataset 312. The trained machine learning algorithm 310 can be applied to new datasets to generate annotated data.
[0041] Machine learning algorithm 310 can be configured to identify specific features in raw source data 315. Raw source data 315 may include multiple instances or input datasets for which results need to be annotated. For example, machine learning algorithm 310 can be configured to identify the presence of a pedestrian in a video image and annotate that presence. Machine learning algorithm 310 can be programmed to process raw source data 315 to identify the presence of specific features. Machine learning algorithm 310 can be configured to identify features in raw source data 315 as predetermined features (e.g., pedestrians). Raw source data 315 can be derived from various sources. For example, raw source data 315 may be actual input data collected by a machine learning system. Raw source data 315 may be machine-generated for testing a system. As an example, raw source data 315 may include raw video images from a camera.
[0042] In this example, machine learning algorithm 310 can process the raw source data 315 and output an indication of the image representation. The output may also include an enhanced representation of the image. Machine learning algorithm 310 can generate a confidence level or factor for each generated output. For example, a confidence value exceeding a predetermined high confidence threshold may indicate that machine learning algorithm 310 is confident that the identified feature corresponds to a specific feature. A confidence value below a low confidence threshold may indicate that machine learning algorithm 310 has some uncertainty regarding the existence of a specific feature.
[0043] Figure 4 This is a graphical representation of a multimodal fusion system 400. The multimodal fusion network 402 receives input modes 404a, 404b, and 404c, extracts features 406a, 406b, and 406c from each mode, and mixes them in a fusion layer 408 and subsequent downstream layers 410 to produce an output. This multimodal fusion system 400 can be implemented on an electronic computing system. The system 400 operates well under ideal conditions; however, if one of the modes experiences an adversarial perturbation (e.g., input mode 404b), the system may provide an invalid output.
[0044] Example machine architecture and machine-readable media. Figure 5 It is a block diagram of an electronic computing system suitable for implementing the system or for performing the methods disclosed herein. Figure 5 The machines are shown as standalone devices suitable for implementing the concepts within this disclosure. Regarding the servers described above, multiple such machines operating within data centers, as part of a cloud architecture, etc., can be used. In the server aspect, not all shown functions and devices are utilized. For example, while systems, devices, etc., used by users to interact with servers and / or cloud architectures may have screens, touchscreen inputs, etc., servers typically do not have screens, touchscreens, cameras, etc., and typically interact with users through connection systems with appropriate input and output aspects. Therefore, the architecture described below should be considered to include multiple types of devices and machines, and aspects may or may not be present in any particular device or machine, depending on their form factor and purpose (e.g., servers rarely have cameras, and wearable devices rarely include disks). However, Figure 5 The exemplary explanations are adapted to allow those skilled in the art to determine how the previously described embodiments can be implemented through appropriate combinations of hardware and software, or through appropriate modifications to the illustrated embodiments of the particular devices, machines, etc. used.
[0045] Although only a single machine is shown, the term "machine" should also be understood to include any collection of machines that individually or jointly execute a set (or more) of instructions to perform any one or more of the methods discussed herein.
[0046] Examples of machine 500 include at least one processor 502 (e.g., a controller, microcontroller, central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), advanced processing unit (APU), or a combination thereof), and one or more memories, such as main memory 504, static memory 506, or other types of memory, which communicate with each other via link 508. Link 508 may be a bus or other type of connection channel. Machine 500 may include other optional aspects, such as a graphics display unit 510 including any type of display. Machine 500 may also include other optional aspects such as alphanumeric input device 512 (e.g., keyboard, touchscreen, etc.), user interface (UI) navigation device 514 (e.g., mouse, trackball, touch device, etc.), storage unit 516 (e.g., disk drive or other storage device), signal generation device 518 (e.g., speaker), sensor 521 (e.g., GPS sensor, accelerometer, microphone, camera, etc.), output controller 528 (e.g., wired or wireless connection for connecting to and / or communicating with one or more other devices, such as Universal Serial Bus (USB), Near Field Communication (NFC), Infrared (IR), serial / parallel bus, etc.), and network interface device 520 (e.g., wired and / or wireless) for connecting to and / or communicating through one or more networks.
[0047] Various memories (i.e., the memories of 504, 506, and / or the processor 502) and / or storage units 516 may store one or more sets of instructions and data structures (e.g., software) 524 that embody or are utilized by any or more of the methods or functions described herein. These instructions, when executed by one or more processors 502, cause various operations to implement the disclosed embodiments.
[0048] Figure 6 This is a graphical representation of a multimodal fusion system with an outlier removal network. The multimodal fusion network 602 receives input modalities 604a, 604b, and 604c, and extracts features 606a, 606b, and 606c from each modality; these features are feature vectors. The output of the feature extractor 606 is fed into the outlier removal network 612. The outlier removal network 612 generates "inconsistent" modality predictions, which, along with the output of the feature extractor 606, are fed into a robust fusion layer 608. The robust fusion layer 608 outputs a fused feature vector, which is then fed into a downstream layer 610 to produce an output. This multimodal fusion system 600 can be implemented on an electronic computing system.
[0049] Figure 7 Such as Figure 6 A graphical representation of the outlier removal network 700 of the outlier removal network 612 is provided. Network 700 receives features 702, such as the outputs from feature extractors 602a, 602b, and 602c, and generates modality prediction weights 704, such that modality prediction weights 704a, 704b, and 704c are associated for each feature channel. These modality prediction weights 704a, 704b, and 704c produce an outlier removal vector, which is forwarded to the robust feature fusion layer.
[0050] Figure 8 This is a graphical representation of a robust feature fusion layer 800 with inputs that have outliers removed. The fusion layer 800 receives features 802 from each modality and performs fusion 804 on each modality to produce fused features 806 for each modality. The fused features 806 are then combined with features from... Figure 7 Modal prediction 704 is fused to produce output.
[0051] consider Figure 2 The multimodal neural network shown here fuses data from... k The system receives input from several different sources to identify objects in the autonomous driving system. If one of the modalities (e.g., a red-green-blue camera) receives a worst-case or adversarial disturbance, the model fails to detect the truck in the scene. Alternatively, the model uses the remaining k-1 Make robust predictions using unperturbed modes (e.g., LiDAR sensors, audio microphones, etc.). This example illustrates the importance of single-source adversarial robustness in avoiding catastrophic failures in real-world multimodal systems. In practical settings, any single mode may be affected by worst-case perturbations, while multiple modes typically do not fail simultaneously, especially when physical sensors are not coupled.
[0052] In the field of adversarial robustness, most research has focused on unimodal settings rather than multimodal settings. An effective strategy for defending unimodal models against adversarial attacks is adversarial training. adversarial training (i.e., end-to-end training of the model on adversarial examples). In principle, adversarial training can also be extended to multimodal models, but it has several disadvantages: (1) it is resource-intensive and may not scale well to large multimodal models with more parameters than their unimodal counterparts; (2) it significantly degrades performance on clean data. For these reasons, end-to-end adversarial training may be impractical for multimodal systems used in real-world tasks.
[0053] This disclosure has three modes ( k=3Multimodal robustness against single-source adversarial perturbations is demonstrated on different benchmark tasks: action recognition in EPIC-Kitchens, object detection in KITTI, and sentiment analysis in CMU-MOSI. While this disclosure uses three modalities as examples, it is not limited to three modalities and can be extended to more than three. This disclosure will illustrate that standard multimodal fusion practices are susceptible to single-source adversarial perturbations. Even when multiple unperturbed modalities exist that can produce correct predictions, the natural confusion between features from the perturbed modality and features from the clean modality may not automatically produce robust predictions. Figure 4 As shown, in a multimodal model, the worst-case input to any single modality can be more significant than that to other modalities, leading to model failure. In fact, contrary to expectations, in some cases, under the same attack, the multimodal model under a single-source perturbation... k = 3 ) is not superior to the single-mode model ( k = 1 ).
[0054] This disclosure proposes an adversarially robust fusion strategy applicable from mid-stage to late-stage fusion models to defend against this vulnerability without compromising clean performance. It is based on the assumption that a multimodal model can be trained to detect correspondences (or lack thereof) between features from different modalities, and this information is used to perform robust feature fusion against the perturbed modality. The method extends existing work on adaptive gating strategies by leveraging a robust fusion training process based on Odd-One-Out Learning to improve single-source adversarial robustness without compromising clean performance. Extensive experiments demonstrate that the method is even effective against adaptive white-box attacks using the robust fusion strategy. Exemplary embodiments of the system significantly outperform prior art methods in terms of single-source robustness. Test results of the exemplary system and method show gains of 7.8–25.2% on action recognition for EPIC Kitchens, 19.7–48.2% on 2D object detection for KITTI, and 1.6–6.7% on sentiment analysis for CMU-MOSI.
[0055] Multimodal models are typically not inherently robust to single-source adversarial attacks, but this disclosure demonstrates how to improve the robustness of multimodal models without the drawbacks associated with end-to-end adversarial training in unimodal models. The combination of robust fusion architecture and robust fusion training could be a practical strategy for defending real-world systems against adversarial attacks and opens up promising directions for future research.
[0056] Adversarial Robustness. Deep learning-based vision systems are vulnerable to adversarial attacks, which are additional, worst-case, and imperceptible perturbations to the input that lead to incorrect predictions. Many defenses against adversarial attacks have been proposed, with two of the most effective being end-to-end adversarial training, which synthesizes adversarial examples and includes them in the training data, and demonstrably robust training, which provides theoretical constraints on performance. However, these methods focus on unimodal settings where the input is a single image. In contrast to these works, we consider single-source adversarial perturbations in multimodal settings and leverage consistent information between modalities to improve the robustness of the model fusion step. This training process is related to adversarial training in a sense, as it also uses perturbed inputs, but instead of end-to-end training of model parameters, the focus is on robustly designing and training feature fusion. This strategy yields the benefits of adversarial training while maintaining performance on clean data and significantly reducing the number of parameters that need to be trained on perturbed data.
[0057] Multimodal fusion models. Multimodal neural networks exhibit good performance on various visual tasks, such as scene understanding, object detection, sentiment analysis, speech recognition, and medical imaging. In terms of fusion methods, gating networks adaptively weight the sources based on the input. These fusion methods leverage multiple modalities to improve clean task execution, but do not evaluate or extend these methods to improve single-source robustness, which is one of the focuses of this disclosure.
[0058] Single-source robustness. Recent works have provided robustness against single-source damage (e.g., occlusion, missing data, and Gaussian noise) with two modes ( k=2 This disclosure provides an important understanding of the impact of object detection systems on object detection. In contrast, this disclosure considers single-source adversarial perturbations, exploring the worst-case failure of multimodal systems due to a single perturbation mode. This disclosure considers tasks other than object detection and utilizes three modes (…). k=3 The evaluation model is used where there are more clean sources than perturbation sources. Regarding defense strategies, a robust multimodal fusion method based on end-to-end robust training and an adaptive gated fusion layer improves robustness against single-source perturbations. This disclosure extends this by developing a robust fusion strategy that balances the correspondence between undisturbed modes to defend against perturbed modes and effectively resists more challenging adversarial perturbations.
[0059] Single-source counter-perturbation.
[0060] set up Indicates having k One input mode (i.e., A multimodal model. Consider... f Performance due to any single mode (in The worst-case perturbation on ) reduces to the level of, while others k- 1 The modes remain undisturbed. Therefore, this will be addressed regarding the modes. i On f The single-source counter-perturbation is defined as Equation 1.
[0061]
[0062] in It is a loss function, and The perturbation was defined The allowable range. Assuming multimodal input x and output y are sampled from distribution D, then... f Relative to mode The single-source adversarial performance is given by the following formula:
[0063]
[0064] f represents the performance of undisturbed data, i.e. The average difference between it and the single-source countermeasure performance specified in equation (2) is expressed as f For its modality i The worst-case sensitivity of the input. Ideally, a multimodal model that can access multiple input modalities with redundant information should not be sensitive to perturbations on a single input; it should be able to utilize the remaining... k-1 Unperturbed modes are used to make correct predictions. However, it can be seen that standard multimodal fusion models are surprisingly susceptible to these perturbations in various multimodal benchmark tasks, even when the number of clean modes exceeds the number of perturbed modes. Experiments and results are provided later in this disclosure, but this weakness necessitates a solution.
[0065] Anti-robust fusion strategies.
[0066] set up It is a standard multimodal neural network, pre-trained to achieve acceptable performance on undisturbed data, i.e., it minimizes... The robust fusion strategy disclosed in this paper aims to improve performance by utilizing the correspondence between undisturbed modes to detect and defend against disturbed modes. Single-source robustness. Assume... It features a mid-to-late stage fusion architecture, which consists of modality-specific feature extractors applied to their respective modalities. And the composition of the fusion subnetwork h:
[0067]
[0068] In order to To enhance robustness, it is equipped with an auxiliary outlier removal network and a robust feature fusion layer, replacing the default feature fusion operation, such as... Figure 2 As shown in the diagram. Robust training is then performed based on outlier removal learning and adversarial training focused on these new modules. The outlier removal network is trained when presented with feature representations of different modalities (e.g., outlier removal learning). o To detect inconsistent or perturbed modalities, the robust feature fusion layer uses a multimodal fusion operation with different output sets from the outlier removal network, ensuring that only consistent modalities are passed to downstream layers (e.g., the robust feature fusion layer). The fusion subnetwork equipped with the robust feature fusion layer is then used. h Represented as And represent the complete, enhanced multimodal model as As shown in Equation 4,
[0069]
[0070] Finally, the outlier removal network was jointly trained. o and fusion sub-network While maintaining the feature extractor Weights and architecture from Fixed (e.g., robust training process).
[0071] Learning by eliminating outliers.
[0072] Outlier removal learning aims to learn from a set of originally consistent elements (e.g., Figure 7 This is a self-supervised task to identify inconsistent elements in a multimodal model. To leverage shared information between modalities, an outlier removal network is used to augment the multimodal model. Given a feature set extracted from k modal inputs... The outlier removal network predicts whether multimodal features are consistent with each other (i.e., all inputs are clean) or whether one modality is inconsistent with the others (i.e., some inputs have been perturbed). To perform this task, the outlier removal network must compare features from different modalities, identify shared information among them, and detect any modality inconsistent with the others. For convenience, the features are treated as feature extractor networks applied to their respective modalities. The final output. However, in principle, these features can also come from any intermediate layer of the feature extractor.
[0073] Specifically, the outlier removal network maps the feature z to a size of k+1 Neural networks of vectors o ,like Figure 7 As shown in the diagram. The i-th entry of this vector indicates the mode. i The probability that has been disturbed, i.e. Inconsistent with other features. The vector's first... k+1 Each entry indicates the probability that no mode is perturbed. The outlier removal network is trained by minimizing the following cross-entropy loss. o To perform outlier removal prediction:
[0074]
[0075] in It is the perturbation input generated during training. Features extracted from [the data].
[0076] Robust feature fusion layer.
[0077] To make the outlier removal network o The output is integrated into a multimodal model, considering a feature fusion layer excited by an expert mixture layer (e.g., Figure 8 This layer consists of... k +1 Feature fusion operation The set consists of, where each operation specifically excludes a mode, such as Figure 8 As shown in the diagram. Formally, each fusion operation takes a multimodal feature z as input and performs the fusion of a subset of features as follows:
[0078]
[0079] in This represents the concatenation operation, and NN represents a shallow neural network. By definition, Responsible for executing from except for those targeting i The fusion of features from all modalities except those mentioned above, and only Features from all modalities are fused.
[0080]
[0081] If features Because it is generated by a disturbance input and is different from other... k-1 If the modal characteristics are inconsistent, then The output reception of the outlier removal network has more weights than other fusion operations:
[0082]
[0083] Then by merging subnetworks h This robust feature fusion layer is used to form a robust fusion subnetwork. Then training. and o To optimize clean performance, as shown in Equation 8,
[0084]
[0085] And single-source robustness, as expressed in Equation 9.
[0086]
[0087] For each mode, where It is the input from the perturbation generated during training. Extracted features. Note the connection to the fusion network. One of the parameters is now o The output of .
[0088] Spatiotemporal dimension. Formula assumptions. It is a one-dimensional feature representation. In this case, the outlier removal network... o and fusion operation It can be implemented as a shallow, fully connected network (e.g., two fully connected layers). In many multimodal models, features also have a spatiotemporal dimension of alignment between different modalities, i.e., ,in It is the number of feature channels, and This involves a shared spatiotemporal dimension (e.g., audio and visual features extracted from a video are aligned along the time axis, while features extracted from different visual modalities are aligned along the spatial axis). In those cases, outlier removal networks and fusion operations are implemented more efficiently with... The filter is a convolutional neural network. This enables the loss in equations (5) and (7) to be computed in parallel in the spatiotemporal dimensions.
[0089] Robust training program
[0090] Equipped with an anomaly removal network o and fusion sub-network multimodal A mechanism is included that compares information from all input sources, detects inconsistencies between perturbed modes and other unperturbed modes, and only allows information from unperturbed modes to pass through. During training, perturbed inputs are generated using a single-source adversarial perturbation from Equation 1. That is, let
[0091]
[0092] Note that this counter-perturbation is aimed at In other words, this method performs adversarial training of the fusion network and also utilizes adversarial examples to provide self-supervised labels for learning outlier removal. As shown in Algorithm 1, the loss optimization of the outlier removal network in equations (5), (8), and (9) is described. o and fusion sub-network The parameters. Note that it is not necessary to retrain the feature extractor that has already been pre-trained on clean data. .
[0093] Figure 9 This is a flowchart of a robust training strategy 900 for a feature fusion and outlier removal network. This flowchart corresponds to Algorithm 1 described above. In step 902, the controller initializes the outlier removal loss (odd-one-out loss) as shown in line 2 of Algorithm 1. In step 904, the controller initializes the task loss as shown in line 3 of Algorithm 1. In step 906, the controller receives samples from the training dataset as shown in line 4 of Algorithm 1 and proceeds to step 908, where the controller processes the samples using the function g as shown in line 5 of Algorithm 1. In step 910, the controller updates the outlier removal loss using unperturbed samples as shown in line 6 of Algorithm 1, and in step 912, the controller updates the task loss using unperturbed samples as shown in line 7 of Algorithm 1. In step 914, the controller generates perturbations for each modality. In step 916, the controller updates the outlier removal loss using samples with adversarial perturbations as shown in line 11 of Algorithm 1. In step 918, the controller updates the task loss using samples with adversarial perturbation as shown in line 12 of Algorithm 1. And in step 920, in response to satisfying the stopping criterion, the controller branches back to step 914 to provide an iteration for another perturbation. And in response to satisfying the stopping criterion, the controller branches to step 924. In step 924, the controller calculates the total loss, including the outlier removal loss and the task loss as shown in line 13 of Algorithm 1. In step 926, the controller updates the fusion function and the outlier removal network as shown in line 14 of Algorithm 1. The stopping criterion in step 920 may include a predetermined number of iterations, a predetermined runtime, convergence to a threshold, or a combination thereof.
[0094] Exemplary experimental data.
[0095] Exemplary evaluations of the single-source adversarial robustness of multimodal models were performed on three benchmark tasks: action recognition on EPIC-Kitchen, 2D object detection on KITTI, and sentiment analysis on MOSI. The benchmarks considered involve three input modalities and span a wider range of tasks and data sources, ensuring the generality of the conclusions drawn. A summary can be found in Table 1.
[0096]
[0097] Exemplary multimodal benchmark task.
[0098] An exemplary action recognition example from EPIC-Kitchens. EPIC-Kitchens is a large egocentric video dataset consisting of 39,596 video clips. The goal is to predict actions occurring in the videos, which consist of a verb and a noun from classes 126 and 331, respectively. Three modalities are obtained from the raw dataset: visual information (RGB frames), motion information (optical flow), and audio information. Figure 10A This is a graphical representation of an example action recognition result.
[0099] Exemplary object detection on KITTI. KITTI is an autonomous driving dataset that contains stereo camera and LiDAR information for 2D object detection, where the goal is to draw bounding boxes around objects of interest from predefined categories, such as cars, pedestrians, cyclists, etc. Existing work uses different combinations and processing variants of available data modalities for object detection. For the proposed benchmark, consider the following three modalities: (1) RGB frames used by most detection methods, (2) LiDAR points projected onto a sparse depth map, and (3) a depth map estimated from a stereo view. Figure 10B This is a graphical representation of an exemplary two-dimensional object detection result.
[0100] An Exemplary Sentiment Analysis Using CMU-MOSI. The Multimodal Opinion-Level Sentiment Intensity Corpus (CMU-MOSI) is a multimodal dataset for sentiment analysis, consisting of 93 video clips of movie reviews, each clip divided into an average of 23.2 segments. Each segment is labeled with a continuous sentiment intensity between [−3, 3]. The aim is to predict sentiment at either a binary scale (i.e., negative versus positive) or a 7-class scale (i.e., rounded to the nearest integer). MOSI contains three modalities: text, video, and audio. Figure 10C This is a graphical representation of an example sentiment analysis result.
[0101] Exemplary implementation details.
[0102] Exemplary model architecture and training. For each task, consider a mid- to late-stage multimodal model using the architectures summarized in column 4 of Table 1. Perform a first training baseline multimodal model for each task on clean data to obtain... Then, based on the adversarial robust fusion strategy, these models are extended with an outlier removal network and a robust feature fusion layer to obtain... And robust training is performed according to Algorithm 1.
[0103] Exemplary adversarial attacks. Column 5 of Table 1 summarizes the adversarial perturbations for each task, using Projected Gradient Descent (PGD) attacks on various modalities except for text, where word substitution was used. Note that these perturbations are white-box adaptive attacks, i.e., attacks that are fully understood by the target audience. It also generates attacks. Other types of attacks are also executed, such as redirection attacks, target attacks, and signature-based attacks.
[0104] Exemplary evaluation metrics. Column 6 of Table 1 summarizes the metrics used for each task. For action recognition, classification accuracy for verbs, nouns, and actions is considered. For object detection, average accuracy for car, pedestrian, and cyclist detection is considered at the intersection-over-union (IoU) threshold shown in the table and at three difficulty levels after the KITTI evaluation server. For sentiment analysis, binary and 7-level prediction accuracy are considered. For each metric, clean performance and performance under single-source attacks are considered.
[0105]
[0106]
[0107] Baseline
[0108] In addition to the methods presented in this disclosure, two types of methods were evaluated: a standard multimodal model trained with clean data (standard training) and a prior art robust multimodal model with robust training, evaluated using the following fusion method.
[0109] Concatenation Fusion with Standard Training (“Concat Fusion”). This is a standard approach for feature fusion using a multimodal model with the same feature extractor and concatenated features up to the final layer.
[0110] Mean fusion ("Mean Fusion") is used with standard training. For each modality, a single-modality model is trained using the same feature extractor and final layer as the multimodal model on clean data. Then, the mean fusion is achieved by taking their average. The outputs of the single-modal models are fused. For action recognition and sentiment analysis, average fusion is performed on the logits layer. For object detection, fusion is performed before the YOLO layer. Average fusion is a common fusion practice used in recent fusion models, and in the context of defense against adversarial perturbations, it is equivalent to a soft voting strategy between different modalities.
[0111] A robustly trained latent ensemble layer (“LEL + Robust”) is provided. The method comprises (1) training on clean data and data in which each single-source perturbation is performed alternately, and (2) using a concatenation fusion followed by a linear network to create a multimodal feature set. This strategy is adapted to the model presented in this disclosure by training these multimodal models on data augmented with single-source perturbations using the LEL + Robust fusion layer.
[0112] Information-Gated Fusion (“Gating + Robust”) with Robust Training. This method applies a multiplicative gating function to features from different modalities before combining feature sets from different modalities. An adaptive gating function is trained on clean data and data with single-source perturbations. This robust strategy is adapted to the model presented in this disclosure by training these multimodal models with gated feature fusion layers of these multimodal models on data enhanced with single-source adversarial perturbations.
[0113] The upper bound (“Oracle (Upper Bound)”) is used to obtain an empirical upper bound on robust performance against attacks targeting each modality. A 2-modal model excluding the perturbation modality is trained and evaluated. This model is called “Oracle” because it assumes complete knowledge of which modality is being attacked (i.e., a fully anomaly-removing network), which is not feasible in practice.
[0114]
[0115] Table 4. Binary and seven-level classification results on MOSI (%).
[0116]
[0117] Table 5. Detection rate (%) of the outlier removal network using a comparison of unaligned and aligned representations of features from each modality.
[0118]
[0119] Table 6. Number of parameters (in millions) in the feature extractor and the fusion network of the multimodal model of the present invention.
[0120] Figures 11-16Exemplary embodiments are shown; however, the concepts of this disclosure can be applied to other embodiments. Some exemplary embodiments include: industrial applications where the modality may include video, weight, IR, 3D camera, and sound; power tool or appliance applications where the modality may include torque, pressure, temperature, distance, or sound; medical applications where the modality may include ultrasound, video, CAT scan, MRI, or sound; robotic applications where the modality may include video, ultrasound, LIDAR, IR, or sound; and security applications where the modality may include video, sound, IR, or LIDAR. Modalities may have different datasets; for example, a video dataset may include images, a LIDAR dataset may include point clouds, and a microphone dataset may include time series data.
[0121] Figure 11 This is a schematic diagram of a control system 1102 configured to control a vehicle, which may be at least partially autonomous or at least partially autonomous. The vehicle includes sensors 1104 and actuators 1106. Sensors 1104 may include one or more wave energy-based sensors (e.g., charge-coupled device CCD or video), radar, LiDAR, microphone arrays, ultrasonic, infrared, thermal imaging, acoustic imaging, or other technologies (e.g., positioning sensors such as GPS). One or more of the specific sensors may be integrated into the vehicle. Alternatively, or in addition to the one or more specific sensors identified above, control module 1102 may include software modules configured to determine the state of actuator 1104 during execution.
[0122] In embodiments where the vehicle is at least partially autonomous, actuator 1106 may be implemented in the vehicle's braking system, propulsion system, engine, drivetrain, or steering system. Actuator control commands can be determined to control actuator 1106 so that the vehicle avoids collisions with detected objects. Detected objects can also be classified based on what a classifier deems most likely, such as pedestrians or trees. Actuator control commands can be determined based on these classifications. For example, control system 1102 may segment images (e.g., optical, acoustic, thermal) or other inputs from sensor 1104 into one or more background classes and one or more object classes (e.g., pedestrians, bicycles, vehicles, trees, traffic signs, traffic lights, road debris, or building barrels / cones, etc.) and send control commands to actuator 1106 to avoid collisions with objects, in which case the actuator is implemented in the braking or propulsion system. In another example, the control system 1102 can segment the image into one or more background classes and one or more marker classes (e.g., lane markings, guardrails, road edges, vehicle tracks, etc.) and send control commands to the actuator 1106 implemented herein in the steering system to cause the vehicle to avoid crossing the markers and remain within the lane. In scenarios where adversarial attacks may occur, the system described above can be further trained to better detect objects or recognize changes in lighting conditions or angles of sensors or cameras on the vehicle.
[0123] In other embodiments where vehicle 1100 is at least partially autonomous, vehicle 1100 may be a mobile robot configured to perform one or more functions such as flying, swimming, diving, and stepping. The mobile robot may be at least partially autonomous lawnmower or at least partially autonomous cleaning robot. In such embodiments, actuator control command 1106 may be determined such that the mobile robot's propulsion unit, steering unit, and / or braking unit can be controlled to prevent the mobile robot from colliding with identified objects.
[0124] In another embodiment, vehicle 1100 is at least partially autonomous in the form of a gardening robot. In such an embodiment, vehicle 1100 may use an optical sensor as sensor 1104 to determine the state of plants in the environment approaching vehicle 1100. Actuator 1106 may be a nozzle configured to spray chemicals. Based on the identified species and / or the identified state of the plants, actuator control command 1102 may be determined to cause actuator 1106 to spray the plants with an appropriate amount of appropriate chemicals.
[0125] Vehicle 1100 may be a partially autonomous robot in the form of a household appliance. Non-limiting examples of household appliances include washing machines, stoves, ovens, microwave ovens, or dishwashers. In such vehicle 1100, sensor 1104 may be an optical or acoustic sensor configured to detect the state of an object to be processed by the household appliance. For example, in the case of a washing machine, sensor 1104 may detect the state of the clothes inside the washing machine. Actuator control commands may be determined based on the detected state of the clothes.
[0126] In this embodiment, the control system 1102 receives image (optical or acoustic) and annotation information from the sensor 1104. It uses these and stores a predetermined number of category k and similarity metrics in the system. The control system 1102 can classify each pixel of the image received from the sensor 1104 using the method described in FIG. 10. Based on the classification, a signal can be sent to the actuator 1106, for example, to brake or steer to avoid a collision with a pedestrian or tree, to steer to stay between detected lane markings, or any action performed by the actuator 1106 as described above. Based on the classification, the signal can also be sent to the sensor 1104, for example, to focus or move the camera lens.
[0127] Figure 12 A schematic diagram is depicted of a system 1200 (e.g., a manufacturing machine) configured to control a manufacturing system 102, such as a stamping and cutting machine, a cutting machine, or a deep hole drill, which is part of a production line. The control system 1202 may be configured to control actuator 14, which is configured to control the control system 100 (e.g., the manufacturing machine).
[0128] The sensor 1204 of system 1200 (e.g., a manufacturing machine) may be a wave energy sensor, such as an optical or acoustic sensor or sensor array configured to capture one or more attributes of the manufactured product. Control system 1202 may be configured to determine the state of the manufactured product based on one or more of the captured attributes. Actuator 1206 may be configured to control system 1202 (e.g., the manufacturing machine) based on the state of the manufactured product 104 determined for subsequent manufacturing steps of the manufactured product. Actuator 1206 may be configured to control based on the determined state of previously manufactured products. Figure 11 The function of a system (e.g., a manufacturing machine) in the subsequent manufacture of products.
[0129] In this embodiment, the control system 1202 receives images (e.g., optical or acoustic) and annotation information from the sensor 1204. It uses these and stores a predetermined number of category k and similarity metrics in the system. The control system 1202 can use the method described in Figure 10 to classify each pixel of the image received from the sensor 1204, for example, segmenting an image of a manufactured object into two or more categories, detecting anomalies in the manufactured product to ensure the presence of objects such as barcodes on the manufactured product. Based on this classification, a signal can be sent to the actuator 1206. For example, if the control system 1202 detects an anomaly in the product, the actuator 1206 can mark or remove the abnormal or defective product from the production line. In another example, if the control system 1202 detects the presence of barcodes or other objects placed on the product, then the actuator 1106 can apply or remove these objects. Based on this classification, a signal can also be sent to the sensor 1204, for example, to focus or move a camera lens.
[0130] Figure 13 A schematic diagram of a control system 1302 configured to control a power tool 1300, such as an electric drill or actuator, having at least a partially autonomous mode is shown. The control system 1302 may be configured to control an actuator 1306 configured to control the power tool 1300.
[0131] The sensor 1304 of the power tool 1300 may be a wave energy sensor, such as an optical or acoustic sensor, configured to capture one or more properties of the work surface and / or the fastener driven into the work surface. The control system 1302 may be configured to determine the state of the work surface and / or the fastener relative to the work surface based on one or more of the captured properties.
[0132] In this embodiment, the control system 1302 receives images (e.g., optical or acoustic) and annotation information from the sensor 1304. These are used in conjunction with a predetermined number of category k and similarity metrics stored in the system. The control system 1302 can use the method described in Figure 10 to classify each pixel of the image received from the sensor 1304 to segment the image of the work surface or fastener into two or more categories, or to detect anomalies in the work surface or fastener. Based on this classification, a signal can be sent to the actuator 1306, such as pressure or speed to the tool, or any action performed by the actuator 1306 as described above. Based on this classification, a signal can also be sent to the sensor 1304, for example, to focus or move a camera lens. In another example, the image can be a time-series image of signals from the power tool 1300, such as pressure, torque, revolutions per minute, temperature, current, etc., where the power tool is a hammer drill, drill, hammer (rotary or demolition), impact drive, reciprocating saw, oscillating multi-tool, and the power tool is wireless or wired.
[0133] Figure 14 A schematic diagram depicts a control system 1402 configured to control an automated personal assistant 1401. The control system 1402 may be configured to control an actuator 1406, which is also configured to control the automated personal assistant 1401. The automated personal assistant 1401 may be configured to control household appliances, such as a washing machine, stove, oven, microwave oven, or dishwasher.
[0134] In this embodiment, the control system 1402 receives images (e.g., optical or acoustic) and annotation information from the sensor 1404. These are used in conjunction with a predetermined number of category k and similarity metrics stored in the system. The control system 1402 can use the method described in FIG10 to classify each pixel of the image received from the sensor 1404, for example, segmenting the image of an appliance or other object to be manipulated or operated. Based on this classification, a signal can be sent to the actuator 1406, for example, to control the moving part of the automated personal assistant 1401 to interact with household appliances, or any action performed by the actuator 1406 as described above. Based on this classification, the signal can also be sent to the sensor 1404, for example, to focus or move a camera lens.
[0135] Figure 15 A schematic diagram of a control system 1502 configured as a monitoring system 1500 is depicted. The monitoring system 1500 can be configured to physically control entry through door 252. Sensor 1504 can be configured to detect and determine the relevant scene for access permission. Sensor 1504 can be an optical or acoustic sensor or sensor array configured to generate and transmit image and / or video data. Such data can be used by the control system 1502 to detect a person's face.
[0136] The monitoring system 1500 can also be a surveillance system. In such an embodiment, the sensor 1504 can be a wave energy sensor, such as an optical sensor, an infrared sensor, an acoustic sensor configured to detect the scene under surveillance, and the control system 1502 is configured to control the display 1508. The control system 1502 is configured to determine the classification of the scene, for example, whether the scene detected by the sensor 1504 is suspicious. Disturbance objects can be used to detect certain types of objects to allow the system to identify such objects under suboptimal conditions (e.g., nighttime, fog, rain, background noise, etc.). The control system 1502 is configured to transmit actuator control commands to the display 1508 in response to the classification. The display 1508 can be configured to adjust the displayed content in response to the actuator control commands. For example, the display 1508 can highlight objects that the controller 1502 considers suspicious.
[0137] In this embodiment, the control system 1502 receives image (optical or acoustic) and annotation information from the sensor 1504. A predetermined number of these and stored categories are used and stored in the system. k and similarity measurement The control system 1502 can use the method described in FIG10 to classify each pixel of the image received from the sensor 1504 in order to, for example, detect the presence of suspicious or unwanted objects in the scene, detect the type of lighting or viewing conditions, or detect motion. Based on this classification, signals can be sent to the actuator 1506, for example, to lock or unlock a door or other entrance passage, to activate an alarm or other signal, or any of the actions performed by the actuator 1506 as described above. Based on this classification, signals can also be sent to the sensor 1504, for example, to focus or move a camera lens.
[0138] Figure 16 A schematic diagram is depicted of a control system 1602 configured to control an imaging system 1600, such as an MRI apparatus, an X-ray imaging apparatus, or an ultrasound apparatus. The sensor 1604 may be, for example, an imaging sensor or an array of acoustic sensors. The control system 1602 may be configured to determine a classification of all or part of the sensed image. The control system 1602 may be configured to determine or select actuator control commands in response to a classification obtained by a trained neural network. For example, the control system 1602 may interpret a region of the sensed image (optical or acoustic) as a potential anomaly. In this case, an actuator control command may be determined or selected to cause the display 1606 to display the image and highlight the potentially anomalous region.
[0139] In this embodiment, the control system 1602 receives image and annotation information from the sensor 1604. It then uses and stores a predetermined number of categories of these information within the system. k and similarity measurement The control system 1602 can use the method described in FIG10 to classify each pixel of the image received from the sensor 1604. Based on the classification, a signal can be sent to the actuator 1606, for example, to detect abnormal regions of the image or any action performed by the actuator 1606 as described in the above sections.
[0140] Program code implementing the algorithms and / or methods described herein can be distributed individually or collectively as a program product in various different forms. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to perform aspects of one or more embodiments. Inherently non-transitory computer-readable storage media can include volatile and non-volatile, and removable and non-removable tangible media implemented with any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media may also include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, portable optical disc read-only memory (CD-ROM) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is computer-readable. Computer-readable program instructions can be downloaded from the computer-readable storage medium to a computer, another type of programmable data processing apparatus, or another device, or downloaded via a network to an external computer or external storage device.
[0141] Computer-readable program instructions stored in a computer-readable medium can be used to instruct a computer, other type of programmable data processing apparatus, or other device to operate in a particular manner, causing the instructions stored in the computer-readable medium to produce an article of art including instructions that implement the functions, actions, and / or operations specified in a flowchart or diagram. In some alternative embodiments, the functions, actions, and / or operations specified in the flowcharts and diagrams can be reordered, processed sequentially, and / or processed concurrently, consistent with one or more embodiments. Furthermore, any of the flowcharts and / or diagrams may include more or fewer nodes or blocks than those shown according to one or more embodiments.
[0142] While the invention has been fully described through various embodiments, and while these embodiments have been described in considerable detail, the applicant does not intend to limit the scope of the appended claims or restrict them in any way to such details. Additional advantages and modifications will readily occur to those skilled in the art. Therefore, the invention, in its broader aspects, is not limited to the specific details, representative apparatuses and methods, and illustrative examples shown and described. Consequently, deviations from these details may be made without departing from the spirit or scope of the overall conception of the invention.
Claims
1. A multimodal sensing system, comprising: The controller is configured to, It receives a first signal from a first sensor, a second signal from a second sensor, and a third signal from a third sensor. Extract the first feature vector from the first signal. Extract the second feature vector from the second signal. Extract the third feature vector from the third signal. Based on inconsistent mode prediction, an outlier removal vector is determined from the first feature vector, the second feature vector, and the third feature vector via a machine learning network. The outlier removal network is a neural network o that maps features z to a vector of size k+1. The i-th entry of this vector indicates the probability that mode i has been perturbed, and the (k+1)-th entry indicates the probability that no mode has been perturbed. , The first feature vector, the second feature vector, and the third feature vector are fused with the outlier removal vector to form a fused feature vector. Output the fused feature vector.
2. The multimodal sensing system as described in claim 1, wherein, The first sensor, the second sensor, and the third sensor each have different modes.
3. The multimodal sensing system as described in claim 2, wherein, The length of the outlier removal vector is the mode plus 1, and each mode has a perturbation, where the plus 1 represents an unperturbed mode.
4. The multimodal sensing system as described in claim 3, wherein, The controller determines the outlier removal network and fuses the feature vector with the outlier removal vector into a convolutional neural network (CNN) to align the spatiotemporal dimensions of the different modalities.
5. The multimodal sensing system as described in claim 1, wherein, The controller is further configured to extract the first feature vector from the first signal via a first pre-trained AI model, extract the second feature vector from the second signal via a second pre-trained AI model, and extract the third feature vector from the third signal via a third pre-trained AI model.
6. The multimodal sensing system as described in claim 5, wherein, The controller is also configured to jointly train the outlier removal network in parallel with the task modality based on a loss function, the loss function being expressed as follows: Where z is the feature set extracted from the k-modal input, and g i It is a feature extractor applied to the i-th mode, z -i It is the feature received from the i-th mode. From disturbance input Features extracted from, and o It is a network that removes outliers and, in response to a stopping criterion event, stops the joint training.
7. The multimodal sensing system as described in claim 1, wherein, The controller is further configured to fuse the first feature vector, the second feature vector, the third feature vector, and the outlier vector into a fused feature vector according to the following formula: in, This represents the concatenation operation, NN represents a shallow neural network, z is the input, k is the modality, and z -i It is the feature received from the i-th mode. e i It comes from excluding i The fusion of features from all modalities except those mentioned above, and only e k+1 Features from all modalities are fused.
8. The multimodal sensing system as described in claim 1, wherein, The first sensor is one of video, RADAR, LIDAR, or ultrasound, and the controller is further configured to control the autonomous vehicle based on the fused feature vector.
9. The multimodal sensing system as described in claim 1, wherein, The first sensor is one of video, sound, IR, or LIDAR, and the controller is further configured to control the channel gate based on the fused feature vector.
10. The multimodal sensing system as described in claim 1, wherein, The first sensor is one of video, sound, ultrasound, IR, or LIDAR, and the controller is further configured to control a mechanical system.
11. A multimodal sensing method, comprising: Receive a first signal from a first sensor, a second signal from a second sensor, and a third signal from a third sensor; Extract a first feature vector from the first signal, extract a second feature vector from the second signal, and extract a third feature vector from the third signal; Based on inconsistent mode prediction, an outlier removal vector is determined from the first feature vector, the second feature vector, and the third feature vector via a machine learning network. The outlier removal network is a neural network o that maps features z to a vector of size k+1. The i-th entry of this vector indicates the probability that mode i has been perturbed, and the (k+1)-th entry indicates the probability that no mode has been perturbed. , The first feature vector, the second feature vector, and the third feature vector are fused with the outlier vector to form a fused feature vector; and Output the fused feature vector.
12. The multimodal sensing method as described in claim 11, wherein, The first sensor, the second sensor, and the third sensor each have different modes.
13. The multimodal sensing method as described in claim 12, wherein, The length of the outlier removal vector is the mode plus 1, and each mode has a perturbation, where the plus 1 represents an unperturbed mode.
14. The multimodal sensing method as described in claim 13, wherein, The determination of the outlier removal network and the fusion of the feature vector with the outlier removal vector are achieved by aligning the spatiotemporal dimensions of the different modalities via a convolutional neural network (CNN).
15. The multimodal sensing method as described in claim 11, wherein, The first feature vector is extracted from the first signal by a first pre-trained AI model, the second feature vector is extracted from the second signal by a second pre-trained AI model, and the third feature vector is extracted from the third signal by a third pre-trained AI model.
16. The multimodal perception method of claim 15, further comprising jointly training the outlier removal network in parallel with the task modality according to a loss function expressed by the following formula: Where z is the feature set extracted from the k-modal input, and g i It is a feature extractor applied to the i-th mode, z -i It is the feature received from the i-th mode. From disturbance input Features extracted from, and o The network that removes outliers stops the joint training in response to a stopping criterion event.
17. The multimodal sensing method as described in claim 11, wherein, The first feature vector, the second feature vector, the third feature vector, and the outlier vector are fused into a fused feature vector according to the following formula: in, This represents the concatenation operation, NN represents a shallow neural network, z is the input, k is the modality, and z -i It is the feature received from the i-th mode. e i It comes from excluding i The fusion of features from all modalities except those mentioned above, and only e k+1 Features from all modalities are fused.
18. A multimodal perception system for autonomous vehicles, comprising: The first sensor is one of a video, RADAR, LIDAR, or ultrasonic sensor; as well as The controller is configured to, It receives a first signal from a first sensor, a second signal from a second sensor, and a third signal from a third sensor. Extract the first feature vector from the first signal. Extract the second feature vector from the second signal. Extract the third feature vector from the third signal. Based on inconsistent mode prediction, an outlier removal vector is determined from the first feature vector, the second feature vector, and the third feature vector via a machine learning network. The outlier removal network is a neural network o that maps features z to a vector of size k+1. The i-th entry of this vector indicates the probability that mode i has been perturbed, and the (k+1)-th entry indicates the probability that no mode has been perturbed. , The first feature vector, the second feature vector, the third feature vector, and the outlier vector are fused into a fused feature vector. Output the fused feature vector, and The autonomous vehicle is controlled based on the fused feature vector.
19. The multimodal sensing system as described in claim 18, wherein, The first, second, and third sensors are each different modes, and the length of the outlier removal vector is the mode plus 1, and each mode has a perturbation, where the plus 1 represents the unperturbed mode.
20. The multimodal sensing system as described in claim 19, wherein, The controller determines the outlier removal network and fuses the feature vector with the outlier removal vector into a convolutional neural network (CNN) to align the spatiotemporal dimensions of the different modalities.