Guided multispectral examination
By using a pre-trained neural network in the visible domain to detect regions of interest and guide millimeter-wave radar scanning, the trade-off between performance and region size in millimeter-wave imaging is resolved, achieving efficient and accurate multispectral imaging.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTERNATIONAL BUSINESS MACHINE CORPORATION
- Filing Date
- 2021-07-13
- Publication Date
- 2026-04-24
AI Technical Summary
In millimeter-wave imaging, there is a trade-off between performance (speed, accuracy, signal-to-noise ratio, system complexity) and the size of the region of interest, resulting in the ability to only partially optimize the imaging effect.
A pre-trained neural network is used to detect the region of interest in the visible domain, and a controller subsystem guides a second imaging system (such as millimeter-wave radar) to a specific location for scanning. Machine learning techniques are combined to adjust the focus in real time to collect data in different spectral domains.
It achieves efficient and accurate imaging within the region of interest, reduces energy consumption from unnecessary scanning, and improves the overall performance of the imaging system.
Smart Images

Figure CN116368531B_ABST
Abstract
Description
Technical Field
[0001] This invention relates generally to imaging systems, and more specifically to guided multispectral inspection. Background Technology
[0002] Thanks to the advent of small, portable imaging sensors (e.g., cameras, IR cameras, and imaging radar), multispectral imaging of a given scene is now possible. Each part of the spectrum provides different information that may be relevant to a given application. For example, millimeter-wave (mmWave) imaging radar has the ability to obtain information (approximate shape and reflectivity) from objects located behind / inside opaque materials (such as cardboard packaging or fabric).
[0003] Meanwhile, computer vision (CV) algorithms and machine learning (ML) methods have matured significantly and can now automatically extract information from visible field cameras and / or videos. For example, the location of a certain type of object in a scene can be automatically identified.
[0004] Millimeter-wave imagers operate by: forming a light beam, guiding the beam to illuminate a given portion of a scene, receiving reflected signals (using a receiver to point the beam in the same direction), and obtaining the distance to the object and reflectivity in that beam direction using signal processing techniques. An image is formed by repeating this process at multiple locations.
[0005] There is a common trade-off between millimeter-wave imaging performance (speed, accuracy, signal-to-noise ratio, spatial resolution) and the size of the region of interest. Although millimeter-wave images can have a wider total field of view (FoV) than cameras, for optimal results (higher frame rates, higher spatial resolution, and less energy consumption), it is preferable to scan only the most relevant part of the scene in a single scan. Summary of the Invention
[0006] According to an aspect of the present invention, an imaging system is provided. The imaging system includes a first imaging system for acquiring initial sensor data in the form of visible domain data. The imaging system further includes a second imaging system for acquiring subsequent sensor data in the form of second domain data, wherein the initial sensor data and the subsequent sensor data have different spectral domains. The imaging system also includes a controller subsystem that, by applying machine learning techniques to the visible domain data, detects at least one region of interest in real time, locates at least one object of interest within the at least one region of interest to generate location data for at least one object of interest, and, in response to the location data, autonomously directs the focus of the second imaging system to a region of the scene including the object of interest to acquire second domain data.
[0007] According to other aspects of the invention, a method for imaging is provided. The method includes acquiring initial sensor data in the form of visible domain data via a first imaging system. The method further includes detecting at least one region of interest in real time by a controller subsystem applying machine learning techniques to the visible domain data. The method also includes locating at least one object of interest within the at least one region of interest via the controller subsystem to generate location data for at least one object of interest. The method additionally includes the controller subsystem autonomously directing the focus of a second imaging system to the same or similar scene in response to the acquisition of subsequent sensor data in the form of second domain data, wherein the initial and subsequent sensor data have different spectral domains.
[0008] According to another aspect of the invention, a computer program product for imaging is provided. The computer program product includes a non-transitory computer-readable storage medium having program instructions embodied therein. The program instructions are executable by a computing system to cause the computing system to perform a method. The method includes acquiring initial sensor data in the form of visible domain data by a first imaging system of the computing system. The method further includes detecting at least one region of interest in real time by a controller subsystem of the computing system applying machine learning techniques to the visible domain data. The method also includes locating at least one object of interest in the at least one region of interest by the controller subsystem to generate location data of at least one object of interest. The method further includes autonomously directing the focus of a second imaging system of the computing system to the same or similar scene by the controller subsystem in response to the location data to acquire subsequent sensor data in the form of second domain data. The initial and subsequent sensor data have different spectral domains.
[0009] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which will be read in conjunction with the accompanying drawings. Attached Figure Description
[0010] The following description will provide details of preferred embodiments with reference to the following figures, in which:
[0011] Figure 1 This is a block diagram illustrating an exemplary processing system according to an embodiment of the present invention;
[0012] Figure 2 This is a block diagram illustrating an exemplary artificial neural network (ANN) architecture according to an embodiment of the present invention;
[0013] Figure 3 This is a block diagram illustrating an exemplary neuron according to an embodiment of the present invention;
[0014] Figure 4 This is a block diagram illustrating an exemplary system for guided multispectral inspection according to an embodiment of the present invention;
[0015] Figures 5 to 6 This is a flowchart illustrating an exemplary method for guided multispectral inspection according to an embodiment of the present invention;
[0016] Figure 7 This illustrates an embodiment of the invention. Figure 4 A block diagram of an exemplary neural network configuration used by the system;
[0017] Figure 8 This is a block diagram illustrating an illustrative cloud computing environment according to an embodiment of the present invention, having one or more cloud computing nodes communicating with a local computing device used by a cloud consumer; and
[0018] Figure 9 This is a block diagram illustrating a set of functional abstraction layers provided by a cloud computing environment according to an embodiment of the present invention. Detailed Implementation
[0019] The embodiments of the present invention are for guided multispectral inspection.
[0020] One or more embodiments of the present invention enable information extracted in one spectral domain (e.g., the visible domain) to guide the operation of a sensor in a different spectral domain (e.g., imaging radar). Specifically, one or more embodiments of the present invention describe visible domain information having a specific location in a scene that the imaging radar or a second imaging device should focus on.
[0021] One or more embodiments of the present invention may involve using artificial intelligence (AI) driven attention to identify objects of interest.
[0022] As an example of the problem this invention aims to solve, it is important to note the overall trade-off between millimeter-wave imaging performance (speed, accuracy, signal-to-noise ratio, system complexity) and the size of the region of interest. While a millimeter-wave imager (or an exemplary second imaging system) can have a wider total field of view (FoV) than a camera (or an exemplary first imaging system), for optimal results, it is preferable to scan only the most relevant portions of the scene at a given time. One or more embodiments of the invention address this trade-off by processing images acquired using a first imaging system (which can then be rapidly scanned by a millimeter-wave radar) to autonomously detect the region of interest.
[0023] To this end, the present invention uses a pre-trained neural network (which has learned to focus attention on regions of interest in images acquired by the first imaging system) to guide the second imaging system to the same or similar scene.
[0024] In one embodiment, a first imaging system is configured to acquire initial sensor data in the form of visible-domain data, and a second imaging system is configured to acquire subsequent sensor data in the form of second-spectral-domain data. In another embodiment, the subsequent imaging data has different sensing characteristics than the initial imaging data. For example, the second imaging data may come from a millimeter-wave radar with 3D sensing capabilities and the ability to detect objects partially or completely covered in the visible domain by opaque materials (such as fabrics, cardboard boxes, and plastics).
[0025] Figure 1 This is a block diagram illustrating an exemplary processing system 100 according to an embodiment of the present invention. The processing system 100 can be used as... Figure 4 The sensor control and data processing subsystem 430 is described. The processing system 100 includes a set of processing units (e.g., CPUs) 101, a set of GPUs 102, a set of memory devices 103, a set of communication devices 104, and a set of peripheral devices 105. The CPU 101 may be a single-core or multi-core CPU. The GPU 102 may be a single-core or multi-core GPU. One or more memory devices 103 may include cache, RAM, ROM, and other memories (flash memory, optical memory, magnetic memory, etc.). The communication devices 104 may include wireless and / or wired communication devices (e.g., network (e.g., WIFI, etc.) adapters, etc.). The peripheral devices 105 may include display devices, user input devices, printers, imaging devices, etc. The components of the processing system 100 are connected by one or more buses or networks (collectively indicated by reference numeral 110).
[0026] In one embodiment, the processing system 100 further includes a controller 177 for guiding a multispectral inspection. The controller 177 may be implemented using an ASIC, FPGA, or the like. In one embodiment, the controller 177 has on-board memory for storing program code for performing the guided multispectral inspection. In other embodiments, the program code may be stored in a memory device 103. In other embodiments, the controller 177 is embodied by one or more CPUs 101 and / or one or more GPUs 102. These and other variations can be readily implemented, as will be readily understood by those skilled in the art, given the teachings of the invention provided herein.
[0027] In one embodiment, memory device 103 may store specially programmed software modules to transform a computer processing system into a dedicated computer configured to implement various aspects of the present invention. In another embodiment, dedicated hardware (e.g., application-specific integrated circuits, field-programmable gate arrays (FPGAs), etc.) may be used to implement various aspects of the present invention.
[0028] Of course, the processing system 100 may also include other elements (not shown), as readily apparent to those skilled in the art, and some elements may be omitted. For example, as readily understood by those skilled in the art, input devices and / or output devices may be included in the processing system 100 depending on the specific implementation of various other input devices and / or output devices in the processing system 100. For example, various types of wireless and / or wired input and / or output devices may be used. Furthermore, additional processors, controllers, memory, etc., may be utilized in various configurations. Further, in another embodiment, a cloud configuration may be used (e.g., see...). Figure 8-9 Given the teachings of the invention provided herein, those skilled in the art will readily conceive of these and other variations of the processing system 100.
[0029] Furthermore, it is understood that the various figures described below with respect to the various elements and steps related to the present invention may be implemented, in whole or in part, by one or more elements of system 100.
[0030] As used herein, the terms "hardware processor subsystem" or "hardware processor" can refer to a processor, memory, software, or a combination thereof that cooperate to perform one or more specific tasks. In useful embodiments, a hardware processor subsystem may include one or more data processing elements (e.g., logic circuitry, processing circuitry, instruction execution devices, etc.). These one or more data processing elements may be included in a central processing unit, a graphics processing unit, and / or a separate processor- or computing element-based controller (e.g., logic gates, etc.). A hardware processor subsystem may include one or more on-board memories (e.g., cache, dedicated memory array, read-only memory, etc.). In some embodiments, a hardware processor subsystem may include one or more memories (e.g., ROM, RAM, basic input / output system (BIOS), etc.) that may be on-board or off-board, or may be dedicated to use by the hardware processor subsystem.
[0031] In some embodiments, the hardware processor subsystem may include and execute one or more software elements. The one or more software elements may include an operating system and / or one or more applications and / or specific code for implementing a specified result.
[0032] In other embodiments, the hardware processor subsystem may include dedicated, special-purpose circuitry that performs one or more electronic processing functions to achieve a specified result. Such circuitry may include one or more application-specific integrated circuits (ASICs), FPGAs, and / or PLAs.
[0033] According to embodiments of the present invention, these and other variations of the hardware processor subsystem are also conceivable.
[0034] This hardware processor can be used to perform guided multispectral inspections using data from multiple sensors.
[0035] According to embodiments of the present invention, these and other variations of the hardware processor subsystem are also conceivable.
[0036] Figure 2 This is a block diagram illustrating an exemplary artificial neural network (ANN) architecture 200 according to an embodiment of the present invention. It is understood that this architecture is purely exemplary and other architectures or types of neural networks may be used alternatively. Specifically, while hardware embodiments of ANNs are described herein, it is understood that neural network architectures can be implemented or simulated in software. The hardware embodiments described herein are intended to illustrate the general principles of neural network computation in a high-level generalization and should not be construed as limiting in any way.
[0037] Furthermore, the neuron layers and the weights connecting them described below are described in a general manner and can be replaced by any type of neural network layer with any appropriate degree or type of interconnectivity. For example, layers may include convolutional layers, pooling layers, fully connected layers, softmax layers, or any other suitable type of neural network layer. Additionally, layers may be added or removed as needed, and weights may be omitted for more complex forms of interconnection.
[0038] During the feed-forward operation, a set of input neurons 202 each provides an input voltage in parallel with the weights 204 of the corresponding row. In the hardware embodiment described herein, each weight 204 has a configurable resistance value such that a current output flows from the weight 204 to the corresponding hidden neuron 206 to represent a weighted input. In the software embodiment, the weight 204 can be simply represented as a coefficient value multiplied by the output of the associated neuron.
[0039] Following the hardware implementation, the current output by the given weight 204 is determined as follows: Where V is the input voltage from input neuron 202, and r is the set resistance of weight 204. The current from each weight is summed column-wise and flows to hidden neuron 206. A set of reference weights 207 with fixed resistances combines their outputs into a reference current, which is provided to each of the hidden neurons 206. Because conductance values can only be positive, some reference conductance is needed to encode positive and negative values in the matrix. The current generated by weight 204 is continuously positive, and therefore reference weights 207 are used to provide the reference current, assuming that current above the reference current is positive and current below the reference current is negative. In a software embodiment, reference weights 207 are not required, where the values of the output and weights can be obtained precisely and directly. As an alternative to using reference weights 207, another embodiment can use a separate array of weights 204 to acquire negative values.
[0040] Hidden neuron 206 uses current from weight array 204 and reference weight 207 to perform some calculations. Hidden neuron 206 then outputs its own voltage to another weight array 204. This array performs in the same way, where a column of weights 204 receives voltage from their respective hidden neurons 206 to produce a weighted current output that is added row by row and provided to output neuron 208.
[0041] It is understood that any number of these stages can be achieved by inserting additional layers of the array and hidden neurons 206. It should also be noted that some neurons can be constant neurons 209 that provide a constant output to the array. Constant neurons 209 can be present between input neurons 202 and / or hidden neurons 206, and are used only during forward feed operations.
[0042] During backpropagation, output neurons 208 provide a reverse voltage across the weight array 204. The output layer compares the generated network response with the training data and calculates the error. The error is applied to the array as a voltage pulse, where the pulse height and / or duration is modulated proportionally to the error value. In this example, a row of weights 204 receives the voltage from the corresponding output neuron 208 in parallel and converts it into a column-wise current that provides input to the hidden neuron 206. The hidden neuron 206 combines the weighted feedback signal with the derivative calculated from its forward feed and stores the error value before outputting the feedback signal voltage to its corresponding weight column 204. This backpropagation travels through the entire network 200 until all hidden neurons 206 and input neurons 202 have stored the error value.
[0043] During weight updates, a first weight update voltage is applied forward to input neurons 202 and hidden neurons 206 through network 200, and a second weight update voltage is applied backward to output neurons 208 and hidden neurons 206. The combination of these voltages produces a state change within each weight 204, causing weight 204 to exhibit a new resistance value. In this way, weights 204 can be trained to adapt the neural network 200 to errors in its processing. It should be noted that the three operating modes (forward feeding, backpropagation, and weight update) do not overlap.
[0044] As described above, the weights 204 can be implemented in software or hardware, for example, using relatively complex weighting circuits or resistive crosspoint devices. This resistive device can have switching characteristics, which have nonlinearity that can be used to process data. The weights 204 can belong to a class of devices called resistive processing units (RPUs) because their nonlinear characteristics are used to perform computations in the neural network 200. The RPU device can be implemented using resistive random access memory (RRAM), phase-change memory (PCM), programmable metallized cell (PMC) memory, or any other device with nonlinear resistive switching characteristics. The RPU device can also be considered a memristor system.
[0045] Figure 3 This is a block diagram illustrating an exemplary neuron 300 according to an embodiment of the present invention. This neuron may represent any one of input neuron 202, hidden neuron 206, or output neuron 208. It should be noted that... Figure 3 The components that handle all three operational phases are shown: forward feeding, backpropagation, and weight update. However, since the different phases do not overlap, there must be some form of control mechanism within neuron 300 to control which components are active. Therefore, it is understandable that switches and other structures, not shown in neuron 300, may exist to handle mode switching.
[0046] In the forward feed mode, interpolation block 302 determines the value of the input by comparing the input from the array with a reference input. This sets the amplitude and sign (e.g., + or -) of the input from the array to the neuron 300. Block 304 performs a computation based on the input, and its output is stored in memory 305. In particular, it is contemplated that block 304 computes a nonlinear function and can be implemented as an analog or digital circuit or can be executed in software. The value determined by function block 304 is converted into a voltage at forward feed generator 306, which applies the voltage to the next array. The signal propagates through multiple layers of the array and the neuron until it reaches the final output layer of the neuron. The input is also applied to the derivative of the nonlinear function in block 308, and its output is stored in memory 309.
[0047] During backpropagation, an error signal is generated. This error signal can be generated at output neuron 208 or computed by a separate unit that receives input from output neuron 208 and compares the output with the correct output based on training data. Alternatively, if neuron 300 is hidden neuron 206, it receives backpropagation information from weight array 204 and compares the received information with a reference signal at difference box 310 to provide a signed error signal with continuously taking values. This error signal is multiplied by multiplier 312 with the derivative of a nonlinear function of the previous forward feed step stored in memory 309, and the result is stored in memory 313. The value determined by multiplier 312 is converted into a backpropagation voltage pulse proportional to the error computed at backpropagation generator 314, which applies the voltage to the previous array. The error signal thus propagates through multiple layers of the neuron and array until it reaches the input layer 202 of the neuron.
[0048] During the weight update mode, after the forward and reverse propagation are completed, each weight 204 is updated proportionally to the product of the signals passed through the weights during the forward and reverse propagation. The update signal generator 316 provides voltage pulses in both directions (but note that only one direction will be available for both input and output neurons). The shape and amplitude of the pulses from the update generator 316 are configured to change the state of the weights 204, causing the resistance of the weights 204 to be updated.
[0049] Compared to forward and reverse loops, implementing weight updates locally and entirely in parallel on a 2D cross-array of resistive processing units, independent of array size, is challenging. This requires computing vector-vector cross products and may require multiplication operations and incremental weight updates performed locally at each cross-array.
[0050] Figure 4 This is a block diagram illustrating an exemplary system 400 for guided multispectral inspection according to an embodiment of the present invention.
[0051] System 400 includes a first imaging system 410, a second imaging system 420, and a sensor control and data processing subsystem 430. In embodiments, one or both of the first imaging system 410 and the second imaging system 420 may include a delivery system for delivering the system to a desired scene with potentially interesting targets. For example, after the first imaging system finds a region of interest, a delivery system such as a drone may then deploy the second imaging system for additional scanning of the region of interest. However, other embodiments place the first and second imaging systems in a co-location and operate simultaneously to minimize latency. Within the fraction of the time of the second imaging system, inferences about the region of interest identified by the first imaging system can be obtained.
[0052] In one embodiment, the sensor control and data processing subsystem 430, the first imaging system 410, and the second imaging system 420 are capable of wireless communication therebetween. In other embodiments, other types of connections may be used.
[0053] Images from the first imaging system 410 and the second imaging system 420 share a common set of x and y coordinates from the same or similar scene (here shown as the same scene 450). In this embodiment, the origins in the two scenes are the same for quick reference relative to each other.
[0054] exist Figure 4 In the example, the scenario could be a train platform where there are unattended bags.
[0055] The object of interest (in the above example, the unattended bag) and its location can be identified by applying computer vision (CV) or machine learning (ML) algorithms to the data acquired by the first imaging system.
[0056] Shared coordinates are used to control the imaging area oriented (pointed to) by the second imaging system 420. It is assumed that the second imaging system 420 includes electron beam scanning capabilities or other means for obtaining imaging data from a specific FoV illumination.
[0057] In one embodiment, the first imaging system 410 includes a camera. In another embodiment, the camera is an RGB camera. In yet another embodiment, the camera is an infrared (IR) camera. Of course, other types of cameras and imaging devices may be included in the first imaging system 410.
[0058] In this embodiment, the second imaging system is a radar imager. Specifically, the radar imager is a millimeter-wave radar imager with beamforming and beamguiding capabilities. Of course, other types of radar imagers and imaging systems can also be used as the second imaging device.
[0059] In embodiments, embodiments of the invention are configured to use coordinates of an image acquired by the first imaging system 410 to define an effective FoV or to control a target point of the second imaging system 420. This is achieved by sharing coordinates between the two imaging devices / systems 410 and 420. In practical implementation, the first imaging system 410 and the second imaging system 420 include electronic or mechanical control mechanisms to control the effective FoV or target point of the respective system.
[0060] It can be envisioned that the first imaging system and the second imaging system are different imaging systems capable of acquiring corresponding images in different domains.
[0061] As an example, the invention is not limited to this example. The imaging system according to the invention can be used for autonomous driving, defensive driving and obstacle avoidance vehicles, for robot control in warehouses or manufacturing (cars, machines, processor-controlled systems, etc.) facilities, and for a variety of other applications that are readily apparent to those skilled in the art given the teachings of the invention provided herein.
[0062] Figures 5 to 6 This is a flowchart illustrating an exemplary method 500 for guided multispectral inspection according to an embodiment of the present invention.
[0063] In box 505, a neural network is trained offline (i.e., the connections between neurons are optimized) on a training dataset that includes objects of interest of predefined types.
[0064] At frame 510, initial imaging data of the scene acquired by the first imaging system 410 in the form of visible field imaging data is received.
[0065] At frame 520, objects of a predefined type in the scene acquired by the first imaging system 410 are detected. In an embodiment, objects of interest of a predefined type can be detected by applying machine learning techniques to the visible domain. In an embodiment, a visual attention-based neural network can be used, which has been trained to equip the neural network with the ability to focus on at least one region of interest.
[0066] At box 530, in response to the detection result from box 520, the coordinates of relevant objects in the scene acquired by the first imaging system 410 are extracted. At box 540, coordinates are shared between the images acquired by both the first imaging system 410 and the second imaging system 420 (hereinafter referred to as "shared coordinates").
[0067] In box 540, the second imaging system 420 receives shared coordinates, controls the target orientation (effective field of view (FoV)) of the second imaging system in response to the shared coordinates, and acquires subsequent imaging data of a portion of the scene as initial imaging data including relevant objects. Although the same coordinates are used, it is understood that, depending on the characteristics of the second imaging system and the relative position between the first and second imaging systems, overlapping areas, non-overlapping areas, smaller areas, sampling areas, etc., may be acquired as subsequent imaging data relative to the initial imaging data. In this embodiment, the second imaging system 420 is not activated if the shared coordinates are not received, thereby avoiding any energy consumption associated with unnecessarily activating the second imaging system 420.
[0068] At box 550, actions are selectively performed in response to subsequent imaging data.
[0069] For example, in one embodiment, an autonomous motor vehicle is controlled in response to subsequent imaging data. This control may include braking, steering, acceleration, etc. This control can be performed to drive the vehicle autonomously in a safe manner, and can also be used to autonomously control the vehicle via braking, steering, or acceleration for obstacle avoidance to avoid impending collisions with objects of interest. Further for this purpose, obstacles may be warned to the user so that the user can take control and avoid the obstacle, or the user may simply pay attention to the obstacle while the vehicle is automatically controlled to avoid the object. The vehicle may be a vehicle on a road, a vehicle used in a warehouse or manufacturing facility (e.g., a pallet loader, forklift, a mobile robot, etc.). For the purposes of this application, it is noteworthy that millimeter-wave imaging systems can see through common visible obstacles (such as fog and smoke). For example, a first imaging system in the visible domain can activate a second imaging system in the millimeter-wave domain and direct its target to a part of a scene containing smoke to detect whether an obstacle exists behind the smoke.
[0070] In another embodiment, the first imaging system may be located on a first vehicle (e.g., a drone (UAV) with, for example, RGB and / or IR cameras), and the second imaging system may be located in a second vehicle (e.g., with a millimeter-wave radar imaging system). The second vehicle and the onboard second imaging system may be guided to the same or similar scene acquired by the first vehicle. In this way, a more capable (second) vehicle platform can be selectively deployed based on the initial imaging data and its processing, only when needed.
[0071] Although an example involving two vehicles, one for the first imaging system 410 and the other for the second imaging system 420, has been described, in other embodiments, only one of the imaging systems 410 or 420 is vehicle-mounted, while the other is stationary in a location. In these embodiments, the stationary imaging system can still be able to rotate and perform all 3D movements to acquire the target image.
[0072] In one embodiment, the target environment could be a train platform, where a fixed or movable camera is used to capture images of the platform, from which the location information of the suspicious package is determined. A millimeter-wave scanner can be guided to the location of the package to estimate the nature of its contents, such as the presence of a large metal object.
[0073] In an embodiment, the cloud computing platform can be used to perform initial imaging data processing (e.g., determining whether to initiate a call to the second vehicle, e.g., not initiating a call to the second vehicle at all), and then perform position calculations, etc., to provide the position data to a specific controller with the task of controlling the FoV or target direction of the second imaging system, so as to acquire the region of interest within the imaging data acquired by the first imaging system. For this purpose, see below. Figure 7 and 8 Any of the various types of cloud computing platforms described in more detail can be used to perform at least a portion of the steps of method 500.
[0074] It should be understood that, for each frame 510, the steps of method 500 may be repeated for each or some images acquired by the first imaging system 410.
[0075] A further description of block 520 of the method 500 according to an embodiment of the present invention will now be given.
[0076] Object recognition is a computer vision method that aims to identify the categories of objects in an image and locate them. It is a trainable deep learning method that can also be pre-trained on larger datasets. Object recognition can be associated with segmentation algorithms that locate a given object and associate it with bounding boxes and coordinates. The object detector can then share its region of interest with a radar using appropriate calibration methods.
[0077] A further description of block 320 of method 300 according to another embodiment of the present invention will now be given.
[0078] Attention mechanisms are machine learning methods that allow for the intelligent selection of relevant portions of data (images) given a specific context. An attention mechanism is part of a neural network that needs to be trained to select specific regions of interest (hard or soft, where soft attention refers to building a graph of attention weights, while hard attention refers to picking the exact location). For some applications, attention mechanisms can focus on different regions simultaneously or sequentially. An attention mechanism can be part of a multi-step inference process that drives attention from one part of an image to another to perform a specific task. The target detector can then share the region of interest with a radar using the correct calibration methods.
[0079] Figure 7 This is a block diagram illustrating an exemplary neural network configuration 700 used by system 400 according to an embodiment of the present invention. Specifically, neural network configuration 600 is used by sensor control and data processing subsystem 430.
[0080] Neural network configuration 700 relates to a neural network 710 that is trained offline on a training dataset that includes objects of interest of predefined types (i.e., the connections between neurons have been optimized).
[0081] The neural network configuration 700 also relates to a visual attention-based neural network 710, which has been trained to equip the neural network 720 with the ability to focus on at least one region of interest.
[0082] The neural network 710 is implemented as including a visual attention-based neural network corresponding to the first modality. The neural network 710 is connected to the second modality pipeline 720. Figure 7 In one embodiment, the second mode pipeline is a millimeter-wave radar processing pipeline. In other embodiments, other types of modes may be used for either the first or second mode.
[0083] Neural network 710 receives an input image 711 acquired by an image sensor. The input image 711 is represented by a matrix having one or more wavelength channels (e.g., RGB, etc.). The size of the input image 711 is adjusted in a resizing operation 712, and then the input image 711 is fed into neural network 710. Neural network 710 includes at least one 2D convolutional layer (collectively indicated by reference numeral 713) and multiple repetitions of subsequent 2D pooling layers 714. Neural network 710 also includes a final 2D convolutional layer 715, followed by a fully connected layer 716, which is connected to an output layer 717 for the final result.
[0084] The output layer 717 includes two units, namely yaw 717A and pitch 717B, with positive and negative values that match the equivalent values of the second image sensor (e.g., a millimeter-wave radar imager).
[0085] Relative to the millimeter-wave radar image processing pipeline 720, yaw 717A and pitch 717B values from the neural network 710 are used to center the effective FoV of the millimeter-wave radar 721. The millimeter-wave imager operates by scanning the beam across the FoV. A recurrent neural network is used as the millimeter-wave radar processing pipeline 720 to process the radar readout sequence 722. Therefore, the second modality-based neural network 720 includes one or more recurrent layers (collectively indicated by reference numeral 723 in the accompanying drawings) and a fully connected layer 724 connected to the output layer 725 for output classification.
[0086] Therefore, for the first modality related to the image, a first-mode acquisition device (e.g., an RGB camera) and visual attention can be used to detect packets on a detection platform within the first region of interest (ROI) as the first ROI. Location information associated with the ROI can be used to guide a second-mode acquisition device (e.g., millimeter-wave radar). The output layer can determine whether the packet content represents a potential threat or not, and classify it as an output.
[0087] It is understood that while this disclosure includes a detailed description of cloud computing, implementations of the teachings cited herein are not limited to cloud computing environments. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.
[0088] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage devices, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0089] The features are as follows:
[0090] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring human interaction with the service provider.
[0091] Extensive network access: Capabilities are available through networks and accessed via standard mechanisms that facilitate the use of heterogeneous thin client platforms or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0092] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated as needed. There is a sense of location independence because consumers typically do not have control or knowledge of the exact location of the resources provided, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0093] Rapid flexibility: The ability to provide capacity quickly and flexibly, automatically scaling down and up rapidly in some situations to scale up rapidly. For consumers, the available supply capacity often appears unlimited and can be purchased in any quantity at any time.
[0094] Measuring services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0095] The service model is as follows:
[0096] Software as a Service (SaaS): This provides consumers with the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage devices, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.
[0097] Platform as a Service (PaaS): This provides consumers with the ability to deploy applications created or acquired by the consumer using programming languages and tools supported by the provider onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environment.
[0098] Infrastructure as a Service (IaaS): This provides consumers with the capability to deliver processing, storage, networking, and other basic computing resources that enable them to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but rather have control over the operating system, storage devices, deployed applications, and potentially limited control over selected networking components (e.g., host firewalls).
[0099] The deployment model is as follows:
[0100] Private cloud: A cloud infrastructure for organization operations only. It can be managed by the organization or a third party and can exist on-site or off-site.
[0101] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.
[0102] Public cloud: Makes cloud infrastructure available to the public or large industry groups and is owned by an organization that sells cloud services.
[0103] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported (e.g., cloud bursting for load balancing between clouds).
[0104] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure comprising a network of interconnected nodes.
[0105] See now Figure 8 This describes an illustrative cloud computing environment 850. As shown, the cloud computing environment 850 includes one or more cloud computing nodes 810, whose local computing devices used by cloud consumers (such as, for example, personal digital assistants (PDAs) or cellular phones 854A, desktop computers 854B, laptop computers 854C, and / or automotive computer systems 854N) can communicate with the cloud computing nodes 810. The nodes 810 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 850 to provide infrastructure, platform, and / or software as services that cloud consumers do not need to maintain on their local computing devices. It is understood that... Figure 8 The types of computing devices 854A-N shown are merely illustrative, and the computing node 810 and cloud computing environment 850 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).
[0106] See now Figure 9 This demonstrates the 850 (cloud computing environment) Figure 8 This provides a set of functional abstractions. It is understandable in advance. Figure 9 The components, layers, and functions shown are merely illustrative, and embodiments of the invention are not limited thereto. As described, the following layers and corresponding functions are provided:
[0107] The hardware and software layer 960 includes hardware and software components. Examples of hardware components include: a mainframe 961; a server 962 based on a RISC (Reduced Instruction Set Computer) architecture; a server 963; a blade server 964; a storage device 965; and a network and network components 966. In some embodiments, software components include network application server software 967 and database software 968.
[0108] The virtualization layer 970 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 971; virtual storage 972; virtual networks 973, including virtual private networks; virtual applications and operating systems 974; and virtual clients 975.
[0109] In one example, management layer 980 can provide the functions described below: Resource Provisioning 981 Provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 982 Provides cost tracking as resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security Provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User Portal 983 Provides consumers and system administrators with access to the cloud computing environment. Service Level Management 984 Provides cloud resource allocation and management to ensure that required service levels are met. Service Level Agreement (SLA) Planning and Fulfillment 985 Provides pre-scheduling and procurement of cloud resources, anticipating future requirements for those resources according to the SLA.
[0110] Workload layer 990 provides examples of functionalities that can be leveraged in a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 991; software development and lifecycle management 992; virtual classroom education delivery 993; data analytics and processing 994; transaction processing 995; and guided multispectral inspection 996.
[0111] This invention can be a system, method, and / or computer program product with any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the invention.
[0112] Computer-readable storage media can be tangible means for retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punched cards, or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.
[0113] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0114] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this invention.
[0115] The present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0116] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in the boxes or blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in the boxes or blocks of a flowchart and / or block diagram.
[0117] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, such that the instructions executed on the computer, other programmable apparatus, or other device perform the functions / actions specified in the boxes or blocks of a flowchart and / or block diagram.
[0118] References to the invention in this specification as "one embodiment" or "embodiment" and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment of the invention. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing in various places throughout the specification, and any other variations, do not necessarily refer to the same embodiment. However, it is understood that features of one or more embodiments can be combined in light of the teachings of the invention provided herein.
[0119] Understandably, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of” is intended to include selecting only the first listed item (A), or only the second listed item (B), or selecting both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed items (A and C), or only the second and third listed items (B and C), or selecting all three options (A, B, and C). This can be extended to up to the listed items.
[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the figures. For example, two consecutively shown blocks may actually be completed as a single step, executed simultaneously, substantially simultaneously, or with partial or complete temporal overlap, or the blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0121] Preferred embodiments of the system and method have been described (these are intended to be illustrative and not restrictive), and it should be noted that modifications and variations can be made by those skilled in the art based on the above teachings. Therefore, it is understood that changes may be made to the specific embodiments disclosed within the scope of the invention as outlined in the appended claims. Various aspects of the invention, having the details and features required by patent law, have thus been described, and the claimed and desired protection by a patent certificate is set forth in the claims.
Claims
1. An imaging system, comprising: The first imaging system is used to acquire initial sensor data in the form of visible field data; A second imaging system mounted on the vehicle is used to acquire subsequent sensor data in the form of second-domain data, wherein the initial sensor data and the subsequent sensor data have different spectral domains; and A controller subsystem, operably coupled to the first imaging system and the second imaging system, detects at least one region of interest in the scene in real time by applying a visual attention-based neural network to the visible domain data. The visual attention-based neural network is used to locate at least one object of interest in the at least one region of interest, thereby generating location data for the at least one object of interest. In response to the location data, the vehicle is autonomously deployed into the scene, and the focus of the second imaging system is directed to the area of the scene including the object of interest, in order to acquire the second domain data; The first imaging system and the second imaging system are configured to share a common set of x-coordinates and y-coordinates of any scene acquired therefrom, wherein the common set is used to generate position information, the position information is used to control the effective field of view (FoV) of the second imaging system, and the controller subsystem is further configured to dynamically adjust the effective field of view (FoV) of the second imaging system based on real-time coordinate changes detected by the first imaging system.
2. The imaging system of claim 1, further comprising autonomously providing depth information and material properties of the at least one object.
3. The imaging system according to claim 1, wherein, The first imaging system includes an imaging device selected from the group consisting of an RGB camera, an infrared camera, and a thermal camera.
4. The imaging system according to claim 3, wherein, The second imaging system includes a radar imaging system.
5. The imaging system according to claim 3, wherein, The second imaging system includes a millimeter-wave imaging system.
6. The imaging system according to claim 1, wherein, Locating the at least one object of interest includes extracting the coordinates of the at least one object of interest.
7. The imaging system according to claim 1, wherein, The second domain data includes information from scenes that cannot be observed in the visible domain.
8. The imaging system of claim 1, further comprising controlling an autonomous driving vehicle in response to subsequent imaging data.
9. The imaging system according to claim 1, wherein, The first imaging system is located on a drone, and the first imaging system includes an imaging device selected from the group consisting of an RGB camera, an infrared camera, and a thermal camera.
10. The imaging system according to claim 1, wherein, The first imaging system is a thermal camera, which is used to detect temperature hotspots in the scene as at least one region of interest.
11. The imaging system according to claim 1, wherein, The controller subsystem is configured for cloud computing.
12. A method for imaging, comprising: Initial sensor data, acquired by the first imaging system in the form of visible field data; Subsequent sensor data, acquired in the form of second domain data, by a second imaging system mounted on the vehicle, wherein the initial sensor data and the subsequent sensor data have different spectral domains; The controller subsystem detects at least one region of interest in the scene in real time by applying a vision-focused neural network to the visible domain data. The controller subsystem uses the visual attention-based neural network to locate at least one object of interest in the at least one region of interest, thereby generating location data for the at least one object of interest; and In response to the location data, the controller subsystem autonomously deploys the vehicle into the scene and directs the focus of the second imaging system to the area of the scene that includes the object of interest, in order to acquire the second domain data; The first imaging system and the second imaging system are configured to share a common set of x-coordinates and y-coordinates of any scene acquired therefrom, wherein the common set is used to generate position information, the position information is used to control the effective field of view (FoV) of the second imaging system, and the controller subsystem is further configured to dynamically adjust the effective field of view (FoV) of the second imaging system based on real-time coordinate changes detected by the first imaging system.
13. The method according to claim 12, wherein, The second domain data includes information from scenes that cannot be observed in the visible domain.
14. The method of claim 12, further comprising autonomously providing depth information and material properties of the at least one object.
15. The method according to claim 12, wherein, Locating the at least one object of interest includes extracting the coordinates of the at least one object of interest.
16. The method of claim 12, further comprising controlling the autonomous driving motor vehicle in response to subsequent imaging data.
17. A computer program product for imaging, the computer program product comprising program instructions executable by a computing system to cause the computing system to perform the method as claimed in any one of claims 12-16.
Citation Information
Patent Citations
Video system and methods for operating a video system
US20030210329A1
Vehicle control device mounted on vehicle and method for controlling the vehicle
US20190144001A1
Low-profile multi-band hyperspectral imaging for machine vision
US20200019039A1
Spatial and temporal attention-based deep reinforcement learning of hierarchical lane-change policies for controlling an autonomous vehicle
US20200139973A1