Apparatus and method for determining physical properties of physical objects

CN114067166BActive Publication Date: 2026-09-25ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110855734.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-29
Filing Date
2021-07-28
Publication Date
2026-09-25
Estimated Expiration
2041-07-28

AI Technical Summary

Technical Problem

由于它们的根本不同的工作方式,传统的图像处理方法、尤其是用于图像处理的神经网络的传统架构对于处理由基于事件的摄像机提供的图像数据来说无效

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067166B_ABST
    Figure CN114067166B_ABST
Patent Text Reader

Abstract

The invention relates to a method of determining a physical property of a physical object, having: for each input time point of a sequence of input time points, detecting sensor data with an event-based sensor, which contains information of the physical object; for each sub-sequence of input time points: feeding the detected sensor data to a spiking neural network, generating a first processing result by processing these sensor data with spiking neurons of the spiking neural network, wherein the spiking neurons integrate values of the sensor data, respectively, feeding the first processing result to a non-spiking neural network, generating a second processing result by processing the first processing result with non-spiking neurons of a first layer of the non-spiking neural network; feeding the second processing result to a second layer, combining the processing results of the sub-sequences by non-spiking neurons of the second layer in such a way that these non-spiking neurons calculate a weighted sum of values of the second processing result of different sub-sequences, respectively; and determining the physical property of the physical object from an output of the non-spiking neurons.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In general, the various embodiments relate to apparatus and methods for determining the physical properties of a physical object. Background Technology

[0002] In the case of event-based cameras, the entire image—that is, all pixels of a frame—is not always transmitted; instead, only the pixel values ​​that have changed are transmitted. This eliminates the need for processing to wait for the entire image (frame) to be constructed. Such event-based cameras are significantly more suitable than traditional cameras in certain applications, such as motion estimation in autonomous driving scenarios, due to their high temporal resolution. Because of their fundamentally different operating methods, traditional image processing methods, especially the traditional architectures of neural networks used for image processing, are ineffective for processing image data provided by event-based cameras.

[0003] Correspondingly, for different tasks, such as monitoring in production (e.g., detecting defective components) or identifying objects in the context of autonomous driving, efficient processing of image data provided by event-based cameras (or sensor data from event-based sensors) is desirable. Summary of the Invention

[0004] According to different implementations, a method for determining the physical characteristics of a physical object is provided, the method comprising: for each input time point of an input time point sequence, detecting sensor data using an event-based sensor, wherein the sensor data contains information about one or more physical objects; and for each subsequence in which the input time point sequence is decomposed into a plurality of subsequences, the method comprising: - The sensor data detected at the input time point for this subsequence is fed into the spiking neural network; - The sensor data at the input time points of the subsequence is processed by spiking neurons of a spiking neural network to generate the first processing result of the subsequence, wherein these spiking neurons integrate the values ​​of the sensor data at different input time points of the subsequence; - The processing result of this subsequence is fed into a non-spiking neural network; and - The processing result of the subsequence is processed by one or more non-pulsing neurons in the first layer of a non-pulsing neural network to generate a second processing result of the subsequence; Moreover, this method also has The second processing results of multiple subsequences are fed into one or more second layers of a non-spiking neural network; the processing results of the multiple subsequences are combined by non-spiking neurons of one or more second layers of the non-spiking neural network, in that the non-spiking neurons of one or more second layers calculate a weighted sum of the values ​​of the second processing results of different subsequences; and the physical properties of the physical object are determined based on the output of the non-spiking neurons of one or more second layers.

[0005] The method described above can efficiently process sensor-based sensor data, particularly eliminating the need for accumulation of sensor data over time, for example, at the input of a spiking neural network. This avoids or at least reduces the loss of temporal information. Furthermore, the method can utilize efficient and proven architectures of non-spiking neural networks, such as convolutional networks, to perform, for example, classification or estimation of optical flow in sensor data (e.g., for motion estimation).

[0006] Different implementations are described below.

[0007] Example 1 is a method for determining the physical properties of a physical object as described above.

[0008] Example 2 is the method according to Example 1, wherein an event-based sensor provides sensor data for multiple components of a sensor data vector and feeds it to non-pulsing neurons of a non-pulsing neural network to generate a second processing result of the subsequence, wherein the non-pulsing neurons calculate a weighted sum of the values ​​of the second processing result of different components of the sensor data vector.

[0009] Therefore, in non-spiking neural networks, analysis is performed across the components of sensor data (e.g., spatial components or frequency components), for example, using convolutional networks. This allows for the use of efficient classification or regression techniques.

[0010] Example 3 is the method according to Example 1 or 2, wherein the input time point is the time point of the event, and the event-based sensor reacts to these events by outputting sensor data.

[0011] This avoids a reduction in temporal resolution relative to the sensor's temporal resolution, thereby preserving the maximum amount of information.

[0012] Example 4 is a method according to any one of Examples 1 to 3, the method comprising: for each subsequence, feeding at least a portion of the output of a spiking neuron to another neural network, wherein the other neural network performs classification, the classification determining the end of the subsequence.

[0013] Example 5 is a method according to any one of Examples 1 to 4, the method comprising: for each subsequence, determining at least a portion of the output of these spiking neurons in each time unit; and terminating the subsequence if the output of these spiking neurons in each time unit exceeds a threshold.

[0014] Examples 4 and 5 intuitively provide a "trigger" for processing via a non-spiking neural network (i.e., a trigger for forward-passing via a non-spiking neural network). This achieves the goal that resources are only allocated to processing via the neural network when it is "worthwhile," that is, only when new information is expected. Thus, a large number of spiking signals in the SNN can, for example, indicate that a new object has entered the sensor's range and therefore processing via the ANN may provide new information (e.g., about the physical properties of the new object).

[0015] Example 6 is a method for controlling an actuator, the method comprising: determining the physical characteristics of a physical object according to any one of Examples 1 to 5; and controlling the actuator according to the determined physical characteristics of the physical object.

[0016] Example 7 is a method for detecting anomalies in a physical object, the method comprising: determining physical characteristics of the physical object according to any one of Examples 1 to 5; comparing the determined physical characteristics of the physical object with reference characteristics of the physical object; and if the determined physical characteristics are different from the reference characteristics, then identifying that the physical object has an anomaly.

[0017] Example 8 is an apparatus configured to implement the method according to any one of Examples 1 to 7.

[0018] Example 9 is a computer program having program instructions that, when implemented by one or more processors, cause the one or more processors to perform the method according to any one of Examples 1 to 7.

[0019] Example 10 is a computer-readable storage medium having stored thereon program instructions that, when implemented by one or more processors, cause the one or more processors to perform the method according to any one of Examples 1 to 7. Attached Figure Description

[0020] Embodiments of the invention are shown in the accompanying drawings and described in detail below. In the drawings, the same reference numerals generally refer to the same parts in various views. These drawings are not necessarily to scale, but their emphasis is generally on illustrating the principles of the invention.

[0021] Figure 1 A vehicle according to an embodiment is shown.

[0022] Figure 2 A neural network according to an embodiment is shown.

[0023] Figure 3 The processing of sensor data from an event-based camera (or typically an event-based sensor) over time according to an implementation method is explained.

[0024] Figure 4 A flowchart is shown, which presents a method for determining the physical properties of a physical object according to an embodiment. Detailed Implementation

[0025] Different implementations, particularly the embodiments described below, can be implemented using one or more circuits. In one implementation, "circuit" can be understood as any type of logical implementation entity, which can be hardware, software, firmware, or a combination thereof. Thus, in one implementation, "circuit" can be hardwired logic circuitry or programmable logic circuitry, such as a programmable processor, for example, a microprocessor. "Circuit" can also be software implemented or carried out by a processor, such as any type of computer program. According to an alternative implementation, any other type of implementation of the corresponding functionality can be understood as "circuit," and these corresponding functionalities are described in more detail below.

[0026] In the context of machine learning, a function is learned that maps input data (e.g., sensor data) to output data. In the case of learning (i.e., training a neural network or other model), this function is determined based on an input dataset (also called a training dataset) for which the desired output (e.g., the desired classification of the input data) is given in advance for each input, such that the function maps the assignment from input to output as well as possible.

[0027] Examples of applying this machine learning capability include object classification or motion estimation for autonomous driving, such as... Figure 1 As explained in [the document / document].

[0028] Figure 1 Vehicle 101 is shown.

[0029] It should be noted that, in the following text, an image or image data is very generally understood as a set of data representing one or more objects or patterns. Image data can be provided by sensors (especially event-based sensors) that measure visible or invisible light, such as infrared or ultraviolet light, ultrasonic or radar waves, or other electromagnetic or sound signals.

[0030] exist Figure 1 In the example, vehicle 101, such as a passenger vehicle (PKW) or a freight vehicle (LKW), is equipped with vehicle control device 102.

[0031] The vehicle control unit 102 has data processing components, such as a processor (e.g., a CPU (central unit)) 103 and a memory 104, which stores the control software that the vehicle control unit 102 operates according to and the data processed by the processor 103.

[0032] For example, the stored control software has (computer program) commands that, when executed by the processor, cause the processor 103 to implement one or more neural networks 107.

[0033] The data stored in memory 104 may include, for example, image data detected by one or more cameras 105. The one or more cameras 105 may, for example, take one or more grayscale or color photographs of the environment surrounding vehicle 101.

[0034] The vehicle control unit 102 can determine, based on image data, whether there are objects and what objects, such as stationary objects like traffic signs or road markings or moving objects like pedestrians, animals and other vehicles, exist in the environment surrounding the vehicle 101 and at what speed these objects are moving relative to the vehicle 101.

[0035] Next, vehicle 101 can be controlled by vehicle control device 102 according to the result of object determination. In this way, vehicle control device 102 can, for example, control actuator 106 (e.g., brake) to control the speed of the vehicle, for example, to brake the vehicle.

[0036] Camera (or one of these cameras) 105 is, for example, an event-based camera. Event-based cameras are a novel camera technology that can produce a fast, high temporal resolution, and efficient representation of the real world. Unlike conventional cameras, event-based cameras output an event stream, where each event represents a binary change in brightness at a specific pixel at a specific point in time. This is particularly well-suited for high-speed motion estimation and low-light navigation, applications where conventional cameras (which provide image data frame-by-frame) face difficulties.

[0037] Currently, non-spiking artificial neural networks (ANNs) are the standard for processing image data. However, these ANNs are only conditionally suited for processing event data because they typically process input data synchronously and using densely padded matrices, whereas event data is sparsely padded and asynchronous.

[0038] One possibility for processing event data is to reduce the temporal resolution by accumulating a certain number of events. In this way, the raw data provided by the event-based camera 105 is transformed into a sequence of densely padded matrices. This representation can be used to teach a conventional neural network (ANN) capable of reconstructing grayscale images. Common machine vision algorithms can then be applied to it. However, this approach limits the temporal resolution due to the accumulation process. Consequently, some information is lost.

[0039] The following describes an embodiment that achieves high temporal resolution and is suitable for general tasks.

[0040] According to different implementation methods, a two-part neural network (e.g., as neural network 107) is used, such as... Figure 2 As shown in the image.

[0041] Figure 2 A neural network 200 according to an embodiment is shown.

[0042] The neural network 200 uses an input layer 202 to obtain data from an event-based sensor (e.g., camera 105). The output of the neural network 200 is achieved through an output layer 206.

[0043] Between the input layer 202 and the output layer 206, the neural network 200 has a sequence of layers 203, 204, and 205.

[0044] The neural network 200 consists of two parts because its layer sequence has a first set of layers 203 (the first subsequence of the layers) and a second set of layers 205 (the second subsequence of the layers). In this example, the two sets are connected by an (optional) intermediate layer 204.

[0045] The first group of 202 layers consists of spiking (peak) neurons, while the second group of 202 layers has traditional artificial neurons (that is, non-spiking neurons).

[0046] A spiking neuron has an internal state V(t), which can change over time t. Communication between a spiking neuron and subsequent neurons occurs through pulses (so-called spikes) emitted by the spiking neuron, typically binary signals, although signals with multiple bits of information are also possible.

[0047] If the internal state of a spiking neuron exceeds or falls below a certain threshold, the spiking neuron emits a pulse. This method of information transmission is inspired by how the human brain processes information.

[0048] Neural networks can, for example, output estimates of optical flow in a scene captured by an event-based camera. End-to-end training can be performed using backpropagation (for example, for this task, but also for classification or regression), where an approximation of gradient computation can be used for spiking neurons to enable stable computation of gradients.

[0049] In continuous time, spiking neurons can be described using the Leaky Integrated-and-Fire (LIF) model: in this case, It is the membrane time constant. It is the membrane voltage. It is the resting potential. It is a resistor, and Input current l It is the index of the layer where the neuron is located. i It is the index of that neuron.

[0050] According to this model, the spiking neuron does not lose voltage over time, that is... This is called the integrated distribution model.

[0051] The input current is typically a sum of multiple pulses, obtained through a Dirac delta function and weights. To model this, the weights are weighted according to the synapse from another (previous) neuron to this neuron: Activation function Determine whether the neuron sends a pulse to one or more subsequent neurons connected to it.

[0052] To simulate this model, time discretization is performed (where U... rest In this example, it is set to zero. This results in the following formula: Equation (1) describes the input current, and equation (2) describes (with index) The layer (with index) The voltage of the neuron, which is measured in time steps. With neurons (with corresponding indexes) Each neuron is connected to the network.

[0053] In the sum, and It is the weight, where the weight It should be noted that neurons in one layer can also connect to themselves.

[0054] Having a positive time constant item It is the decreasing intensity, where 0 < < 1. Applicable ,in It is the membrane voltage Nonlinear functions, such as , where Θ is the Theta function (or Heaviside function).

[0055] constant By having a positive time constant of This is given. According to the implementation method, the limiting cases α = 1 and β = 1 are also specifically used.

[0056] To train a neural network using backpropagation (more precisely, BPTT, Backpropagation through time), the derivative of the activation function needs to be calculated. However, since the activation function is described by the Theta function, a simple approximation of this derivative will result in the derivative being zero at every position, thus making it impossible to calculate weight updates.

[0057] If an auxiliary function is used for the calculation of this update, for example This can avoid the situation.

[0058] This enables the computation of updates to the weights of the SNN (i.e., the spiking neurons of the first layer 203) (i.e., training updates). In this way, the neural network 200 can be trained, where the weights of the ANN (i.e., the weights of the neurons in the second layer 205) are updated using common backpropagation, while the updates to the weights of the SNN neurons are computed using the aforementioned approximation.

[0059] The event stream from an event-based camera can be partitioned into multiple input vectors (or input tensors), where the partitioning is based on a pre-given time step dt. These input vectors are used to train a neural network using a hybrid SNN-ANN architecture, such as... Figure 2 As shown in the image.

[0060] According to the implementation method, the architecture has at least some of the following features: - An SNN consists of one or more layers that can be arbitrarily connected to each other. The delay of each connection can also be arbitrarily adjusted. - The characteristic of SNN lies in the communication between event packets and neurons of the SNN that have internal states. - An ANN outputs predictions based on the output vector of an SNN. This could be, for example, the internal state of the last layer of the SNN, or accumulated spikes from different layers. Between an SNN and an ANN, there can be a so-called trigger that determines when to implement the ANN. This can be achieved, for example, by training an ANN or SNN, in addition to its existing structure, to predict, for example, whether a new object is seen in the input vector or whether an object has moved. - The bandwidth between SNN layers and between SNN and ANN can be adjusted, for example, by integrating additional loss terms used to generate spikes / events into the optimization or by always using only a portion of the network.

[0061] Alternatively, instead of or in combination with backpropagation, SNNs can be taught using local learning methods known from neuroscience, such as spike timing-related plasticity (STDP). Neural networks can also be taught end-to-end.

[0062] For training, a collection of event data from event-based sensors is collected. The neural network is implemented and initialized (e.g., with random values).

[0063] Data output from an SNN can be preprocessed before being processed by an ANN. For example, spiking or (e.g., in transition layer 204) the internal states of SNN neurons can be accumulated.

[0064] In different implementations, the SNN processes the input data from an event-based camera, which provides this input data within a certain time frame—that is, for a subsequence of input time points (of the total sequence of input time points)—and outputs the results to the ANN. The ANN processes the input data for different subsequences separately (that is, independently), and then combines these processing results (for different subsequences) together, such as... Figure 3 As shown in the image.

[0065] Figure 3 It clarifies the processing of sensor data from event-based cameras (or typically event-based sensors) over time.

[0066] Time flows from left to right along time axis 305.

[0067] From top to bottom, the layer depth increases, meaning that the blocks shown below, representing data within neural network 200, are located in deeper layers of neural network 200 (i.e., in...). Figure 2 (The rightmost part of the illustration).

[0068] The top three rows of 306 represent the processing in SNN, and the bottom four rows of 307 represent the processing in ANN.

[0069] The topmost layer corresponds to the first layer of the SNN. Each block in the first layer corresponds to receiving sensor input data from an event-based sensor at the input time point where an event occurs, with the sensor providing input data as a response to the event. Alternatively, input data can be collected over a duration of a (small) time step dt, such that the input time points are arranged regularly on the time axis, for example, at intervals dt. The value dt is, for example, in the range of 0.1 ms to 1 ms, but depending on the application, it can be in the range of 0.01 ms to 10 ms in more extreme cases.

[0070] In this example, these input time points are grouped into groups 301, 302, 303, and 304. The input data received by the SNN at the input time points of groups 301, 302, 303, and 304 is processed by the SNN, and then the common processing results of the corresponding groups are forwarded to the ANN in 308, 309, 310, or 311.

[0071] As shown, the input data of the previous group can also affect the processing of the input data of the next group (as indicated by arrows extending beyond the group boundaries). However, at certain time points 308, 309, 310, and 311, the results are transferred from the SNN to the ANN, and the input data is added to the ANN until the end of the group to which it belongs (i.e., the first group 301 for the first forwarding time point 308).

[0072] ANN combines the processing results of multiple groups, such as the processing results of the first group 301 and the second group 302 in 312, the processing results of the third group 303 and the fourth group 304 in 313, and finally the processing results of all four groups in 314.

[0073] The final result is the processing of sensor data from four groups, 301 to 304.

[0074] The delay between two SNN layers is, for example, a time step. In an ANN, the delay is zero. As shown, there is a connection between the first and third layers of the SNN in this example. The amount of data may also vary between layers, as in an encoder-decoder architecture. For example, the amount of data decreases in the third SNN layer.

[0075] In general, depending on the implementation method, it provides, for example, in Figure 4 The method shown in the figure.

[0076] Figure 4 Flowchart 400 is shown, which illustrates a method for determining the physical properties of a physical object.

[0077] In 401, for each input time point in the input time point sequence, event-based sensors are used to detect sensor data, which contains information about one or more physical objects.

[0078] In 402, for each subsequence in the process of decomposing the input time-point sequence into multiple subsequences... - In 403, the sensor data detected at the input time point for this subsequence is fed into a spiking neural network (SNN). In step 404, the sensor data at the input time points of the subsequence is processed by spiking neurons of a spiking neural network to generate the first processing result of the subsequence. These spiking neurons integrate the values ​​of the sensor data at different input time points of the subsequence. - In 404, the processing result of this subsequence is fed into a non-spiking neural network (ANN), and - In 405, the processing result of the subsequence is processed by one or more non-spiking neurons of the first layer of a non-spiking neural network to produce a second processing result of the subsequence.

[0079] In step 406, the second processing results of the plurality of subsequences are fed into one or more second layers of a non-spiking neural network.

[0080] In 407, the processing results of the multiple subsequences are combined by one or more non-pulsing neurons in the second layer of a non-pulsing neural network, in that the one or more non-pulsing neurons in the second layer calculate a weighted sum of the values ​​of the second processing results of different subsequences.

[0081] In 408, the physical properties of the physical object are determined based on the output of the non-spiking neurons of the one or more second layers.

[0082] Event-based sensors include event-based cameras or other event-based sensors, such as neuromorphic cochlea or IoT (Internet of Things) devices, which send information triggered by events.

[0083] The output of a neural network (that is, information output regarding the determined physical characteristics of a physical object) can be used to control devices or systems, such as robots, based on the output of the neural network. A "robot" can be understood as any physical system (with mechanical parts whose movement is controlled), such as computer-controlled machines, vehicles, household appliances, power tools, production machines, personal assistants, or access control systems. The output of a neural network can also be used in systems that transmit information, such as those used in medical (imaging) methods.

[0084] Neural networks can be used for regression or classification of data provided by event-based sensors, meaning that physical characteristics can be determined based on the classification or regression results. This data is provided, for example, by event-based cameras and represents a scene. Here, the term classification also includes semantic segmentation, such as semantic segmentation of an image (which can also be considered pixel-by-pixel classification). Similarly, the term classification includes detection, such as object detection (which can be considered a classification of the presence or absence of an object). Examples of applications include: real-time location and tracking of cars and pedestrians in road traffic; tracking people on surveillance cameras in industrial plants; or determining the position of oil droplets moving at high speed on a water film.

[0085] The output of a neural network (e.g., a classification) can be used (e.g., by a corresponding control device) to detect anomalies, such as to detect defects in manufactured components.

[0086] This physical property can be determined by correspondingly interpreting the outputs of the one or more non-spiking neurons in the second layer. For example, the outputs of the one or more non-spiking neurons in the second layer describe the softmax value of the category attribute. The one or more non-spiking neurons in the second layer may also output values ​​that specify motion vectors, such as motion maps, that is, specifying the direction and speed of motion for each pixel of the image. Therefore, this physical property can be determined by reading from the outputs of the one or more non-spiking neurons in the second layer. However, further processing can also be provided, such as averaging within certain regions (which are given by the segmentation output by the neural network), etc.

[0087] According to the implementation method, this method is computer-based.

[0088] Although the invention has been shown and described primarily with reference to specific embodiments, those skilled in the art will understand that numerous changes can be made to its design and details without departing from the spirit and scope of the invention as defined by the following claims. Therefore, the scope of the invention is determined by the appended claims and is intended to cover all changes falling within the literal meaning or equivalent scope of the claims.

Claims

1. A method for determining the physical properties of a physical object, the method comprising: For each input time point in the input time point sequence, event-based sensors are used to detect sensor data, wherein the sensor data contains information about one or more physical objects; For each subsequence in which the input time point sequence is decomposed into multiple subsequences The sensor data detected at the input time points of the sub-sequence are fed into the spiking neural network. The sensor data at the input time points of the subsequence are processed by the spiking neurons of the spiking neural network to generate a first processing result for the subsequence, wherein the spiking neurons integrate the values ​​of the sensor data at different input time points of the subsequence. The processing results of the subsequence are fed into a non-spiking neural network, and The processing result of the subsequence is processed by one or more non-pulsing neurons in the first layer of the non-pulsing neural network to generate a second processing result of the subsequence; The second processing results of the plurality of subsequences are fed into one or more second layers of the non-spiking neural network; The processing results of the multiple sub-sequences are combined using one or more non-pulsing neurons in the second layer of the non-pulsing neural network, wherein the one or more non-pulsing neurons in the second layer respectively calculate a weighted sum of the values ​​of the second processing results of different sub-sequences; and The physical properties of the physical object are determined based on the outputs of the one or more non-spiking neurons in the second layer. The trigger determines when to implement the non-spiking neural network processing, wherein the trigger corresponds to a spiking neural network that generates a threshold number of pulses.

2. The method of claim 1, wherein the event-based sensor provides sensor data for multiple components of a sensor data vector and provides it to non-pulsing neurons of the non-pulsing neural network to generate a second processing result of the subsequence, wherein the non-pulsing neurons calculate a weighted sum of the values ​​of the second processing result for different components of the sensor data vector.

3. The method according to claim 1 or 2, wherein the input time point is the time point of the event, and the event-based sensor responds to the event by outputting sensor data.

4. The method according to any one of claims 1 to 3, the method comprising: for each subsequence, feeding at least a portion of the output of the spiking neuron to another neural network, wherein the other neural network performs classification, the classification determining the end of the subsequence.

5. The method according to any one of claims 1 to 4, the method comprising: for each subsequence, determining at least a portion of the output of the spiking neuron in each time unit; and terminating the subsequence if the output of the spiking neuron in each time unit exceeds a threshold.

6. A method for controlling an actuator, the method comprising: The method according to any one of claims 1 to 5 is used to determine the physical properties of a physical object; and The actuator is controlled based on the determined physical characteristics of the physical object.

7. A method for detecting anomalies in a physical object, the method comprising: The method according to any one of claims 1 to 5 is used to determine the physical properties of a physical object; The determined physical characteristics of the physical object are compared with the reference characteristics of the physical object; and If the determined physical characteristics differ from the reference characteristics, then the physical object is found to be abnormal.

8. An apparatus configured to perform the method according to any one of claims 1 to 7.

9. A computer program product having program instructions that, when implemented by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon program instructions that, when implemented by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 7.