Methods and systems for verification

EP4652537A1Pending Publication Date: 2025-11-26OPENORIGINS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024701393
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-20
Filing Date
2024-01-19
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

Current authentication systems are vulnerable to spoofing attacks, particularly the Picture-of-Picture (PoP) attack, where an adversary creates or modifies visual content to deceive viewers or applications, as they cannot differentiate between authentic and spoofed content captured by trusted devices.

Method used

A method and system that detect movement of an image sensing device before, during, and after image or video capture using sensors like gyroscopes, accelerometers, and depth sensors to determine if the device is stationary or being held, thereby identifying potential spoofing attacks by analyzing movement thresholds.

Benefits of technology

Effectively prevents spoofing attacks by distinguishing between real-world scene captures and pre-recorded or displayed content, enhancing the authenticity verification of visual information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024051261_25072024_PF_FP_ABST
    Figure EP2024051261_25072024_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method comprises detecting a movement of an image sensing device in a time period immediately before, during and / or immediately after a capture of at least one image or video by the image sensing device, detecting a spoofing attack based on analysis of the detection of any movement of the image sensing device.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR VERIFICATIONBackground

[0001] Embodiments of the present invention relate to methods, apparatus and systems for verification in general, and, in particular, verification of image captures for the purpose of preventing spoofing attacks.

[0002] Digital media comprising visual information, such as images and videos, have wide applications in the information age. An authentication system normally assumes that content of an image or video is genuine, especially if the image or video is captured by a device or application trusted by the viewer.

[0003] However, capturing of visual information is subject to various ‘spoofing’ attacks, in which visual information may be created, modified or reproduced by an unauthorised adversary or program to gain an unfair or even illegitimate advantage. These attacks attempt to present a false scene to a viewer. The false scene may be created by artificially setting up a scene which mimics the scene expected by a viewer or authenticator. There has been an ongoing need for methods for verifying whether images are captured in a manner for spoofing a viewer.

[0004] The embodiments described below are not limited to implementations which solve any or all of the disadvantages of known techniques for verification or spoofing detection.Summary

[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0006] The invention is set out in the appended set of claims.

[0007] A first aspect provides a computer-implemented method comprising: detecting a movement of an image sensing device in a time period immediately before, during and / or immediately after a capture of at least one image or video by the image sensing device, detecting a spoofing attack based on analysis of the detection of any movement of the image sensing device. A second aspect provides an apparatus, comprising at least one processor and at least one memory, the memory storing computer-implementable instructions, when executed by the at least one processor, causing the at least one processor to perform the above method. A third aspect provides a computer-readable media storing computer-executable instructions configured to perform the method.

[0008] The methods described herein may be performed by software in machine readable form on a tangible storage medium e.g. in the form of a computer program comprising computer program code means adapted to perform all the steps of any of the methods described herein when the program is run on a computer and where the computer program may be embodied on a computer-readable medium. Examples of tangible (or non-transitory) storage media include disks, thumb drives, memory cards etc and do not include propagated signals. The software can be suitable for execution on a parallel processor or a serial processor such that the method steps may be carried out in any suitable order, or simultaneously.

[0009] This acknowledges that firmware and software can be valuable, separately tradable commodities. It is intended to encompass software, which runs on or controls “dumb” or standard hardware, to carry out the desired functions. It is also intended to encompass software which “describes” or defines the configuration of hardware, such as HDL (hardware description language) software, as is used for designing silicon chips, or for configuring universal programmable chips, to carry out desired functions.

[0010] The preferred features may be combined as appropriate, as would be apparent to a skilled person, and may be combined with any of the aspects of the invention.Brief Description of the Drawings

[0011] Embodiments of the invention will be described, by way of example, with reference to the following drawings, in which:

[0012] Figure 1 is a schematic diagram of an environment where a user device being held in a user’s hand is taking an image or a video;

[0013] Figure 2 is a schematic diagram of an environment where a user device being mounted on a stand is taking an image or a video of content displayed on a screen;

[0014] Figure 3 is a block diagram of an exemplary set of components of a user device in which the embodiments of the present invention may be implemented;

[0015] Figure 4 is a flow diagram of a method for verifying image capture according to some embodiments of the invention;

[0016] Figure 5 is a flow diagram of a method for verifying image capture according to some embodiments of the invention.

[0017] Common reference numerals are used throughout the figures to indicate similar features.Detailed Description

[0018] Embodiments of the present invention are described below by way of example only. These examples represent the best ways of putting the invention into practice that are currently known to the Applicant although they are not the only ways in which this could be achieved. The description sets forth the functions of the example and the sequence of steps for constructing and operating the example. However, the same or equivalent functions and sequences may be accomplished by different examples.

[0019] A common spoofing technique is the Picture-of-Picture (PoP) attack. In this attack, an adversary first obtains or creates an image or a video containing visual features, possibly edits it, and displays it on a screen. Such an image or video may represent a scene expected by a viewer or application program, such as an authenticator program. Then, using a capture device trusted by the viewer or application program, the adversary can take an image or a video of the screen displaying the pre-recorded or pre-created image or video of the scene containing the visual features and upload it for viewing by the viewer or for processing by the application program. Without any effective detection mechanism in place, even with the use of a trusted capture device, the viewer or application program will not be able to differentiate whether the scene is directly captured by the trusted device (meaning that authentic content is provided) or the scene is created or pre-recorded by a device untrusted by the viewer or application program, possibly with some modification, and then being displayed to the trusted device (meaning that the viewer or application program is spoofed with unauthentic content).

[0020] To make the PoP attack more difficult to detect, the adversary may keep the capture device stable at a particular position, for example using a stationary stand, a handheld gimbal or a stable drone that can hover at a one spot, so as to align the field of view of the capture device to the screen to capture an image or video of the screen. This arrangement facilitates the PoP attack by allowing the capture device to be stationary at a position and an orientation carefully chosen (for example, the capture device may be facing a direction approximately perpendicular to the screen) to minimise reflections from the screen and any distortion and to maximise the screen capture area, such that the captured image or video does not show any border of the screen. Otherwise, if the capture device is being held in the adversary’s hand, any slight movement of the hand may cause reflections from the screen and / or a border of the screen to be captured in an image or video, which will make the PoP attack easily transpire, as a viewer can more easily recognise the presence of a screen used for the PoP attack from the reflections or the border of the screen.

[0021] Without any effective detection mechanism in place, the viewer or application program may not be able to differentiate the capture of a real-world scene by a capture device beingheld in a user’s hand from a capture of a virtual scene displayed on a 2-dimensional screen by a capture device fixed at a particular position. Accordingly, there is a need for a method and a system for assessing how a capture device takes images and videos, for example, for determining whether the capture device is being kept stable at a particular position when taking an image or video, and whether the image or video is taken directly from a real-world scene or being pre-recorded and then replayed on a screen to deceive the viewer or application program.

[0022] In some of the embodiments I examples, there is provided a computer-implemented method comprising: detecting a movement of an image sensing device in a time period immediately before, during and / or immediately after a capture of at least one image or video by the image sensing device, detecting a spoofing attack based on analysis of the detection of any movement of the image sensing device. There is also provided an apparatus, comprising at least one processor and at least one memory, the memory storing computer-implementable instructions, when executed by the at least one processor, causing the at least one processor to perform the above method. There is also provided a computer-readable media storing computer-executable instructions configured to perform the method.

[0023] Figure 1 illustrates an environment where a user is taking an image or a video of a scene. As shown in Figure 1 , a system 100 comprises a user device 102, a communication network 104 and server device 106.

[0024] In the environment illustrated by Figure 1 , the user device 102 comprises an image sensing device configured to capture visual information of a scene 108. In various embodiments, the scene 108 is a three-dimensional, real-world environment. Although the scene 108 in Figure 1 is depicted as a car, it can be appreciated that the car is only a nonlimiting example and that the scene can be any three-dimensional, real-world environment and can include any three-dimensional, real-world object(s). In one embodiment, the user device 102 comprises at least one sensor, e.g. a camera, for capturing images and / or videos. The user device 102 may also comprise at least one processor for processing the data relating to the captured images or videos and at least one memory for storing raw data and / or processed data relating to the captured images or videos. The user device 102 may also comprise a communication interface for sending data to and / or receiving data from the authenticator device 106 through the communication network 104. Optionally, the user device can also comprise a display screen for displaying the scene 108 being captured.

[0025] The communication network 104 may include any wired or wireless connection, the internet, or any other form of communication. Although one network 104 is shown in FIG. 1 , the communication network 104 may include any number of different communication networks between the user device 102 and the server device 106. The communication network 104 isconfigured to enable communication between the user device 102 and the server device 106. Various implementations of communication network 120 may employ different types of networks, for example, but not limited to, computer networks, telecommunications networks (e.g., cellular), mobile wireless data networks, and any combination of these and / or other networks.

[0026] The server device 106 is a computing device for displaying or processing the captured images or videos received by the server device 106 from the user device 102. The server device 106 comprises at least one processor and at least one memory for storing instructions and / or data to be processed by the at least one processor. The server device 106 may be configured to display the images or videos captured by the user device 102. Alternatively or additionally, the server device 106 can be configured to analyze the images or videos captured by the user device 102. The server device 106 may also be configured to display an outcome of the analysis of the images or videos captured by the user device 102. In some examples, the server device 106 may have access to information for authenticating a user, a device, an application and / or a process and / or to digital content. For example, the information may include pre-stored information, such as identity information and / or account information of authorised users. The server device 106 may also comprise a communication interface for sending data to and / or receiving data from the user device 102 through the communication network 104.

[0027] Figure 2 illustrates a scenario where a user device 202 is taking an image or a video of a scene being displayed on a screen. As shown in Figure 2, a system 200 comprises a user device 202 being mounted on a stationary stand 210, a communication network 204 and a server device 206. The user device 202, the communication network 204 and the authenticator device 206 may be identical to, or perform functions similar to those performed by, the user device 102, the communication network 104 and the authenticator device 106 respectively.

[0028] The system 200 also comprises a screen 208. The screen 208 is configured to display an image or a video. The image or video may represent a scene. Such image or video may be synthesized or pre-recorded, possibly with some modifications, and then displayed on the screen 208 by a malicious user, in an attempt to deceive a user or an application of the server device. For example, the screen 208 may display an image or a video of a secondhand car being put up for sale online by a user. The user has a strong motive to tamper with the image or video to make it look nicer than it is in real life. Thus, instead of directly taking an image or video of the car using the user device 202 trusted by the server device 206, the user can modify or synthesize an image or video of the car, display the modified I synthesized image or video on the screen 208 and then take an image or video of the modified image or video on the screen 208 using the user device 202. A malicious user may also misrepresents the car for sale by taking an image or video of a car that is not the actual car for sale and display such an imageor video on the screen 208. The stationary stand 210 makes the PoP attack more difficult to detect, as it minimizes movement of the user device 202 during a capture, making any reflections from the screen 208 or the border of the screen 208 less likely to appear in a captured image or video. The stationary stand 210 may be an apparatus in any form suitable for allowing the user device to be fixed in a stationary position. Even if the user device 202 is a device trusted by a user or by an application of the server device 206, without any effective detection mechanism in place, the user or application of the server device 206 will not be able to determine whether the user device 202 is being mounted on a stationary stand when taking an image or video and to determine whether the scene shown in the image I video are directly captured by the user device 202 or are pre-recorded and then displayed to the user device 202. The stationary stand is just one exemplary way of keeping the capture device stable at a particular position. There are alternative ways to achieve a similar effect, such as using a handheld gimbal or a stable drone that can hover at a one spot, to align the field of view of the capture device to a display screen and capture an image or video of the screen.

[0029] Figure 3 is a block diagram of an exemplary set of components for a system 300 in which the embodiments of the present invention may be implemented. The user device 102 I 202 in system 100 1200 may be implemented as system 300.

[0030] The system 300 may be implemented as one or more computing and / or electronic devices. The system 300 comprises one or more processors 302 which may be microprocessors, controllers or any other suitable type of processors for processing computer executable instructions to control the operation of the system 300. Platform software comprising an operating system 306 or any other suitable platform software may be provided on the system to enable application software 308 to be executed on the system. In some embodiments, the application software 308 may comprise a software program for processing images, deriving data from the images, and processing the data derived from the images according to various methods described herein. The components of the system 300 described herein may be enclosed in a casing.

[0031] Computer executable instructions may be provided using any computer-readable media that are accessible by the system 300. Computer-readable media may include, for example, computer storage media such as a memory 304 and communications media. Computer storage media, such as the memory 304, include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storagedevices, or any other non-transmission medium that can be used to store information for access by a computing device. In contrast, communication media may embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transport mechanism. Although the computer storage medium (the memory 304) is shown within the system 300, it will be appreciated, by a person skilled in the art, that at least a part of the storage may be distributed or located remotely and accessed via a network or other communication link (e.g. using a communication interface 312).

[0032] The system 300 may comprise an input / output controller 314 arranged to receive and process input from one or more input devices 318 which may be separate from or integral to the system 300 and may also be arranged to output information to one or more optional output devices 316 which may be separate from or integral to the system 300. In some embodiments, the input devices 318 may comprise input devices for controlling the operation of the system 300, such as a set of buttons or keys. For example, the input devices 318 may comprise keys for controlling at least one sensing device, such as adjusting an orientation and / or a zoom of a camera, and / or for manipulating an image being displayed on a screen. In some embodiments, the input devices may comprise at least one image sensing device, such as a camera.

[0033] The input devices 318 may further comprise at least one movement detection sensor 320 for detecting physical movements of the system 300, which may be implemented in the form of the user device 102 or 202 in system 100 or 200. In some embodiments, the at least one movement detection sensor 320 includes at least one gyroscope, at least one accelerometer, at least one magnetometer and / or at least one pressure sensor. Gyroscopes are known for measuring changes in orientation and angular velocity. Accelerometers are known for measuring directions and magnitudes of acceleration. Magnetometers are devices that measure the direction, strength, or relative change of a magnetic field at a particular location. They can provide absolute angular measurements relative to the Earth's magnetic field. Pressure sensors, such as capacitive pressure sensors and piezoresistive strain gauge pressure sensors, can be used to detect a force, a tension and / or a movement applied to an object as a result of a pressure applied to the object. These sensors are all known for their use in detecting movements of objects on which they are mounted.

[0034] The at least one movement detection sensor 320 exemplified above are for detecting changes in the position and I or acceleration of the system 300 so as to detect movement of the system. In another example, such at least one movement detection sensor 320 may be at least one depth sensor.

[0035] The at least one depth sensor can provide depth information relating to a surface or a scene from a viewpoint. A depth sensor measures the distance, or view “depth”, between thedepth sensor and the object or surface. For example, depth sensors can measure depth by using stereo image sensing technologies or a round trip time of reflected depth sensing signals. Any change in the depth indicates a relative movement between the depth sensor comprised in the system 300 and the object or surface which reflects depth sensing signals back to the depth sensor. This type of depth sensor can be used to measure relative movements between the image sensing device 102 1 202 and the object 108 1 208 that the image sensing device is capturing.

[0036] In some embodiments, the at least one depth sensor comprises at least one of a radio wave based depth sensor, such as a radar sensor, a light based depth sensor, such as a LiDAR sensor, an acoustic depth sensor, such as a sonar sensor, and a multi-perspective camera setup system.

[0037] LiDAR, also known as "light detection and ranging" or "laser imaging, detection, and ranging", is a time-of-flight technique for determining ranges (variable distance) by targeting an object or a surface with a laser and measuring the time for the reflected light to return to a sensor, e.g. a LiDAR camera.

[0038] Radar, which stands for radio detection and ranging, is a detection technique that uses radio waves to determine the distance (ranging), angle, and radial velocity of objects relative to a site. A radar system consists of a transmitter producing electromagnetic waves in the radio or microwaves domain, a transmitting antenna, a receiving antenna (often the same antenna is used for transmitting and receiving) and a receiver and processor to determine properties of an object. Radio waves from the transmitter reflect off the object and return to the receiver, giving information about the object's location and speed. Radar signals can obtain a distance measurement based on the time-of-flight by transmitting a short pulse of radio signal and measure the time it takes for the reflection of the radio signal caused by a target object or surface to return.

[0039] Sonar, also known as sound navigation and ranging, is an acoustic location technique that uses sound propagation and reflection to measure distances to sound reflective surfaces or detect objects. Sonar can be used to derive a contour of a surface by emitting pulses of sounds and detecting echoes reflected from the surface.

[0040] The time-of-flight sensor is a range imaging camera system employing time-of-flight techniques to calculate a distance between its camera and a point on a target surface, by measuring the round-trip time of a signal, such as an artificial light signal emitted by a light source, and reflected by the point on the target surface back to the camera. The time-of-flight sensor can be used to make a digital 3-D representation of a surface by collecting depthinformation from a plurality of points on the surface using the time of flight of the light (i.e., the time it takes each light signal from the light source to hit a target point and return to the camera).

[0041] A depth sensor based on the multi-perspective camera setup system measures depth by capturing images of the same scene from different perspectives using different cameras. It uses stereophotogrammetry where the depth data of the pixels are determined from data acquired using the multi-perspective camera setup system. To solve the depth measurement problem using the multi-perspective camera setup system, corresponding points in the different images captured by the different cameras are identified. A disparity map can be constructed to indicate the apparent pixel difference or motion between a pair of stereo images. Various techniques exist for deriving a depth map from a disparity map.

[0042] The optional output devices 316 may include a display screen. In some embodiments, the output device 316 may also act as an input device, for example, when the output device 316 is a touch screen. The input / output controller 314 may also output data to devices other than the output device, for example to a locally connected computing device. According to some embodiments, image processing and calculations based on data derived from images captured by the input device 318 and / or any other functionality as described in the embodiments, may be implemented by software or firmware, for example, the operating system 306 and the application software 308 working together and / or independently, and executed by the processor 202.

[0043] The communication interface 312 enables the system 300 to communicate with other devices and systems. The communication interface 312 may include any type of signal transceivers, such as 3G, 4G and / or 5G wireless mobile telecommunications transceivers, WiFi™ signal transceivers and / or Bluetooth™ transceivers, as well as any wired telecommunications transceivers, such as the Ethernet and the thunderbolt interface.

[0044] The functionality described herein in the embodiments may be performed, at least in part, by one or more hardware logic components. According to an embodiment, the computing device 300 is configured by programs 306, 308 stored in the memory 304 when executed by the processor 302 to execute the embodiments of the operations and functionality described. Alternatively, or in addition, the functionality described herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), Graphics Processing Units (GPUs).

[0045] Figure 4 is a flow diagram of a method 400 according to some embodiments of the invention. The method aims to verify authentic capturing of a real-world scene by an image capture device. The method prevents spoofing attacks and, in particular, PoP attacks.

[0046] Method 400 starts with step 402, in which a movement of an image sensing device in a time period immediately before, during and / or immediately after a capture of at least one image and / or video by the image sensing device is detected. The movement can be detected by any movement detection sensor, such as by at least one of a gyroscope, an accelerometer, a magnetometer, a pressure sensor and a depth sensor, which is comprised within or attached to the image sensing device. The time period immediately before, during and / or immediately after a capture of at least one image can be a pre-defined time period of any suitable length. For example, the time period may have a length of 0.5, 1 , 1 .5 or 2 seconds. The time period being immediately before, during and / or immediately after a capture of at least one image means that the image capture can take place immediately before or after this time period or at any time within this time period.

[0047] Then the method 400 proceeds to step 404, which involves detecting a spoofing attack based on analyzing the detection of any movement of the image sensing device. In some embodiment, this step comprises detecting a spoofing attack based on whether the movements of the image sensing device within the time period immediately before, during and / or immediately after a capture of at least one image is detected to be within a pre-defined movement threshold.

[0048] In some embodiments, the pre-defined movement threshold is a threshold value for a maximum rotation, maximum acceleration or maximum displacement of the image sensing device within the time period. Alternatively, the pre-defined movement threshold can be a threshold value for an average rotation, average acceleration or average displacement of the image sensing device within the time period, or a threshold value for a standard deviation of the rotation, acceleration or displacement of the image sensing device within the time period. If the maximum or average rotation, acceleration or displacement of the image sensing device within the time period does not exceed the threshold value, then it can be determined that a spoofing attack is likely to be present. On the other hand, if the maximum or average rotation, acceleration or displacement of the image sensing device within the time period exceeds the threshold value, then it can be determined that a spoofing attack is not likely to be present. This is because an adversary, who is trying to take an image of a virtual scene displayed on a display screen and try to avoid the image showing any indication of being taken from a display screen, is likely to fix the camera at a position and carefully align the field of view of the image sensing device to the screen, so that reflections from the screen is minimised in the image and that the captured image or video does not show any border of the screen. If the image sensingdevice is held by a user’s hand(s) without being mounted to any other apparatus for keeping the image sensing device constantly stable at a particular position when capturing the image, normally there will be natural movements due to the user’s hands not being able to be completely stationary at a fixed position in the air within a period of time. On the other hand, if the image sensing device is being mounted to another apparatus for keeping the image sensing device constantly stable at a particular position when capturing an image, the image sensing device will have very little or virtually no movements within the period of time immediately before and / or after the image capture.

[0049] The threshold value for the maximum or average displacement of the image sensing device can be an empirical value obtained by measuring displacements of an image sensing device, such as a camera or a mobile phone comprising a camera, being held by various people taking a picture. For example, if it is determined (e.g. by measuring displacements of various image sensing devices held by a number of people while taking an image) that human users will almost certainly cause the image sensing device to displace by at least 2 mm while taking the image, then the empirical value, and hence the threshold value, can be set to 2 mm. Then in use, if the maximum displacement of an image sensing device during an image or video capture is within the pre-determined threshold value of 2 mm, the image sensing device is likely to be fixed to a stationary position, e.g. mounted on a stand and it can be determined that a spoofing attack is likely to be present for the reason mentioned in the preceding paragraph.

[0050] Optionally, the outcome of step 404 is used for determining whether any subsequent process is to be activated. If it is determined in step 404 that a spoofing attack is likely to be present, a remedial process can be activated. In one example, such a remedial process is to restrict the user’s access or further access to a particular function, application or process. This is to prevent the adversary making the spoofing attack from accessing such function, application or process for a malicious purpose. In another example, such a remedial process may be to send an alert or notification to an administrator, so that the administrator can take further actions for verifying whether there is a spoofing attack. Such further actions may include, for example, reviewing the content of the image or any video being recorded by the image sensing device immediately before, during and / or immediately after the capture of the image to see if there is any further indication of the image being taken from a screen display. Such further indication may be reflections from the screen and / or any border of the screen being shown in the image and / or video.

[0051] Figure 5 is a flow diagram of a method 500 according to further embodiments of the invention. The method aims to verify authentic capturing of a real-world scene by an image capture device. The method prevents spoofing attacks and, in particular, PoP attacks.

[0052] Step 502 in Figure 5 is identical to step 402 in Figure 4. As with method 400, method 500 starts with a step (step 502), in which a movement of an image sensing device in a first time period immediately before, during and / or immediately after a capture of at least one image is detected by the image sensing device. As with step 402, in step 502 the movement can be detected by any movement detection sensor, such as by at least one of a gyroscope, an accelerometer, a magnetometer, a pressure sensor and a depth sensor, which is comprised within or attached to the image sensing device. The first time period immediately before, during and / or immediately after a capture of at least one image can be a pre-defined time period of any suitable length. For example, the first time period may have a length of 0.5, 1 , 1.5 or 2 seconds. The first time period being immediately before, during and / or immediately after a capture of at least one image means that the image capture can take place immediately before or after the first time period or at any time within the first time period.

[0053] Method 500 further comprises a step 504. Step 504 can be performed before, after or simultaneously with step 502. Step 504 comprises receiving a video captured in a second time period immediately before, during and / or immediately afterthe capture of the at least one image by the image sensing device. In some embodiments, the video is captured by the same image sensing device which captured the at least one image. Alternatively, the video can be captured by a different image sensing device. The first time period and the second time period may overlap with each other, although it is not necessary forthem to have any overlap. The capture of the at least one image can take place immediately before or after the second time period or at any time within the second time period. As an optimisation, the video being captured in the second time period has a low resolution, so that it does not necessarily takes up much storage space.

[0054] Subsequent to steps 502 and 504, the method 500 proceeds to step 506, in which a detection of spoofing attack is performed based on analyzing the detection of any movement of the image sensing device in step 502 and the video received in step 504. In some embodiments, step 506 comprises detecting a spoofing attack based on whether the movements of the image sensing device within the first time period detected in step 502 are to be within a pre-defined movement threshold. This may comprise detecting a spoofing attack based on whether the movements of the image sensing device within the first time period in step 502 are detected to be within a pre-defined movement threshold. In some examples, the pre-defined movement threshold is a threshold value for a maximum rotation, maximum acceleration or maximum displacement of the image sensing device within the first time period. Alternatively, the pre-defined movement threshold can be a threshold value for an average rotation, average acceleration or average displacement of the image sensing device within the first time period. If the maximum or average rotation, acceleration or displacement of the image sensing device within the first time period does not exceed the threshold value, which indicatesthat the image sensing device is likely to be mounted on another apparatus for keeping the sensing device stable at a particular position, then it can be determined that a spoofing attack is very likely to be present. On the other hand, if the maximum or average rotation, acceleration or maximum displacement of the image sensing device within the first time period exceeds the threshold value, which indicates that the image sensing device is likely to be held in a user’s hand(s), then it can be determined that a spoofing attack is not likely to be present.

[0055] Additionally, step 506 also comprises detecting a spoofing attack based on the video received in step 504. In some embodiments, algorithms can be used to detect a border of a screen in the video so as to spot the presence of a display screen in the video, which can indicate the likelihood of a spoofing attack. This can be done by detecting contrast between the display area of a screen and the bezel around the display area of the screen, as the display area is typically a lot brighter than the bezel around it. Alternatively or additionally, as borders of a screen usually consist straight lines and right angles, algorithms can be used to detect straight line edges and / or right-angle corners for identifying borders of a screen. If such borders of a screen are detected, it can be determined that there is a high likelihood for the video to be taken from a scene displayed on a screen, instead of a real-world scene, and hence there is a high likelihood of spoofing attack. Three-dimensional objects with pronounced edge features or blob features can be effectively recognized by detection algorithms, such as the Harris affine region detection and the scale-invariant feature transform (SIFT) method. There are also various open-source software programs for detecting three-dimensional objects, an example of which can be found at: https: / / docs.opencv.Org / 4.x / d5 / d54 / aroup objdetect.html.

[0056] Optionally, in step 506 the detection of spoofing attack comprises estimating a risk of a spoofing attack based on the detection of the movement of the image sensing device in step 502 and the video received in step 504. For example, if both of the detection of the movement of the image sensing device in step 502 and the video received in step 504 indicate a high likelihood of a spoofing attack, it can be estimated that the risk of a spoofing attack is high; if only one of the detection of the movement of the image sensing device in step 502 and the video received in step 504 indicates a high likelihood of a spoofing attack, it can be estimated that the risk of a spoofing attack is medium; and if neither of the detection of the movement of the image sensing device in step 502 and the video received in step 504 indicates a high likelihood of a spoofing attack, it can be estimated that the risk of a spoofing attack is low.

[0057] Optionally, the outcome of step 506 is used for determining whether any subsequent process is to be performed. If it is determined in step 506 that a spoofing attack is likely to be present, e.g. a high and / or medium risk of spoofing attack, a remedial process can be performed. In one example, such a remedial process can be to restrict the user’s access or further access to a particular function, application or process. This is to prevent the adversarymaking the spoofing attack from accessing such function, application or process for a malicious purpose. In another example, such a remedial process may be to send an alert or notification to an administrator, so that the administrator can take further actions for verifying whether there is a spoofing attack. Such further actions may include, for example, reviewing the content of the image or any video being recorded by the image sensing device immediately before, during and / or immediately after the capture of the image to see if there is any further indication of the image being taken from a screen display. Such further indication may be reflections, in particular aberrant reflections, from the screen and / or any border of the screen being shown in the image and / or video. Various methods exist for detecting reflections in images. An exemplary method for detecting reflection is published in paper “Reflection Detection in Image Sequences”, by Mohamed Abdelaziz Ahmed, Francois Pitie and Anil Kokaram from Sigmedia, Electronic and Electrical Engineering Department, Trinity College Dublin, published in July 2011 , Proceedings I CVPR, IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE Computer Society Conference on Computer Vision and Pattern Recognition. This paper presents an automated technique for detecting reflections in image sequences. It is based on analyzing spatio-temporal profiles of feature point trajectories and focuses on examining three main features of reflections: 1) the ability of decomposing an image into two-independent layers, 2) image sharpness, 3) the temporal behavior of image patches.

[0058] In an embodiment alternative to Fig.5, instead of receiving a video automatically captured in a second time period immediately before, during and / or immediately after the capture of the at least one image by the image sensing device as set out in step 504, in the alternative embodiment, instructions can be given to a user of the image sensing device to move the image sensing device in a certain way when a video is captured. Such instructions may be in the form of an on-screen prompt, or any other visual or audio instructions suitable for directing the user to move the image sensing device. For example, there could be arrows on the screen instructing the user to rotate the image sensing device in a clockwise I anticlockwise direction and / or move the image sensing device up, down, to the left or to the right. The purpose of instructing the user to move the camera while the video is captured is to increase the chance that the video captured can record and reveal any border or corner of a display screen which may be used in a PoP attack. Such a video can then be used in step 506 as described above. Optionally, the image sensing device can verify, using the movement detection sensors, that the user is indeed moving the image sensing device as instructed and simultaneously capture a video during the movement. If the verification concludes that the movement of the image sensing device does not conform to the instructions given to the user, the video can be disregarded, and new instructions can be given to the user to move the image sensing device again and a video can be recorded again during the new movement of the sensing device. Such steps may repeat until the movement of the image sensing device conforms to the instructions given to the user.

[0059] The term 'computer' or ‘computing device’ is used herein to refer to any device with processing capability such that it can execute instructions. Those skilled in the art will realize that such processing capabilities are incorporated into many different devices and therefore the term 'computer' includes PCs, servers, mobile telephones, personal digital assistants and many other devices.

[0060] Those skilled in the art will realize that storage devices utilized to store program instructions can be distributed across a network. For example, a remote computer may store an example of the process described as software. A local or terminal computer may access the remote computer and download a part or all of the software to run the program. Alternatively, the local computer may download pieces of the software as needed, or execute some software instructions at the local terminal and some at the remote computer (or computer network). Those skilled in the art will also realize that by utilizing conventional techniques known to those skilled in the art that all, or a portion of the software instructions may be carried out by a dedicated circuit, such as a DSP, programmable logic array, or the like.

[0061] Any range or device value given herein may be extended or altered without losing the effect sought, as will be apparent to the skilled person.

[0062] It will be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments. The embodiments are not limited to those that solve any or all of the stated problems or those that have any or all of the stated benefits and advantages.

[0063] Any reference to 'an' item refers to one or more of those items. The term 'comprising' is used herein to mean including the method blocks or elements identified, but that such blocks or elements do not comprise an exclusive list and a method or apparatus may contain additional blocks or elements.

[0064] The steps of the methods described herein may be carried out in any suitable order, or simultaneously where appropriate. Additionally, individual blocks may be deleted from any of the methods without departing from the spirit and scope of the subject matter described herein. Aspects of any of the examples described above may be combined with aspects of any of the other examples described to form further examples without losing the effect sought.

[0065] It will be understood that the above description of a preferred embodiment is given by way of example only and that various modifications may be made by those skilled in the art. Although various embodiments have been described above with a certain degree of particularity, or with reference to one or more individual embodiments, those skilled in the art could makenumerous alterations to the disclosed embodiments without departing from the spirit or scope of this invention.

Claims

Claims1 . A computer-implemented method comprising: detecting a movement of an image sensing device in a time period immediately before, during and / or immediately after a capture of at least one image or video by the image sensing device, detecting a spoofing attack based on analysis of the detection of any movement of the image sensing device.

2. The method of claim 1 , wherein said detecting a spoofing attack comprises determining the presence of a spoofing attack if it is detected that the movement of the image sensing device in the time period is within a predefined movement threshold.

3. The method of claim 1 or 2, wherein said time period is a first time period, wherein the method further comprises receiving a video captured in a second time period, and wherein said detecting a spoofing attack further comprises detecting a spoofing attack based on the video.

4. The method of claim 3, wherein the second time period is immediately before, during and / or immediately after the capture of the at least one image or video by the image sensing device.

5. The method of claim 3, further comprising providing an instruction for directing a user of the image sensing device to move the image sensing device while the video is being taken.

6. The method of claim 5, further comprising detecting movement of the image sensing device while the video is being taken, and disregarding the video if the detected movement does not correspond to the instruction.

7. The method of any of claims 3-6, wherein said detecting spoofing determines the presence of a spoofing attack, if it is detected that the movement of the image sensing device in the first time period is outside a predefined movement threshold, and / or if it is determined that the captured video includes an edge of a flat surface or it is detected that the captured video includes aberrant reflections.

8. The method of claim 2 or 7, wherein said predefined movement threshold is a threshold for a maximum rotation, a maximum acceleration or a maximum displacement of the image sensing device, a threshold for an average rotation, average acceleration or average displacement of the image sensing device within the time period, or a threshold for changes in the rotation, acceleration or displacement of the image sensing device within the time period.

9. The method of any of claims 1-4, wherein the movement is detected based on data received from at least one movement detection sensor.

10. The method of claim 9, wherein the at least one movement detection sensor comprises at least one of a gyroscope, an accelerometer, a magnetometer, a pressure sensor and a depth sensor.11 . The method of claim 10, wherein the depth sensor comprises at least one of a time- of-flight sensor, a LiDAR sensor, a radar sensor, a sonar sensor and a multi-perspective camera setup depth sensor.

12. An apparatus, comprising at least one processor and at least one memory, the memory storing computer- implementable instructions, when executed by the at least one processor, causing the at least one processor to perform the method of any of claims 1-11.

13. A computer-readable media storing computer-executable instructions configured to perform the method of any of claims 1-11.

14. A method substantially as described with reference to figures 1-5 of the drawings.

15. A system substantially as described with reference to figures 1-5 of the drawings.