Method and system for validation

The method uses device movement analysis to detect spoofing attacks by comparing stable and unstable image capture devices, enhancing authentication systems' resistance to picture-of-picture attacks.

JP2026507430APending Publication Date: 2026-03-04OPENORIGINS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing authentication systems are vulnerable to spoofing attacks where visual information is created or reproduced to deceive viewers, especially in picture-of-picture attacks where a capture device is stabilized to minimize screen reflections and distortions, making it difficult to distinguish between real-world and virtual scenes.

Method used

A computer-implemented method that detects movement of an image sensing device before, during, and after image capture, using sensors like gyroscopes and accelerometers to analyze movement patterns and detect spoofing attacks by determining if the device is stationary or held by a user.

Benefits of technology

Effectively identifies spoofing attacks by distinguishing between stable and unstable capture devices, preventing unauthorized image or video content from being passed off as authentic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507430000001_ABST
    Figure 2026507430000001_ABST
Patent Text Reader

Abstract

The computer-implemented method includes detecting movement of the image sensing device during a time period immediately before, during, and / or immediately after capture of at least one image or video by the image sensing device, and detecting a spoof attack based on an analysis of the detection of any movement of the image sensing device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION Embodiments of the present invention relate generally to methods, apparatus, and systems for verification, and more particularly, for verification of image capture for the purpose of preventing spoofing attacks. [Background technology]

[0002] Digital media containing visual information, such as images and videos, have wide application in the information age. Authentication systems typically assume that the content of an image or video is authentic, especially if the image or video is captured by a device or application that the viewer trusts.

[0003] However, the capture of visual information is subject to a variety of "spoofing" attacks, in which visual information can be created, modified, or reproduced by an unauthorized adversary or program to gain an unfair or even illegal advantage. These attacks attempt to present a false scene to the viewer. A false scene can be created by artificially setting up a scene that mimics the scene expected by the viewer or authenticator. There is a continuing need for methods to verify whether an image is being captured in a way that deceives the viewer.

[0004] The embodiments described below are not limited to implementations that address any or all of the shortcomings of known techniques for verification or spoof detection. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] https: / / docs.opencv.org / 4.x / d5 / d54 / group__objdetect.html [Non-patent document 2] "Reflection Detection in Image Sequences" by Mohamed Abdelaziz Ahmed, Francois Pitie, and Anil Kokaram of Sigmedia, Electronic and Electrical Engineering Department, Trinity College Dublin, July 2011, Proceedings / CVPR, IEEE Computer Society Conference on Computer Vision and Pattern Recognition Summary of the Invention [Means for solving the problem]

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0007] The invention is set out in the accompanying set of claims.

[0008] A first aspect provides a computer-implemented method including detecting movement of an image sensing device during a time period immediately prior to, during, and / or immediately after capture of at least one image or video by the image sensing device, and detecting a spoof attack based on an analysis of the detection of any movement of the image sensing device. A second aspect provides an apparatus comprising at least one processor and at least one memory, the memory storing computer-executable instructions that, when executed by the at least one processor, cause the at least one processor to implement the above method. A third aspect provides a computer-readable medium storing computer-executable instructions configured to implement the method.

[0009] The methods described herein may be implemented by software in machine-readable form on a tangible storage medium, for example, in the form of a computer program comprising computer program code means adapted to perform all the steps of any of the methods described herein when the program is run on a computer and the computer program may be embodied on a computer-readable medium. Examples of tangible (or non-transitory) storage media include disks, thumb drives, memory cards, etc., and do not include propagated signals. The software may be suitable for execution on a parallel or serial processor such that the method steps may be performed in any suitable order, or simultaneously.

[0010] This recognizes that firmware and software can be valuable, separately tradable commodities. It is intended to encompass software that runs on or controls "dumb" or standard hardware to perform a desired function. It is also intended to encompass software that "describes" or defines the configuration of hardware, such as HDL (Hardware Description Language) software such as that used to design silicon chips or configure universal programmable chips, to perform a desired function.

[0011] As will be apparent to one skilled in the art, the preferred features may be combined as desired and may be combined with any of the aspects of the present invention.

[0012] Embodiments of the present invention will now be described, by way of example, with reference to the following drawings, in which: [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a schematic diagram of an environment in which a user device held in a user's hand is taking an image or video. [Figure 2]FIG. 1 is a schematic diagram of an environment in which a user device mounted on a stand is taking images or footage of content displayed on a screen. [Figure 3] FIG. 2 is a block diagram of an exemplary set of components of a user device in which embodiments of the present invention may be implemented. [Figure 4] 1 is a flow diagram of a method for verifying image capture according to some embodiments of the present invention. [Figure 5] 1 is a flow diagram of a method for verifying image capture according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0014] Common reference numbers are used throughout the figures to denote similar features.

[0015] Embodiments of the present invention are described below by way of example only. These examples represent the best ways of practicing the invention currently known to the applicant, but are not the only ways in which this may be achieved. The specification describes the functions of the examples and the sequence of steps for constructing and operating the examples. However, the same or equivalent functions and sequences may be achieved by different examples.

[0016] A common spoofing technique is the picture-of-picture (PoP) attack. In this attack, an adversary first acquires or creates an image or video containing visual features, possibly editing it, and displays it on a screen. Such an image or video may represent a scene expected by an application program, such as a viewer or authenticator program. Then, using a capture device trusted by the viewer or application program, the adversary can take an image or video of a screen displaying a pre-recorded or pre-created image or video of the scene containing the visual features and upload it for viewing by the viewer or processing by the application program. Without some effective detection mechanism in place, even with a trusted capture device, the viewer or application program cannot distinguish whether the scene was captured directly by a trusted device (meaning authentic content is provided) or whether the scene was created or pre-recorded by a device not trusted by the viewer or application program, possibly with some modification, and then displayed on the trusted device (meaning the viewer or application program is spoofed with inauthentic content).

[0017] To make PoP attacks more difficult to detect, an adversary can stabilize a capture device in a specific position using, for example, a stationary stand, a handheld gimbal, or a stabilized drone that can hover in one spot to focus the capture device's field of view on the screen and capture an image or video of the screen. This configuration facilitates PoP attacks by allowing the capture device to remain stationary in a carefully chosen position and orientation (e.g., the capture device may be oriented nearly perpendicular to the screen) to minimize reflections from the screen and any distortions and maximize the screen capture area so that the captured image or video does not show any edges of the screen. Otherwise, if the capture device is held in an adversary's hand, any slight movement of the hand could cause reflections from the screen and / or the edges of the screen to be captured in the image or video, making the PoP attack more easily noticeable because a viewer can more easily recognize the presence of a screen used in a PoP attack from the screen's reflection or edges.

[0018] Unless some effective detection mechanism is implemented, a viewer or application program may not be able to distinguish between the capture of a real-world scene by a capture device held in a user's hand and the capture of a virtual scene displayed on a two-dimensional screen by a capture device fixed in a particular position. Thus, there is a need for methods and systems for evaluating how a capture device captures images and videos, for example, to determine whether the capture device is stabilized in a particular position when taking the image or video, and whether the image or video is taken directly from a real-world scene or is pre-recorded and then played back on a screen to deceive the viewer or application program.

[0019] In some embodiments / examples, a computer-implemented method is provided that includes detecting movement of an image sensing device during a time period immediately before, during, and / or after capture of at least one image or video by the image sensing device, and detecting a spoof attack based on an analysis of the detection of any movement of the image sensing device. Also provided is an apparatus that includes at least one processor and at least one memory, the memory storing computer-executable instructions that, when executed by the at least one processor, cause the at least one processor to implement the method. Also provided is a computer-readable medium that stores computer-executable instructions configured to implement the method.

[0020] 1 illustrates an environment in which a user is taking an image or video of a scene. As shown in FIG. 1, a system 100 includes a user device 102, a communication network 104, and a server device 106.

[0021] In the environment illustrated by FIG. 1 , the user device 102 comprises an image sensing device configured to capture visual information of a scene 108. In various embodiments, the scene 108 is a three-dimensional, real-world environment. While the scene 108 in FIG. 1 is illustrated as a car, it can be appreciated that a car is merely a non-limiting example and that the scene can be any three-dimensional, real-world environment and can include any three-dimensional, real-world object. In one embodiment, the user device 102 comprises at least one sensor, e.g., a camera, for capturing images and / or video. The user device 102 may also comprise at least one processor for processing data associated with the captured image or video and at least one memory for storing raw and / or processed data associated with the captured image or video. The user device 102 may also comprise a communication interface for sending data to and / or receiving data from the server device 106 over the communication network 104. Optionally, the user device 102 may also comprise a display screen for displaying the scene 108 being captured.

[0022] The communication network 104 may include any wired or wireless connection, the Internet, or any other form of communication. While one network 104 is shown in FIG. 1 , the communication network 104 may include any number of different communication networks between the user device 102 and the server device 106. The communication network 104 is configured to enable communication between the user device 102 and the server device 106. Various implementations of the communication network 104 may employ different types of networks, for example, but not limited to, computer networks, telecommunications networks (e.g., cellular), mobile wireless data networks, and any combination of these and / or other networks.

[0023] The server device 106 is a computing device for displaying or processing captured images or videos received by the server device 106 from the user device 102. The server device 106 comprises at least one processor and at least one memory for storing instructions and / or data to be processed by the at least one processor. The server device 106 may be configured to display images or videos captured by the user device 102. Alternatively or additionally, the server device 106 may be configured to analyze images or videos captured by the user device 102. The server device 106 may also be configured to display results of the analysis of images or videos captured by the user device 102. In some examples, the server device 106 can access information for authenticating users, devices, applications, and / or processes and / or digital content. For example, the information may include pre-stored information, such as identity information and / or account information of authorized users. The server device 106 may also comprise a communication interface for sending data to and / or receiving data from the user device 102 over the communication network 104.

[0024] 2 illustrates a scenario in which a user device 202 is taking an image or video of a scene being displayed on a screen. As shown in FIG. 2, system 200 comprises user device 202 mounted on a stationary stand 210, a communication network 204, and a server device 206. User device 202, communication network 204, and server device 206 may be identical to or perform similar functions as those performed by user device 102, communication network 104, and server device 106, respectively.

[0025] System 200 also includes a screen 208. Screen 208 is configured to display an image or video. The image or video may represent a scene. Such an image or video may be synthesized or pre-recorded, possibly with some modification, and then displayed on screen 208 by a malicious user in an attempt to deceive a user or an application on the server device. For example, screen 208 may display an image or video of a used car posted for sale online by a user. The user has a strong incentive to alter the image or video to make it appear better than it actually is. Thus, instead of directly taking an image or video of the car using user device 202, which server device 206 trusts, the user can modify or composite the image or video of the car, display the modified / composite image or video on screen 208, and then use user device 202 to take an image or video of the modified image or video on screen 208. A malicious user can also misrepresent a car for sale by taking an image or video of a car that is not actually for sale and displaying such image or video on screen 208. The stationary stand 210 minimizes movement of the user device 202 during capture, making any reflections from or edges of the screen 208 less likely to appear in the captured image or video, making PoP attacks more difficult to detect. The stationary stand 210 can be any form of apparatus suitable for allowing the user device to be fixed in a stationary position. Even if the user device 202 is a device trusted by the user or an application on the server device 206, unless some effective detection mechanism is implemented, the user or an application on the server device 206 cannot determine whether the user device 202 is attached to the stationary stand when taking an image or video, and whether the scene shown in the image / video was captured directly by the user device 202 or pre-recorded and then displayed on the user device 202.A stationary stand is just one exemplary method of keeping a capture device stable in a particular position. There are alternative ways to achieve a similar effect, such as using a handheld gimbal or a stabilized drone that can hover in one spot to line the field of view of the capture device with the display screen and capture an image or footage of the screen.

[0026] 3 is a block diagram of an exemplary set of components for a system 300 in which embodiments of the present invention may be implemented. A user device 102 / 202 in system 100 / 200 may be implemented as system 300.

[0027] System 300 may be implemented as one or more computing and / or electronic devices. System 300 includes one or more processors 302, which may be a microprocessor, a controller, or any other suitable type of processor for processing computer-executable instructions to control the operation of system 300. Platform software including an operating system 306 or any other suitable platform software may be provided on the system to enable application software 308 to run on the system. In some embodiments, application software 308 may include software programs for processing images, deriving data from images, and processing data derived from images according to various methods described herein. The components of system 300 described herein may be housed in a casing.

[0028] Computer-executable instructions may be provided using any computer-readable media accessible by system 300. Computer-readable media may include, for example, computer storage media, such as memory 304, and communication media. Computer storage media, such as memory 304, include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that may be used to store information for access by a computing device. In contrast, communication media may embodi computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism. While computer storage media (memory 304) is shown in system 300, those skilled in the art will appreciate that at least a portion of the storage may be distributed or remotely located and accessed over a network or other communications link (e.g., using communications interface 312).

[0029] System 300 may include an input / output controller 314 configured to receive and process input from one or more input devices 318, which may be separate from or integral with system 300, and to output information to one or more optional output devices 316, which may be separate from or integral with system 300. In some embodiments, input device 318 may include input devices for controlling the operation of system 300, such as a set of buttons or keys. For example, input device 318 may include keys for controlling at least one sensing device and / or for manipulating an image displayed on a screen, such as adjusting the orientation and / or zoom of a camera. In some embodiments, input device 318 may include at least one image sensing device, such as a camera.

[0030] The input device 318 may further include at least one motion detection sensor 320 for detecting physical movement of the system 300, which may be implemented in the form of a user device 102 or 202 in the system 100 or 200. In some embodiments, the at least one motion detection sensor 320 includes at least one gyroscope, at least one accelerometer, at least one magnetometer, and / or at least one pressure sensor. Gyroscopes are known for measuring changes in orientation and angular velocity. Accelerometers are known for measuring the direction and magnitude of acceleration. Magnetometers are devices that measure the direction, strength, or relative change of a magnetic field at a specific location. Magnetometers can provide absolute angular measurements with respect to the Earth's magnetic field. Pressure sensors, such as capacitive pressure sensors and piezoresistive strain gauge pressure sensors, can be used to detect force, tension, and / or movement applied to an object as a result of pressure applied to the object. All of these sensors are known for use in detecting movement of the object to which the sensor is attached.

[0031] The at least one motion detection sensor 320 exemplified above is for detecting changes in position and / or acceleration of the system 300 to detect system movement. In another example, such at least one motion detection sensor 320 may be at least one depth sensor.

[0032] At least one depth sensor can provide depth information related to a surface or scene from a certain viewpoint. A depth sensor measures the distance, or "depth," between the depth sensor and an object or surface. For example, a depth sensor can measure depth by using stereo image sensing techniques or the round-trip time of a reflected depth sensing signal. Any change in depth indicates relative motion between the depth sensor included in the system 300 and the object or surface that reflects the depth sensing signal back to the depth sensor. This type of depth sensor can be used to measure relative motion between the image sensing device 102 / 202 and the object 108 / 208 that the image sensing device is capturing.

[0033] In some embodiments, the at least one depth sensor includes at least one of a radio-based depth sensor such as a radar sensor, an optical-based depth sensor such as a LiDAR sensor, an acoustic depth sensor such as a sonar sensor, and a multi-view camera setup system.

[0034] LiDAR, also known as "light detection and ranging" or "laser imaging, detection, and ranging," is a time-of-flight technique for determining distance (variable distance) by using a laser to target an object or surface and measuring the time it takes for the reflected light to return to a sensor, e.g., a LiDAR camera.

[0035] Radar, which stands for radio detection and ranging, is a detection technique that uses radio waves to determine the distance (ranging), angle, and radial velocity of an object relative to a site. A radar system consists of a transmitter that produces electromagnetic waves in the radio or microwave range, a transmitting antenna, a receiving antenna (often the same antenna is used for both transmitting and receiving), and a receiver, as well as a processor to determine the nature of the object. Radio waves from the transmitter reflect off objects and return to the receiver, providing information about the object's location and velocity. Radar signals can obtain distance measurements based on time of flight by transmitting short pulses of radio signal, measuring the time it takes for the radio signal reflection caused by a target object or surface to return.

[0036] Sonar, also known as sound navigation and ranging, is an acoustic location technique that uses acoustic propagation and reflection to measure distance to acoustically reflective surfaces or to detect objects. Sonar can be used to derive the contours of a surface by emitting pulses of sound and detecting the echoes reflected from the surface.

[0037] A time-of-flight sensor is a distance imaging camera system that employs time-of-flight techniques to calculate the distance between a camera and a point on a target surface by measuring the round-trip time of a signal, such as an artificial light signal emitted by a light source and reflected back to the camera by the point on the target surface. Time-of-flight sensors can be used to create a digital 3D representation of a surface by collecting depth information from multiple points on the surface using the time-of-flight of light (i.e., the time it takes for each light signal from the light source to hit the target point and return to the camera).

[0038] Depth sensors based on a multi-view camera setup system measure depth by capturing images of the same scene from different viewpoints using different cameras. Depth sensors use stereo photogrammetry, in which pixel depth data is determined from data acquired using the multi-view camera setup system. To solve the depth measurement problem using the multi-view camera setup system, corresponding points in different images captured by different cameras are identified. A disparity map can be constructed to show apparent pixel differences or motion between pairs of stereo images. Various techniques exist for deriving a depth map from a disparity map.

[0039] The optional output device 316 may include a display screen. In some embodiments, the output device 316 may also serve as an input device, for example, when the output device 316 is a touchscreen. The input / output controller 314 may also output data to a device other than an output device, for example, to a locally connected computing device. According to some embodiments, image processing and calculations based on data derived from images captured by the input device 318 and / or any other functionality described in the embodiments may be implemented by software or firmware, for example, the operating system 306 and application software 308, working together and / or independently, and executed by the processor 302.

[0040] The communication interface 312 allows the system 300 to communicate with other devices and systems. The communication interface 312 may include any type of signal transceiver, such as a 3G, 4G, and / or 5G wireless mobile telecommunications transceiver, a WiFi® signal transceiver and / or a Bluetooth® transceiver, as well as any wired telecommunications transceiver, such as an Ethernet and Thunderbolt interface.

[0041] The functionality described herein in embodiments may be implemented, at least in part, by one or more hardware logic components. According to one embodiment, the system 300 is configured by programs 306, 308 stored in the memory 304 that, when executed by the processor 302, perform embodiments of the described operations and functions. Alternatively, or in addition, the functionality described herein may be implemented, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and graphics processing units (GPUs).

[0042] 4 is a flow diagram of a method 400 according to some embodiments of the present invention. The method aims to verify the authentic capture of a real-world scene by an image capture device. The method prevents spoofing attacks, specifically PoP attacks.

[0043] Method 400 begins at step 402, in which motion of the image sensing device is detected during a time period immediately before, during, and / or immediately after the capture of at least one image and / or video by the image sensing device. The motion may be detected by any motion detection sensor, such as by at least one of a gyroscope, an accelerometer, a magnetometer, a pressure sensor, and a depth sensor included in or attached to the image sensing device. The time period immediately before, during, and / or immediately after the capture of the at least one image may be a predefined time period of any appropriate length. For example, the time period may have a length of 0.5 seconds, 1 second, 1.5 seconds, or 2 seconds. A time period immediately before, during, and / or immediately after the capture of at least one image means that the image capture may occur immediately before or immediately after this time period or at any time within this time period.

[0044] Method 400 then proceeds to step 404, which includes detecting a spoof attack based on analyzing the detection of any movement of the image sensing device. In some embodiments, this step includes detecting a spoof attack based on whether movement of the image sensing device within a time period immediately prior to, during, and / or immediately after the capture of the at least one image is detected to be within a predefined movement threshold.

[0045] In some embodiments, the predefined motion threshold is a threshold of the maximum rotation, maximum acceleration, or maximum displacement of the image sensing device within the time period. Alternatively, the predefined motion threshold may be a threshold of the average rotation, average acceleration, or average displacement of the image sensing device within the time period, or a threshold of the standard deviation of the rotation, acceleration, or displacement of the image sensing device within the time period. If the maximum or average rotation, acceleration, or displacement of the image sensing device within the time period does not exceed the threshold, it may be determined that a spoof attack is likely to exist. On the other hand, if the maximum or average rotation, acceleration, or displacement of the image sensing device within the time period exceeds the threshold, it may be determined that a spoof attack is unlikely to exist. This is because an adversary attempting to capture an image of a virtual scene displayed on a display screen and wanting to avoid an image that shows any indication that it was taken from the display screen would likely fix the camera in a position and carefully align the field of view of the image sensing device with the screen so that reflections from the screen are minimized in the image and so that the captured image or video does not show any edges of the screen. When an image sensing device is held in a user's hand without being attached to any other device to keep the image sensing device stable in a particular position at all times when capturing an image, there is typically natural movement due to the user's hand being unable to remain perfectly still in a fixed position in the air for any period of time. On the other hand, when an image sensing device is attached to another device to keep the image sensing device stable in a particular position at all times when capturing an image, the image sensing device will have little or substantially no movement during the period of time immediately before and / or after image capture.

[0046] The maximum or average displacement threshold for an image sensing device may be an empirical value obtained by measuring the displacement of an image sensing device, such as a camera or a mobile phone equipped with a camera, held by various people taking pictures. For example, if it is determined (e.g., by measuring the displacement of various image sensing devices held by several people while taking images) that a human user will almost certainly displace the image sensing device by at least 2 mm while taking an image, the empirical value, and therefore the threshold, may be set to 2 mm. Then, in use, if the maximum displacement of the image sensing device during image or video capture is within the predetermined threshold of 2 mm, it may be determined that the image sensing device is likely fixed in a stationary position, e.g., mounted on a stand, and a spoof attack is likely present for the reasons described in the previous paragraph.

[0047] Optionally, the result of step 404 is used to determine whether any subsequent processes should be activated. If it is determined in step 404 that a spoofing attack is likely, a remediation process may be activated. In one example, such a remediation process may be to restrict a user's access or further access to a particular function, application, or process. This is to prevent an adversary from accessing such function, application, or process for malicious purposes by creating a spoofing attack. In another example, such a remediation process may be to send an alert or notification to an administrator so that the administrator can take further steps to verify whether there is a spoofing attack. Such further steps may include, for example, reviewing the content of the image or any video recorded by the image sensing device immediately before, during, and / or immediately after the capture of the image to see if there are any further indications that the image was taken from a screen display. Such further indications may be reflections from the screen and / or any edges of the screen shown in the image and / or video.

[0048] 5 is a flow diagram of a method 500 according to a further embodiment of the present invention. The method aims to verify the authentic capture of a real-world scene by an image capture device. The method prevents spoofing attacks, in particular PoP attacks.

[0049] Step 502 in FIG. 5 is identical to step 402 in FIG. 4. Similar to method 400, method 500 begins with detecting motion of the image sensing device during a first time period immediately prior to, during, and / or immediately after the capture of at least one image by the image sensing device (step 502). Similar to step 402, in step 502, the motion may be detected by any motion detection sensor, such as by at least one of a gyroscope, an accelerometer, a magnetometer, a pressure sensor, and a depth sensor included in or attached to the image sensing device. The first time period immediately prior to, during, and / or immediately after the capture of the at least one image may be a predefined time period of any appropriate length. For example, the first time period may have a length of 0.5 seconds, 1 second, 1.5 seconds, or 2 seconds. The first time period immediately prior to, during, and / or immediately after the capture of at least one image means that the image capture may occur immediately prior to, or immediately after the first time period, or at any time within the first time period.

[0050] Method 500 further includes step 504. Step 504 may be performed before, after, or simultaneously with step 502. Step 504 includes receiving video captured in a second time period immediately before, during, and / or immediately after the capture of the at least one image by the image sensing device. In some embodiments, the video is captured by the same image sensing device that captured the at least one image. Alternatively, the video may be captured by a different image sensing device. The first and second time periods may overlap with each other, but these time periods need not have any overlap. The capture of the at least one image may occur immediately before or after the second time period or at any time within the second time period. As an optimization, the video captured in the second time period has a lower resolution, and as a result, the video does not necessarily occupy a large amount of storage space.

[0051] Following steps 502 and 504, method 500 proceeds to step 506, where detection of a spoof attack is performed based on analyzing the detection of any movement of the image sensing device in step 502 and the video received in step 504. In some embodiments, step 506 includes detecting the spoof attack based on whether the movement of the image sensing device within the first time period detected in step 502 is within a predefined movement threshold. This may include detecting the spoof attack based on whether the movement of the image sensing device within the first time period in step 502 is detected to be within the predefined movement threshold. In some examples, the predefined movement threshold is a threshold of maximum rotation, maximum acceleration, or maximum displacement of the image sensing device within the first time period. Alternatively, the predefined movement threshold may be a threshold of average rotation, average acceleration, or average displacement of the image sensing device within the first time period. If the maximum or average rotation, acceleration, or displacement of the image sensing device within the first time period does not exceed a threshold, indicating that the image sensing device is likely attached to another apparatus to keep the sensing device stable in a particular position, it may be determined that a spoof attack is highly likely to exist. On the other hand, if the maximum or average rotation, acceleration, or displacement of the image sensing device within the first time period exceeds a threshold, indicating that the image sensing device is likely being held in a user's hand, it may be determined that a spoof attack is unlikely to exist.

[0052] Additionally, step 506 also includes detecting a spoof attack based on the video received in step 504. In some embodiments, an algorithm for detecting screen edges in the video may be used to detect the presence of a display screen in the video, which may indicate a possible spoof attack. This may be done by detecting the contrast between the display area of ​​the screen and the bezel around the display area of ​​the screen, since the display area is typically much brighter than the bezel around the display area. Alternatively or additionally, since screen edges are usually composed of straight lines and right angles, an algorithm for detecting straight edges and / or right-angled corners may be used to identify the screen edges. If such a screen edge is detected, it may be determined that the video is likely taken from a scene displayed on the screen rather than a real-world scene, and therefore, a spoof attack is likely. Three-dimensional objects with prominent edge or blob features may be effectively recognized by detection algorithms such as Harris Affine Region Detection and Scale-Invariant Feature Transform (SIFT) methods. There are also various open source software programs for detecting 3D objects, one example of which can be found at https: / / docs.opencv.org / 4.x / d5 / d54 / group__objdetect.html.

[0053] Optionally, in step 506, detecting a spoofing attack includes estimating a risk of a spoofing attack based on the detection of motion of the image sensing device in step 502 and the video received in step 504. For example, if both the detection of motion of the image sensing device in step 502 and the video received in step 504 indicate a high possibility of a spoofing attack, the risk of a spoofing attack may be estimated to be high; if only one of the detection of motion of the image sensing device in step 502 and the video received in step 504 indicates a high possibility of a spoofing attack, the risk of a spoofing attack may be estimated to be medium; and if neither the detection of motion of the image sensing device in step 502 nor the video received in step 504 indicates a high possibility of a spoofing attack, the risk of a spoofing attack may be estimated to be low.

[0054] Optionally, the result of step 506 is used to determine whether any subsequent processes should be implemented. If a spoofing attack is likely to exist, e.g., a high and / or medium risk of a spoofing attack is determined in step 506, a remediation process may be implemented. In one example, such a remediation process may be restricting a user's access or further access to a particular function, application, or process. This is to prevent an adversary from accessing such function, application, or process for malicious purposes by creating a spoofing attack. In another example, such a remediation process may be sending an alert or notification to an administrator so that the administrator can take further measures to verify whether there is a spoofing attack. Such further measures may include, for example, reviewing the content of the image or any video recorded by the image sensing device immediately before, during, and / or immediately after the capture of the image to see if there are any further indications that the image was taken from a screen display. Such further indications may be reflections from the screen, specifically, unusual reflections and / or any edges of the screen shown in the image and / or video. Various methods exist for detecting reflections in images. An exemplary method for detecting reflections is presented in the paper "Reflection Detection in Image Sequences" by Mohamed Abdelaziz Ahmed, Francois Pitie, and Anil Kokaram of Sigmedia, Electronic and Electrical Engineering Department, Trinity College Dublin, presented at Proceedings / CVPR, the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, July 2011. The paper presents an automated technique for detecting reflections in image sequences.The technique is based on analyzing the spatiotemporal profiles of feature point trajectories and focuses on examining three main characteristics of reflectance: 1) the ability to decompose the image into two independent layers, 2) the image sharpness, and 3) the temporal behavior of image patches.

[0055] In an alternative embodiment of FIG. 5 , instead of receiving automatically captured video during a second time period immediately before, during, and / or after the capture of at least one image by the image sensing device as described in step 504, in an alternative embodiment, instructions may be given to a user of the image sensing device to move the image sensing device in a certain direction when the video is captured. Such instructions may be in the form of an on-screen prompt or any other visual or audio instruction suitable for instructing a user to move the image sensing device. For example, there may be an on-screen arrow instructing the user to rotate the image sensing device in a clockwise / counterclockwise direction and / or move the image sensing device up, down, left, or right. The purpose of instructing the user to move the camera while video is being captured is to increase the likelihood that the captured video can record and reveal any edge or corner of the display screen that can be used in a PoP attack. Such video may then be used in step 506 as described above. Optionally, the image sensing device can use a motion detection sensor to verify that the user is actually moving the image sensing device as instructed, while simultaneously capturing video during that movement. If the verification concludes that the movement of the image sensing device does not match the instructions given to the user, the video can be ignored, new instructions can be given to the user to move the image sensing device again, and video can be recorded again during the new movement of the sensing device. Such steps can be repeated until the movement of the image sensing device matches the instructions given to the user.

[0056] The terms "computer" or "computing device" are used herein to refer to any device having processing capability such that it is capable of executing instructions. Those skilled in the art will recognize that such processing capability is incorporated into many different devices, and thus the term "computer" includes PCs, servers, mobile phones, personal digital assistants, and many other devices.

[0057] Those skilled in the art will recognize that storage devices utilized to store program instructions may be distributed across a network. For example, a remote computer may store an example of a process described as software. A local or terminal computer may access the remote computer and download some or all of the software to execute the program. Alternatively, a local computer may download some software as needed, or execute some software instructions at a local terminal and some software instructions at a remote computer (or computer network). Those skilled in the art will also recognize that all or a portion of the software instructions may be executed by dedicated lines such as DSPs, programmable logic arrays, etc., utilizing conventional techniques known to those skilled in the art.

[0058] As will be apparent to one skilled in the art, any ranges or device values ​​given herein may be expanded or modified without losing the desired effect.

[0059] It should be understood that the benefits and advantages described above may relate to one embodiment or to several embodiments, and embodiments are not limited to those that solve any or all of the stated problems or have any or all of the stated benefits and advantages.

[0060] Any reference to "an" item refers to one or more of those items. The term "comprising" is used herein to mean inclusive of specified method blocks or elements, but that such blocks or elements do not comprise an exclusive list and that a method or apparatus may include additional blocks or elements.

[0061] The steps of the methods described herein may be performed in any suitable order, or simultaneously where appropriate. Additionally, individual blocks may be deleted from any of the methods without departing from the spirit and scope of the subject matter described herein. Aspects of any of the above-described examples may be combined with aspects of any of the other described examples to form further examples without losing the desired effect.

[0062] It should be understood that the above description of preferred embodiments is given by way of example only, and that various modifications may be made by those skilled in the art. While various embodiments have been described above with a certain degree of particularity or with reference to one or more individual embodiments, those skilled in the art may make numerous changes to the disclosed embodiments without departing from the spirit or scope of the invention. [Explanation of symbols]

[0063] 100 systems 102 User devices, image sensing devices 104 Communication Network 106 Server Devices 108 Scenes, Objects 200 systems 202 User devices, image sensing devices 204 Communication Network 206 Server Device 208 Screen, Object 210 Stationary Stand 300 System 302 processor 304 memory 306 Operating Systems, Programs 308 Application software and programs 312 Communication Interface 314 Input / Output Controller 316 Output Devices 318 Input Devices 320 Motion Detection Sensor 400 ways 500 ways

Claims

1. detecting movement of the image sensing device during a time period immediately prior to, during, and / or immediately after capture of at least one image or video by the image sensing device; detecting a spoof attack based on an analysis of the detection of any movement of the image sensing device; A computer-implemented method comprising:

2. The step of detecting a spoofing attack comprises: determining the presence of a spoof attack if the movement of the image sensing device during the time period is detected to be within a predefined movement threshold; 2. The method of claim 1, comprising:

3. the time period is a first time period; the method further comprising receiving video captured during a second time period; The method of claim 1 or 2, wherein the step of detecting a spoofing attack further comprises the step of detecting a spoofing attack based on the video.

4. The method of claim 3 , wherein the second period of time is immediately before, during, and / or immediately after the capture of the at least one image or video by the image sensing device.

5. The method of claim 3 , further comprising providing instructions to instruct a user of the image sensing device to move the image sensing device while the video is being captured.

6. The method of claim 5 , further comprising detecting movement of the image sensing device while the video is being captured, and ignoring the video if the detected movement does not correspond to the command.

7. the step of detecting spoofing comprises: if the movement of the image sensing device during the first time period is detected to be outside a predefined movement threshold; and / or If the captured image is determined to contain a flat edge, or if the captured image is detected to contain an anomalous reflection, 7. The method of claim 3, wherein the method determines the presence of a spoofing attack.

8. 8. The method of claim 2 or 7, wherein the predefined motion threshold is a threshold of maximum rotation, maximum acceleration, or maximum displacement of the image sensing device, a threshold of average rotation, average acceleration, or average displacement of the image sensing device within the time period, or a threshold of change in rotation, acceleration, or displacement of the image sensing device within the time period.

9. The method of claim 1 , wherein the movement is detected based on data received from at least one movement detection sensor.

10. The method of claim 9 , wherein the at least one motion detection sensor includes at least one of a gyroscope, an accelerometer, a magnetometer, a pressure sensor, and a depth sensor.

11. The method of claim 10 , wherein the depth sensor comprises at least one of a time-of-flight sensor, a LiDAR sensor, a radar sensor, a sonar sensor, and a multi-view camera setup depth sensor.

12. 12. An apparatus comprising at least one processor and at least one memory, the memory storing computer-implementable instructions that, when executed by the at least one processor, cause the at least one processor to perform the method of any one of claims 1 to 11.

13. A computer-readable medium storing computer-executable instructions configured to perform the method of any one of claims 1 to 11.

14. 10. A method substantially as described with reference to Figures 1 to 5 of the drawings.

15. 10. A system substantially as described with reference to Figures 1 to 5 of the drawings.