Method and system for verification

By detecting the movement of the image sensing device to identify spoofing attacks, the problem of the existing technology that cannot distinguish between the real world and the screen display image is solved, effective defense against picture-in-picture attacks is achieved, and the authenticity verification of image capture is improved.

CN120641896APending Publication Date: 2025-09-12OPENORIGINS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480008328.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-20
Filing Date
2024-01-19
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies cannot effectively distinguish whether an image or video is directly captured by a trusted capture device or is spoofed through a pre-recorded image displayed on the screen, resulting in an inability to prevent picture-in-picture (PoP) attacks.

Method used

By detecting the movement of the image sensing device within a certain period of time before, during and after image capture, and using motion detection sensors such as gyroscopes, accelerometers, magnetometers, pressure sensors and depth sensors, the stability of the device is analyzed to determine whether there is a spoofing attack.

Benefits of technology

Effectively identify and prevent picture-in-picture attacks, ensuring that captured images or videos are directly captured from the real world rather than virtual scenes displayed on the screen, improving the reliability and security of authentication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120641896A_ABST
    Figure CN120641896A_ABST
Patent Text Reader

Abstract

A computer-implemented method comprises: detecting movement of an image sensing device within a time period immediately before, during and / or immediately after capture of at least one image or video by the image sensing device; a spoofing attack is detected based on analysis of the detection of any movement of the image sensing device.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Embodiments of the present invention generally relate to methods, devices, and systems for authentication, and in particular, to methods, devices, and systems for authentication of image capture for the purpose of preventing spoofing attacks.

[0002] Digital media including visual information such as images and videos have widespread applications in the information age. Authentication systems often assume that the content of an image or video is authentic, especially if the image or video was captured by a device or application trusted by the viewer.

[0003] However, the capture of visual information is subject to various "spoofing" attacks, in which unauthorized adversaries or programs can create, modify, or reproduce visual information to gain unfair or even illegal advantages. These attacks attempt to present a false scene to the viewer. Fake scenes can be created by artificially setting scenes that mimic those expected by the viewer or authenticator. There is a continuing need for methods to verify whether images have been captured in a manner that deceives the viewer.

[0004] The embodiments described below are not limited to implementations that solve any or all of the shortcomings of known techniques for authentication or fraud detection. Summary of the Invention

[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0006] The invention is set forth in the appended claims.

[0007] A first aspect provides a computer-implemented method comprising: detecting movement of an image sensing device within a time period immediately before, during, and / or immediately after the image sensing device captures at least one image or video; and detecting a spoofing attack based on an analysis of any detected movement of the image sensing device. A second aspect provides an apparatus comprising at least one processor and at least one memory, the memory storing computer-implemented instructions that, when executed by the at least one processor, cause the at least one processor to perform the aforementioned method. A third aspect provides a computer-readable medium storing computer-executable instructions configured to perform the aforementioned method.

[0008] The methods described herein may be performed by software in a machine-readable form on a tangible storage medium, for example in the form of a computer program comprising computer program code means adapted to perform all the steps of any of the methods described herein when the program is run on a computer, and wherein the computer program may be implemented on a computer-readable medium. Examples of tangible (or non-transitory) storage media include disks, thumb drives, memory cards, and the like, and do not include propagating signals. The software may be suitable for execution on a parallel processor or a serial processor such that the method steps may be performed in any suitable order or simultaneously.

[0009] This recognizes that firmware and software can be valuable, individually tradable commodities. It is intended to cover software that runs on or controls "dumb" or standard hardware to perform a desired function. It is also intended to cover software that "describes" or defines the configuration of hardware, such as HDL (Hardware Description Language) software, when used to design silicon chips or to configure general-purpose programmable chips to perform a desired function.

[0010] It will be apparent to the skilled person that the preferred features may be combined where appropriate and with any of the aspects of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Embodiments of the present invention will be described by way of example with reference to the following drawings, in which:

[0012] Figure 1 is a schematic diagram of an environment in which a user device held in a user's hand captures an image or video;

[0013] Figure 2 is a schematic diagram of an environment in which a user device mounted on a stand captures an image or video of content displayed on a screen;

[0014] Figure 3 is a block diagram of an exemplary set of components of a user device in which embodiments of the present invention may be implemented;

[0015] Figure 4 is a flow chart of a method for verifying image capture according to some embodiments of the present invention;

[0016] Figure 5 is a flow chart of a method for verifying image capture according to some embodiments of the present invention.

[0017] Common reference numbers are used throughout the drawings to indicate similar features. DETAILED DESCRIPTION

[0018] The following description of embodiments of the present invention is by way of example only. These examples represent the best manner currently known to the applicant to put the present invention into practice, although they are not the only manner in which this can be accomplished. The description sets forth the functionality of the examples and the sequence of steps for constructing and operating the examples. However, the same or equivalent functionality and sequence may be achieved using different examples.

[0019] A common spoofing technique is the Picture-in-Picture (PoP) attack. In this attack, an adversary first obtains or creates an image or video containing visual features, optionally edits the image or video, and displays it on a screen. This image or video can represent a scene expected by a viewer or application (such as an authenticator program). Then, using a capture device trusted by the viewer or application, the adversary can capture an image or video of a screen displaying a pre-recorded or pre-created image or video of a scene containing the visual features, and upload the image or video of the screen for viewing by the viewer or processing by the application. Without any effective detection mechanism, even using a trusted capture device, the viewer or application cannot distinguish whether the scene was captured directly by a trusted device (implying authenticated content) or whether the scene was created or pre-recorded by a device untrusted by the viewer or application, possibly with some modifications, and then displayed to a trusted device (implying deceiving the viewer or application with unauthenticated content).

[0020] To make a PoP attack more difficult to detect, the adversary may hold the capture device steady at a particular location, for example using a fixed stand, a handheld gimbal, or a stabilized drone that can hover at one point, to align the capture device's field of view with the screen to capture an image or video of the screen. This arrangement facilitates PoP attacks by allowing the capture device to be fixed at a carefully selected location and orientation (e.g., the capture device may face in a direction roughly perpendicular to the screen) to minimize reflections from the screen and any distortion, as well as maximize the screen capture area so that the captured image or video does not show any borders of the screen. Otherwise, if the capture device is held in the adversary's hand, any slight movement of the hand may cause reflections from the screen and / or the borders of the screen to be captured in the image or video, which would make the PoP attack easily leakable because the viewer can more easily identify the presence of the screen used for the PoP attack from the reflections or borders of the screen.

[0021] Without any effective detection mechanism, a viewer or application may not be able to distinguish between a real-world scene captured by a capture device held in the user's hand and a virtual scene displayed on a 2D screen captured by a capture device fixed at a specific location. Therefore, there is a need for methods and systems for evaluating how a capture device captures images and videos (e.g., for determining whether the capture device remains stable at a specific location when capturing an image or video, and whether the image or video is captured directly from the real-world scene or is pre-recorded and then replayed on the screen to deceive the viewer or application).

[0022] In some embodiments / examples, a computer-implemented method is provided, comprising: detecting movement of an image sensing device within a time period immediately before, during, and / or immediately after the image sensing device captures at least one image or video; and detecting a spoofing attack based on an analysis of any detected movement of the image sensing device. Also provided is an apparatus comprising at least one processor and at least one memory storing computer-implemented instructions that, when executed by at least one controller, cause at least one other processor to perform the above method. Also provided is a computer-readable medium storing the computer-executable instructions configured to perform the method.

[0023] Figure 1 The environment in which the user takes an image or video of a scene is shown. Figure 1 As shown in FIG, system 100 includes a user device 102, a communication network 104, and a server device 106.

[0024] In by Figure 1 In the illustrated environment, user device 102 includes an image sensing device configured to capture visual information of scene 108. In various embodiments, scene 108 is a three-dimensional real-world environment. Figure 1 The scene 108 in FIG. 1 is depicted as a car, but it will be understood that a car is merely a non-limiting example and that the scene can be any three-dimensional real-world environment and can include any three-dimensional real-world object. In one embodiment, the user device 102 includes at least one sensor, such as a camera, for capturing images and / or video. The user device 102 may also include at least one processor for processing data associated with the captured images or video and at least one memory for storing raw data and / or processed data associated with the captured images or video. The user device 102 may also include a communication interface for sending data to and / or receiving data from the authenticator device 106 via the communication network 104. Optionally, the user device may also include a display screen for displaying the scene 108 being captured.

[0025] The communication network 104 may include any wired or wireless connection, the Internet, or any other form of communication. Figure 1 1 , the communication network 104 may include any number of different communication networks between the user device 102 and the server device 106. The communication network 104 is configured to facilitate communication between the user device 102 and the server device 106. Various implementations of the communication network 120 may employ different types of networks, such as, but not limited to, computer networks, telecommunication networks (e.g., cellular), mobile wireless data networks, and any combination of these and / or other networks.

[0026] Server device 106 is a computing device used to display or process captured images or videos received by server device 106 from user device 102. Server device 106 includes at least one processor and at least one memory for storing instructions and / or data to be processed by the at least one processor. Server device 106 can be configured to display images or videos captured by user device 102. Alternatively or additionally, server device 106 can be configured to analyze images or videos captured by user device 102. Server device 106 can also be configured to display the results of the analysis of images or videos captured by user device 102. In some examples, server device 106 can access information used to authenticate users, devices, applications, and / or processes, and / or access digital content. For example, this information can include pre-stored information such as identity information and / or account information of authorized users. Server device 106 can also include a communication interface for sending and / or receiving data to and from user device 102 via communication network 104.

[0027] Figure 2 2 shows a scenario where the user device 202 captures an image or video of a scene being displayed on the screen. Figure 2 As shown in FIG, system 200 includes a user device 202, a communication network 204, and a server device 206 mounted on a fixed support 210. User device 202, communication network 204, and authenticator device 206 may be the same as user device 102, communication network 104, and authenticator device 106, respectively, or perform functions similar to those performed by user device 102, communication network 104, and authenticator device 106, respectively.

[0028] System 200 also includes screen 208. Screen 208 is configured to display an image or video. The image or video may represent a scene. Such an image or video may be synthesized or pre-recorded, possibly with some modifications, and then displayed on screen 208 by a malicious user in an attempt to deceive the user or application of the server device. For example, screen 208 may display an image or video of a used car that a user is selling online. Users have a strong incentive to tamper with the image or video to make it appear better than it actually is in real life. Therefore, instead of directly capturing an image or video of the car using user device 202, which is trusted by server device 206, the user may modify or synthesize an image or video of the car, display the modified / synthesized image or video on screen 208, and then use user device 202 to capture an image or video of the modified image or video on screen 208. A malicious user may also falsely represent the car for sale by capturing an image or video of a car that is not the actual car for sale and displaying such an image or video on screen 208. Fixed mount 210 makes PoP attacks more difficult to detect because it minimizes movement of user device 202 during capture, making any reflections from screen 208 or the borders of screen 208 less likely to appear in the captured image or video. Fixed mount 210 can be any form of device suitable for allowing the user device to be secured in a fixed position. Even if user device 202 is a trusted device by the user or by an application on server device 206, without any effective detection mechanism, the user or application on server device 206 cannot determine whether user device 202 is mounted on a fixed mount when capturing an image or video, and cannot determine whether the scene shown in the image / video was captured directly by user device 202 or pre-recorded and then displayed to user device 202. The fixed mount is merely one exemplary way to stabilize the capture device in a specific position. Alternative methods exist for achieving similar effects, such as using a handheld gimbal or a stabilized drone that can hover in one spot to align the capture device's field of view with the display screen and capture an image or video of the screen.

[0029] Figure 3 3 is a block diagram of an exemplary set of components of a system 300 in which embodiments of the present invention may be implemented. The user device 102 / 202 in the system 100 / 200 may be implemented as the system 300.

[0030] System 300 can be implemented as one or more computing devices and / or electronic devices. System 300 includes one or more processors 302, which can be microprocessors, controllers, or any other suitable type of processor for processing computer-executable instructions to control the operation of system 300. Platform software including an operating system 306 or any other suitable platform software can be provided on the system to enable application software 308 to be executed on the system. In some embodiments, application software 308 can include software programs to process images, derive data from images, and process the data derived from the images according to the various methods described herein. The components of system 300 described herein can be enclosed in a housing.

[0031] Computer-executable instructions can be provided using any computer-readable medium that can be accessed by system 300. Computer-readable media can include, for example, computer storage media such as memory 304 and communication media. Computer storage media such as memory 304 include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVDs) or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by a computing device. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal (such as a carrier wave) or other transmission mechanism. Although the computer storage medium (memory 304) is shown within system 300, those skilled in the art will understand that at least a portion of the storage device can be distributed or located remotely and accessed via a network or other communication link (e.g., using communication interface 312).

[0032] The system 300 may include an input / output controller 314 that is arranged to receive and process input from one or more input devices 318, which may be separate from or integrated with the system 300, and may also be arranged to output information to one or more optional output devices 316, which may be separate from or integrated with the system 300. In some embodiments, the input devices 318 may include input devices such as a set of buttons or keys for controlling the operation of the system 300. For example, the input devices 318 may include keys for controlling at least one sensing device (such as adjusting the orientation and / or zoom of a camera, and / or for manipulating an image displayed on a screen). In some embodiments, the input devices may include at least one image sensing device such as a camera.

[0033] The input device 318 may also include at least one movement detection sensor 320 for detecting physical movement of the system 300, which may be implemented in the form of the user device 102 or 202 in the system 100 or 200. In some embodiments, the at least one movement detection sensor 320 includes at least one gyroscope, at least one accelerometer, at least one magnetometer, and / or at least one pressure sensor. Gyroscopes are known for measuring changes in orientation and angular velocity. Accelerometers are known for measuring the direction and magnitude of acceleration. Magnetometers are devices that measure the direction, strength, or relative changes in a magnetic field at a specific location. They can provide absolute angle measurements relative to the Earth's magnetic field. Pressure sensors (such as capacitive pressure sensors and piezoresistive strain gauge pressure sensors) can be used to detect force, tension, and / or movement applied to an object due to pressure applied to the object. These sensors are known for detecting movement of the object on which they are mounted.

[0034] The at least one motion detection sensor 320 illustrated above is used to detect changes in the position and / or acceleration of the system 300 to detect movement of the system. In another example, the at least one motion detection sensor 320 may be at least one depth sensor.

[0035] At least one depth sensor can provide depth information relative to a surface or scene from a viewpoint. The depth sensor measures the distance, or "depth" of view, between the depth sensor and the object or surface. For example, the depth sensor can measure depth using stereo image sensing techniques or the round-trip time of reflected depth sensing signals. Any change in depth indicates relative movement between the depth sensor included in system 300 and the object or surface that reflected the depth sensing signal back to the depth sensor. This type of depth sensor can be used to measure relative movement between image sensing device 102 / 202 and an object 108 / 208 captured by the image sensing device.

[0036] In some embodiments, the at least one depth sensor includes at least one of a radio wave-based depth sensor (such as a radar sensor), a light-based depth sensor (such as a LiDAR sensor), an acoustic depth sensor (such as a sonar sensor), and a multi-view camera device setting system.

[0037] LiDAR (also known as "Light Detection and Ranging" or "Laser Imaging, Detection, and Ranging") is a time-of-flight technology used to determine range (variable distance) by pointing a laser at an object or surface and measuring the time it takes for the reflected light to return to a sensor (e.g., a LiDAR camera).

[0038] Radar, which stands for Radio Detection and Ranging, is a detection technology that uses radio waves to determine the distance (ranging), angle, and radial velocity of an object relative to a location. A radar system consists of a transmitter that generates electromagnetic waves in the radio or microwave domain, a transmitting antenna, a receiving antenna (usually the same antenna is used for both transmission and reception), a receiver, and a processor to determine the properties of an object. The radio waves from the transmitter reflect off the object and return to the receiver to provide information about the object's position and velocity. Radar signals can be used to obtain time-of-flight distance measurements by transmitting short pulses of radio signals and measuring the time it takes for the reflected radio signal, caused by the target object or surface, to return.

[0039] Sonar (also known as sound navigation and ranging) is an acoustic positioning technology that uses sound propagation and reflection to measure the distance to a sound-reflecting surface or detect objects. Sonar can derive the contours of a surface by emitting pulses of sound and detecting the echoes that reflect from the surface.

[0040] A time-of-flight sensor is a ranging imaging camera system that uses time-of-flight technology to calculate the distance between the camera of the ranging imaging camera system and a point on a target surface by measuring the round-trip time of a signal (such as an artificial light signal emitted by a light source and reflected by a point on the target surface back to the camera). A time-of-flight sensor can be used to collect depth information from multiple points on a surface by using the time of flight of light (i.e., the time it takes for each light signal from the light source to hit the target point and return to the camera) to create a digital three-dimensional representation of the surface.

[0041] Depth sensors based on multi-view camera setups measure depth by using different cameras to capture images of the same scene from different perspectives. Depth sensors based on multi-view camera setups use stereo photogrammetry to determine pixel depth data from data acquired using the multi-view camera setup. To address depth measurement issues using a multi-view camera setup, corresponding points in different images captured by different cameras are identified. A disparity map can be constructed to indicate apparent pixel differences or motion between a pair of stereo images. Various techniques exist for deriving a depth map from a disparity map.

[0042] The optional output device 316 may include a display screen. In some embodiments, for example, when the output device 316 is a touch screen, the output device 316 may also serve as an input device. The input / output controller 314 may also output data to a device other than the output device, such as a locally connected computing device. According to some embodiments, image processing and calculations based on data derived from images captured by the input device 318 and / or any other functions described in the embodiments may be implemented by software or firmware (e.g., an operating system 306 and application software 308 working together and / or independently) and executed by the processor 202.

[0043] The communication interface 312 enables the system 300 to communicate with other devices and systems. The communication interface 312 may include any type of signal transceiver (such as a 3G, 4G and / or 5G wireless mobile telecommunications transceiver, WiFi TM Signal transceiver and / or Bluetooth TM transceivers) and any wired telecommunications transceivers such as Ethernet and Thunderbolt interfaces.

[0044] The functions described herein in the embodiments may be performed at least in part by one or more hardware logic components. According to an embodiment, when executed by the processor 302, the computing device 300 is configured by the programs 306, 308 stored in the memory 304 to perform the described operations and functional embodiments. Alternatively or in addition, the functions described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and graphics processing units (GPUs).

[0045] Figure 4is a flow chart of a method 400 according to some embodiments of the present invention. The method aims to verify an authenticated capture of a real-world scene by an image capture device. The method prevents spoofing attacks, and in particular PoP attacks.

[0046] Method 400 begins with step 402, in which movement of an image sensing device is detected within a time period immediately before, during, and / or immediately after the image sensing device captures at least one image and / or video. Movement can be detected by any motion detection sensor, such as at least one of a gyroscope, accelerometer, magnetometer, pressure sensor, and depth sensor, at least one of which is included in or attached to the image sensing device. The time period immediately before, during, and / or immediately after the capture of the at least one image can be a predefined time period of any suitable length. For example, the time period can be 0.5 seconds, 1 second, 1.5 seconds, or 2 seconds in length. The time period immediately before, during, and / or immediately after the capture of the at least one image means that the image capture can occur immediately before, after, or at any time within the time period.

[0047] Method 400 then proceeds to step 404, which involves detecting a spoofing attack based on analyzing the detection of any movement of the image sensing device. In some embodiments, this step includes detecting a spoofing attack based on whether movement of the image sensing device is detected within a predefined movement threshold within a time period immediately before, during, and / or immediately after the capture of the at least one image.

[0048] In some embodiments, the predefined movement threshold is a threshold for the maximum rotation, maximum acceleration, or maximum displacement of the image sensing device within a time period. Alternatively, the predefined movement threshold may be a threshold for the average rotation, average acceleration, or average displacement of the image sensing device within the time period, or a threshold for the standard deviation of the rotation, acceleration, or displacement of the image sensing device within the time period. If the maximum or average rotation, acceleration, or displacement of the image sensing device within the time period does not exceed the threshold, a spoofing attack may be determined to be likely. On the other hand, if the maximum or average rotation, acceleration, or displacement of the image sensing device within the time period exceeds the threshold, a spoofing attack may be determined to be unlikely. This is because an adversary attempting to capture an image of a virtual scene displayed on a display screen and avoiding the image showing any indication of the image being captured from the display screen may fix the camera device in a certain position and carefully align the image sensing device's field of view with the screen to minimize reflections from the screen in the image and ensure that the captured image or video does not show any boundaries of the screen. If the image sensing device is held by a user's hand, and not mounted to any other device to hold the image sensing device steady at a particular position while an image is captured, there will typically be natural movement, as the user's hand cannot be completely fixed in a fixed position in the air for a period of time. On the other hand, if the image sensing device is mounted to another device to hold the image sensing device steady at a particular position while an image is captured, the image sensing device will experience very little or no movement in the period immediately before and / or after the image is captured.

[0049] The threshold value for the maximum or average displacement of the image sensing device can be an empirical value obtained by measuring the displacement of the image sensing device (such as a camera or a mobile phone including a camera) held by various people taking pictures. For example, if it is determined (for example, by measuring the displacement of various image sensing devices held by multiple people while taking images) that human users almost certainly displace the image sensing device by at least 2 mm while taking images, the empirical value, and therefore the threshold value, can be set to 2 mm. Then, in use, if the maximum displacement of the image sensing device during image or video capture is within the predetermined threshold of 2 mm, the image sensing device is likely to be fixed to a fixed position, such as mounted on a bracket, and it can be determined that a spoofing attack is likely present for the reasons mentioned in the previous paragraph.

[0050] Optionally, the result of step 404 is used to determine whether to activate any subsequent processes. If it is determined in step 404 that a spoofing attack may have occurred, a remediation process may be activated. In one example, such a remediation process limits the user's access to or additional access to specific functions, applications, or processes. This is to prevent an adversary conducting a spoofing attack from accessing such functions, applications, or processes for malicious purposes. In another example, such a remediation process may send an alert or notification to an administrator so that the administrator can take additional action to verify whether a spoofing attack has occurred. Such additional actions may include, for example, reviewing the content of the image or any video recorded by the image sensing device immediately before, during, and / or immediately after the image is captured to see if there are any additional indications that the image was captured from the screen display. Such additional indications may be reflections from the screen and / or any boundaries of the screen shown in the image and / or video.

[0051] Figure 5 is a flow chart of a method 500 according to another embodiment of the present invention. The method aims to verify an authenticated capture of a real-world scene by an image capture device. The method prevents spoofing attacks, and in particular PoP attacks.

[0052] Figure 5 Step 502 in Figure 4 Method 500 is similar to step 402 in method 400. As with method 400, method 500 begins with a step (step 502) of detecting movement of an image sensing device within a first time period immediately before, during, and / or immediately after the image sensing device captures at least one image. As with step 402, in step 502, movement can be detected by any motion detection sensor, such as at least one of a gyroscope, accelerometer, magnetometer, pressure sensor, and depth sensor, at least one of which is included in or attached to the image sensing device. The first time period immediately before, during, and / or immediately after the capture of the at least one image can be a predefined time period of any suitable length. For example, the first time period can be 0.5 seconds, 1 second, 1.5 seconds, or 2 seconds. The first time period being immediately before, during, and / or immediately after the capture of the at least one image means that the image capture can occur immediately before or after the first time period, or at any time within the first time period.

[0053] Method 500 also includes step 504. Step 504 can be performed before, after, or simultaneously with step 502. Step 504 includes receiving a video captured within a second time period immediately before, during, and / or immediately after the image sensing device captures the at least one image. In some embodiments, the video is captured by the same image sensing device that captured the at least one image. Alternatively, the video can be captured by a different image sensing device. The first time period and the second time period can overlap with each other, although it is not necessary for them to have any overlap. Capturing the at least one image can occur immediately before or after the second time period, or at any time within the second time period. As an optimization, the video captured within the second time period has a low resolution so that it does not necessarily take up much storage space.

[0054] After steps 502 and 504, method 500 proceeds to step 506, where detection of a spoofing attack is performed based on analyzing any movement of the image sensing device detected in step 502 and the video received in step 504. In some embodiments, step 506 includes detecting a spoofing attack based on whether the movement of the image sensing device detected in step 502 during the first time period is within a predefined movement threshold. This may include detecting a spoofing attack based on whether the movement of the image sensing device detected in step 502 during the first time period is within the predefined movement threshold. In some examples, the predefined movement threshold is a threshold for the maximum rotation, maximum acceleration, or maximum displacement of the image sensing device during the first time period. Alternatively, the predefined movement threshold may be a threshold for the average rotation, average acceleration, or average displacement of the image sensing device during the first time period. If the maximum or average rotation, acceleration, or displacement of the image sensing device during the first time period does not exceed the threshold (indicating that the image sensing device may be mounted on another device to maintain a stable sensing device position), a spoofing attack may be determined to have occurred. On the other hand, if the maximum or average rotation, acceleration, or maximum displacement of the image sensing device during the first time period exceeds a threshold (which indicates that the image sensing device is likely held in the user's hand), it may be determined that a spoofing attack is unlikely.

[0055] In addition, step 506 also includes detecting a spoofing attack based on the video received in step 504. In some embodiments, an algorithm can be used to detect the boundaries of the screen in the video to note the presence of a display screen in the video, which can indicate the possibility of a spoofing attack. This can be achieved by detecting the contrast between the display area of ​​the screen and the border surrounding the display area of ​​the screen, as the display area is typically much brighter than the border surrounding it. Alternatively or in addition, since the boundaries of the screen are typically composed of straight lines and right angles, an algorithm can be used to detect straight edges and / or right angles to identify the boundaries of the screen. If such boundaries of the screen are detected, it can be determined that there is a high probability that the video was shot from the scene displayed on the screen rather than a real-world scene, and therefore a high probability of a spoofing attack. Three-dimensional objects with obvious edge features or spot features can be effectively identified by detection algorithms (such as Harris affine region detection and scale-invariant feature transform (SIFT) methods). There are also various open source software programs for detecting three-dimensional objects, examples of which can be found at the following websites: https: / / docs.opencv.org / 4.x / d5 / d54 / groupobjdetect.html .

[0056] Optionally, in step 506, the detection of the spoofing attack includes estimating a risk of a spoofing attack based on the movement of the image sensing device detected in step 502 and the video received in step 504. For example, if both the movement of the image sensing device detected in step 502 and the video received in step 504 indicate a high likelihood of a spoofing attack, the risk of a spoofing attack may be estimated to be high; if only one of the movement of the image sensing device detected in step 502 and the video received in step 504 indicates a high likelihood of a spoofing attack, the risk of a spoofing attack may be estimated to be medium; and if neither the movement of the image sensing device detected in step 502 nor the video received in step 504 indicates a high likelihood of a spoofing attack, the risk of a spoofing attack may be estimated to be low.

[0057] Optionally, the result of step 506 is used to determine whether any subsequent processes are to be performed. If it is determined in step 506 that a spoofing attack may exist, such as a high and / or medium risk of a spoofing attack, a remediation process may be performed. In one example, such a remediation process may limit user access to or additional access to specific functions, applications, or processes. This is to prevent an adversary conducting a spoofing attack from accessing such functions, applications, or processes for malicious purposes. In another example, such a remediation process may send an alert or notification to an administrator so that the administrator can take additional action to verify whether a spoofing attack has occurred. Such additional actions may include, for example, reviewing the content of the image or any video recorded by the image sensing device immediately before, during, and / or after the capture of the image to see if there are any additional indications that the image was captured from the screen display. Such additional indications may be reflections, particularly unusual reflections, from the screen and / or any boundaries of the screen shown in the image and / or video. Various methods exist for detecting reflections in images. Mohamed Abdelaziz Ahmed, Francois Pitie, and Anil Kokaram of Sigmedia, Department of Electronic and Electrical Engineering, Trinity College Dublin, published an exemplary method for detecting reflections in their paper, "Reflection Detection in Image Sequences," published in the Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR) in July 2011. The paper proposes an automatic technique for detecting reflections in image sequences. It is based on analyzing the spatiotemporal profiles of feature point trajectories and focuses on examining three key characteristics of reflections: 1) the ability to decompose the image into two independent layers, 2) image sharpness, and 3) the temporal behavior of image patches.

[0058] exist Figure 5In an alternative embodiment, instead of automatically capturing video within a second time period immediately before, during, and / or immediately after the image sensing device captures at least one image, as described in step 504, instructions may be given to the user of the image sensing device to move the image sensing device in a certain manner while capturing video. Such instructions may take the form of on-screen prompts or any other visual or audio instructions suitable for directing the user to move the image sensing device. For example, arrows may be present on the screen to instruct the user to rotate the image sensing device clockwise / counterclockwise and / or move the image sensing device up, down, to the left, or to the right. The purpose of instructing the user to move the camera while capturing video is to increase the chance that the captured video will capture and reveal any edges or corners of the display screen that may be used for a PoP attack. Such video may then be used in step 506 as described above. Optionally, the image sensing device may use a motion detection sensor to verify that the user is indeed moving the image sensing device as instructed, while simultaneously capturing video during the movement. If the verification concludes that the movement of the image sensing device does not comply with the instructions given to the user, the video may be ignored, and new instructions may be given to the user to move the image sensing device again, and the video may be recorded again during the new movement of the sensing device. Such steps may be repeated until the movement of the image sensing device complies with the instructions given to the user.

[0059] The term "computer" or "computing device" is used herein to refer to any device having processing capabilities that enable it to execute instructions. Those skilled in the art will recognize that such processing capabilities are incorporated into many different devices, and thus the term "computer" includes PCs, servers, mobile phones, personal digital assistants, and many other devices.

[0060] Those skilled in the art will recognize that the storage device for storing program instructions can be distributed across a network. For example, a remote computer can store an example of a process described as software. A local or terminal computer can access the remote computer and download a portion or all of the software to run the program. Alternatively, the local computer can download the software as needed, or execute some software instructions at a local terminal, and execute some software instructions at a remote computer (or computer network). Those skilled in the art will also recognize that, by utilizing conventional techniques known to those skilled in the art, all or part of the software instructions can be executed by a dedicated circuit (such as a DSP, a programmable logic array, etc.).

[0061] As will be apparent to the skilled artisan, any range or device value given herein may be expanded or altered without losing the effect sought.

[0062] It should be understood that the benefits and advantages described above may relate to one embodiment, or may relate to multiple embodiments. The embodiments are not limited to embodiments that solve any or all of the problems described or have any or all of the benefits and advantages described.

[0063] Any reference to items modified by the indefinite article an refers to one or more of those items. The term "comprising" is used herein to mean including the identified method blocks or elements, but that such blocks or elements do not comprise an exclusive list and the method or apparatus may contain additional blocks or elements.

[0064] The steps of the methods described herein may be performed in any suitable order, or simultaneously where appropriate. Additionally, individual blocks may be deleted from any of the methods without departing from the spirit and scope of the subject matter described herein. Aspects of any of the examples described above may be combined with aspects of any of the other examples described to form additional examples without losing the effects sought.

[0065] It will be understood that the above description of the preferred embodiments is given by way of example only and that various modifications may be made by those skilled in the art. Although various embodiments have been described above with a certain degree of particularity or with reference to one or more individual embodiments, those skilled in the art may make various changes to the disclosed embodiments without departing from the spirit or scope of the invention.

Claims

1. A computer-implemented method comprising: detecting movement of an image sensing device within a time period immediately before, during, and / or immediately after capture of at least one image or video by the image sensing device, A spoofing attack is detected based on analysis of the detection of any movement of the image sensing device.

2. The method according to claim 1, wherein The detection of spoofing attacks includes, If the detected movement of the image sensing device within the time period is within a predefined movement threshold, it is determined that a spoofing attack exists.

3. The method according to claim 1 or 2, in, The time period is a first time period, The method further comprises receiving a video captured during a second time period, and The detecting of the deception attack further comprises detecting the deception attack based on the video.

4. The method according to claim 3, wherein: The second time period is immediately before, during and / or immediately after the capturing of the at least one image or video by the image sensing device. 5 . The method of claim 3 , further comprising providing instructions for directing a user of the image sensing device to move the image sensing device when capturing the video.

6. The method according to claim 5, further comprising: detecting movement of the image sensing device while capturing the video; And if the detected movement does not correspond to the instruction, ignoring the video.

7. The method according to any one of claims 3 to 6, wherein The detection of deception determines that a deception attack exists when: detecting that the movement of the image sensing device during the first time period is outside a predefined movement threshold, and / or It is determined that the captured video includes an edge of a flat surface, or it is detected that the captured video contains an unusual reflection.

8. The method according to claim 2 or 7, wherein: The predefined movement threshold is a threshold of the maximum rotation, maximum acceleration or maximum displacement of the image sensing device, a threshold of the average rotation, average acceleration or average displacement of the image sensing device within the time period, or a threshold of the change in the rotation, acceleration or displacement of the image sensing device within the time period.

9. The method according to any one of claims 1 to 4, wherein The movement is detected based on data received from at least one movement detection sensor.

10. The method according to claim 9, wherein: The at least one movement detection sensor includes at least one of a gyroscope, an accelerometer, a magnetometer, a pressure sensor, and a depth sensor.

11. The method according to claim 10, wherein: The depth sensor includes at least one of a time-of-flight sensor, a LiDAR sensor, a radar sensor, a sonar sensor, and a depth sensor provided by a multi-view camera device.

12. A device comprising: At least one processor and at least one memory storing computer-implementable instructions which, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 11.

13. A computer-readable medium storing computer-executable instructions configured to perform the method according to any one of claims 1 to 11.

14. A method substantially as hereinbefore described with reference to Figures 1 to 5 of the accompanying drawings.

15. A system substantially as hereinbefore described with reference to Figures 1 to 5 of the accompanying drawings.