Video inspection method and apparatus

By combining environmental image sequences and sensor data matching methods, the problem of deepfake videos being difficult to identify has been solved, achieving more accurate video authenticity detection.

WO2026103524A1PCT designated stage Publication Date: 2026-05-21ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2025-10-30
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing video detection technologies struggle to accurately identify deepfake videos, threatening personal privacy and the authenticity of information.

Method used

By acquiring environmental image sequences and sensor data from the target video, and using methods such as feature point analysis and multi-modal contrastive learning, the motion information of the device is matched to determine the authenticity of the video.

Benefits of technology

This improves the accuracy of detecting fake videos, ensuring the authenticity and security of the videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025131163_21052026_PF_FP_ABST
    Figure CN2025131163_21052026_PF_FP_ABST
Patent Text Reader

Abstract

A video inspection method and apparatus. The method comprises: on the basis of a video frame sequence corresponding to a target video, obtaining an environment image sequence; acquiring target sensor data corresponding to the target video, the target sensor data being provided by a target sensor configured in a target device that acquires the target video, and the target sensor being used for sensing motion of the target device; and, on the basis of the environment image sequence and the target sensor data, determining whether the target video is a fake video.
Need to check novelty before this filing date? Find Prior Art

Description

Video detection method and device

[0001] This application claims priority to Chinese Patent Application No. 202411636918.5, filed on November 14, 2024, entitled "Video Detection Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0002] The embodiments in this specification belong to the field of computer technology, and in particular relate to a video detection method and apparatus. Background Technology

[0003] Deepfake is a simulation technique that uses deep learning algorithms to synthesize videos or audio. It can forge facial features or mimic voices in videos or audio. With the rapid development of deepfake technology, creating fake videos has become increasingly easy and realistic, posing a serious threat to personal privacy, information authenticity, and social trust.

[0004] We hope to find a new technological solution that can more accurately detect whether a video is fake. Summary of the Invention

[0005] The purpose of this invention is to provide a video detection method and apparatus.

[0006] In a first aspect, a video detection method is provided, comprising: obtaining an environmental image sequence based on a video frame sequence corresponding to a target video; acquiring target sensor data corresponding to the target video, wherein the target sensor data is provided by a target sensor configured in a target device that acquires the target video, and the target sensor is used to sense the motion of the target device; and determining whether the target video is a fake video based on the environmental image sequence and the target sensor data.

[0007] Secondly, a video detection device is provided, the device comprising: a video processing unit configured to obtain an environmental image sequence based on a video frame sequence corresponding to a target video; a data acquisition unit configured to acquire target sensor data corresponding to the target video, the target sensor data being provided by a target sensor configured in a target device that acquires the target video, the target sensor being used to sense the motion of the target device; and a detection processing unit configured to determine whether the target video is a fake video based on the environmental image sequence and the target sensor data.

[0008] Thirdly, a computing device is provided, including a memory and a processor, wherein the memory stores computer programs / instructions, and the processor executes the computer programs / instructions to implement the method described in the first aspect.

[0009] Fourthly, a computer-readable storage medium is provided having a computer program / instructions stored thereon, wherein when the computer program / instructions are executed in a computing device, the computing device performs the method described in the first aspect.

[0010] In the technical solutions provided in the embodiments of this specification, the environmental image sequence obtained based on the video frame sequence corresponding to the target video can reflect the motion of the target device during the process of acquiring the target video; the target sensor configured in the target device, which can be used to sense the motion of the target device, and the target sensor data collected during the process of the target device acquiring the target video can also reflect the motion of the target device during the process of acquiring the target video; when the target video is a real video acquired by the target device and not a fake video that has been injected with an attack, the motion of the target device reflected by the environmental image sequence and the motion of the target device reflected by the target sensor data should be consistent. Therefore, it is possible to more accurately determine whether the target video is a fake video based on the environmental image sequence and the target sensor data. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 is a schematic diagram of the technical scenario of the technical solution provided in the embodiments of this specification;

[0013] Figure 2 is a flowchart of a video detection method provided in the embodiments of this specification;

[0014] Figure 3 is a schematic diagram of data flow in a video detection process provided in an embodiment of this specification;

[0015] Figure 4 is a structural schematic diagram of a video detection device provided in the embodiments of this specification. Detailed Implementation

[0016] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0017] With the rapid development of Deepfake technology, the creation of fake videos has become increasingly easy and realistic, posing a serious threat to personal privacy, information authenticity, and social trust. Intruders may use attack methods, including Deepfake technology, to maliciously modify real videos to obtain fake videos, or even directly forge fake videos without any real videos available.

[0018] In technical scenarios involving video capture, such as eKYC (Electronic Know Your Customer), facial recognition, video calls, and iris recognition, it is necessary to effectively identify whether videos from relevant devices are fake in order to defend against identity threats posed by various attack methods, including Deepfake technology.

[0019] The inventors discovered that target devices equipped with image acquisition units typically include target sensors to detect the movement of the target device. During the process of the target device acquiring video through the image acquisition unit, the positional relationship between the image acquisition unit and the target sensor usually remains unchanged. Therefore, in genuine videos acquired by the target device that have not been maliciously modified, changes in the environment are correlated with the movement of the target device; however, in fake videos, this correlation is often disrupted.

[0020] In view of this, the embodiments of this specification provide at least one video detection method, apparatus, computing device, and computer-readable storage medium. The environmental image sequence obtained based on the video frame sequence corresponding to the target video can reflect the motion of the target device during the acquisition of the target video. The target sensor configured in the target device, which can be used to sense the motion of the target device, can also reflect the motion of the target device during the acquisition of the target video by collecting target sensor data. When the target video is a real video acquired by the target device and not a fake video that has been injected with an attack, the motion of the target device reflected by the environmental image sequence and the motion of the target device reflected by the target sensor data should be consistent. Therefore, based on the environmental image sequence and the target sensor data, it is possible to more accurately determine whether the target video is a fake video.

[0021] Figure 1 is a schematic diagram of the technical scenario of the technical solution provided in the embodiments of this specification.

[0022] Referring to Figure 1, the target device equipped with an image acquisition unit also includes a target sensor for sensing the movement of the target device. The target sensor refers to the spatial attitude measurement unit within the target device, and may include, but is not limited to, one or more of various sensors such as accelerometers, gyroscopes, and magnetometers. For example, the target sensor may be an inertial measurement unit (IMU), which integrates three single-axis accelerometers and three single-axis gyroscopes, and may even integrate a magnetometer to assist in determining the target device's own attitude. The aforementioned target device may include, but is not limited to, mobile phones, tablet computers, and personal computers (PCs).

[0023] In various technical scenarios involving video capture, during the process of the target device acquiring the target video through its own configured image acquisition unit, the target sensor configured in the target device can also simultaneously acquire target sensor data corresponding to the target video. When the target video needs to be sent to a relevant recipient, such as a server of a service platform, the target sensor data can also be sent to the recipient of the target video so that the recipient can detect whether the target video is fake.

[0024] Figure 2 is a flowchart of a video detection method provided in an embodiment of this specification. This method can be executed by any device, platform, equipment, or cluster of devices with computing or processing capabilities, such as a server on a related service platform.

[0025] Referring to Figure 2, the method may include, but is not limited to, some or all of the following steps S201 to S205.

[0026] Step S201: Obtain an environmental image sequence based on the video frame sequence corresponding to the target video.

[0027] In one possible implementation, the environmental image sequence may be an unprocessed video frame sequence.

[0028] Depending on the specific tasks required to utilize the target video in subsequent processes, such as biometric identification of the target user (e.g., facial recognition or iris recognition), the target video may contain human image regions. However, if the person within the image acquisition unit's field of view is not stationary during the acquisition of the target video by the relevant target device, the human image region will fail to reflect the movement of the target device during video acquisition, and will also incur significant resource overhead in the subsequent prediction of the target device's movement.

[0029] In one possible implementation, techniques such as object detection and human face segmentation can be used to process the video frame sequence, eliminating human image regions in the video frames and obtaining an environmental image sequence corresponding to the video frame sequence. This allows for more efficient and accurate prediction of the target device's motion based on the environmental image sequence. For example, each video frame in the video frame sequence can first be processed using various object detection models, including Grounding Dino, to identify and mark human face image regions in each video frame. Then, in each video frame, the pixel values ​​of all pixels located within the human face image regions are set to a preset value, such as 0, thereby obtaining an environmental image sequence corresponding to the video frame sequence.

[0030] Step S203: Obtain target sensor data corresponding to the target video, wherein the target sensor data is provided by the target sensor configured in the target device that acquires the target video, and the target sensor is used to sense the motion of the target device.

[0031] Target sensor data refers to the sensor data collected simultaneously by the target device's configured target sensor during the acquisition of target video by its configured image acquisition unit. For example, the video frame sequence corresponding to the target video includes N video frames arranged sequentially, where the acquisition time of the first video frame is time t1, and the acquisition time of the Nth video frame is time t... N The target sensor data can be the data from the target sensor at time t1 to t2. N Sensor data collected continuously.

[0032] As mentioned above, a target sensor refers to a spatial attitude measurement unit in a target device, which may include, but is not limited to, one or more of various sensors such as accelerometers, gyroscopes, and magnetometers. For example, a target sensor may be an inertial measurement unit, which integrates three single-axis accelerometers and three single-axis gyroscopes, and may even integrate a magnetometer.

[0033] Step S205: Determine whether the target video is a fake video based on the environmental image sequence and target sensor data.

[0034] In one possible implementation, the first pose change information corresponding to the target device can be predicted based on the environmental image sequence; the second pose change information corresponding to the target device can be predicted based on the target sensor data; it can be determined whether the first pose change information and the second pose change information match; if not, the target video can be determined to be a fake video.

[0035] The process of predicting the first pose change information corresponding to the target device is described below as an example.

[0036] Various methods, such as the feature point method, the direct method, or training a machine learning model, can be used to process the environmental image sequence to obtain the first pose change information of the target device. Taking the feature point method as an example, it is first assumed that the environmental image sequence includes N environmental images arranged in sequence, and the acquisition time of the video frame corresponding to any i-th environmental image acquired by the target device is denoted as t. i Then: First, feature detection algorithms such as ORB (Oriented Fast and Rotated BRIEF), scale-invariant feature transform (SIFT), or speeded-up robust features (SURF) can be used to extract feature points from N environmental images. Next, brute-force matching or a fast approximate nearest neighbor search based on FLANN (Fast Library for Approximate Nearest Neighbors) can be used to match multiple feature point pairs between the i-th environmental image and the (i-1)-th environmental image. Then, the RANSAC (random sample consensus) algorithm or other possible filtering algorithms can be used to remove potentially anomalous feature point pairs. Finally, based on the filtered feature point pairs, the essential matrix corresponding to the i-th environmental image and the (i-1)-th environmental image is calculated. The third relative pose rotation S of the i-th environmental image relative to the (i-1)-th environmental image is obtained by decomposing the essential matrix. 1i That is, obtaining time t i The corresponding third relative pose S 1i In the aforementioned example, the third relative pose S 1i It can include rotation information R 1i and displacement vector T 1i Rotation information R 1i Used to indicate the target device at time t i Below the target device at time t i-1Below, the rotation angle occurring along the x, y, and z axes in three-dimensional space; displacement vector T 1i Used to indicate the target device at time t i Below the target device at time t i-1 The distance of displacement along the x, y, and z axes in three-dimensional space. It should be noted that when i is 1, the values ​​of each data item in the first relative pose and the third relative pose of the first environmental image relative to the non-existent zeroth environmental image can be recorded as 0.

[0037] The time intervals t1 to t2 are obtained through the process described in the previous example. N After obtaining the third relative pose corresponding to each of the N time points, the change information of the first pose can be determined. For example, the time points t1 to t2 can be used to determine the change information of the first pose. N The third relative pose corresponding to each of the N time points is determined as the first pose change information. For example, the time points t1 to t2 can be used as the first pose change information. N The rotational information included in the third relative pose corresponding to each of the N time points is determined as the first pose change information. For example, the time points t1 to t2 can be used as the first pose change information. N The displacement vectors included in the third relative pose at each of the N time points are used to determine the first pose change information. Alternatively, the first pose change information can be determined based on time points t1 to t2. N By calculating the third relative pose corresponding to each of the N time points, we can determine the N first relative poses of the target device during the acquisition of the first to Nth video frames, relative to the target device during the acquisition of the first video frame. In other words, we can calculate the target device's pose changes at times t1 to t2. N N relative poses relative to the target device at time t1, where each pose changes at time t1; where for any i-th environmental image, any i-th video frame, or time t... i The corresponding i-th first relative pose may include: when the target device acquires the video frame corresponding to the i-th environmental image (i.e., time t) i The rotation angle relative to the target device when it acquires the video frame corresponding to the first environmental image (i.e., time t1), in the x, y, and z axes of three-dimensional space; and / or, the rotation angle of the target device at time t... i The displacement relative to the target device at time t1 along the x, y, and z axes in three-dimensional space. Furthermore, the time interval from t1 to t... N The first relative pose at each of the N time points is determined as the first pose change information of the target device during the acquisition of the target video.

[0038] The process of predicting the second pose change information corresponding to the target device is described below.

[0039] Taking a target sensor including an accelerometer and a gyroscope as an example; the accelerometer can measure the linear acceleration of the target device along the x, y, and z axes in three-dimensional space, while the gyroscope can measure the rotational speed of the target device along the x, y, and z axes. By integrating the data measured by the accelerometer and gyroscope respectively in the target sensor data, the pose changes of the target device relative to the initial state (the position and attitude of the target device when acquiring the first video frame) can be estimated when acquiring N video frames. Therefore, based on the estimated position and attitude changes, the second pose change information of the target device during the acquisition of the target video can be determined.

[0040] Continuing with the previous example, let's denote the acquisition time of the i-th video frame as time t. i Let time t be the time when the (i-1)th video frame is captured. i-1 Then, based on the target sensor data, the acquisition time is located at time t. i-1 ~t i Using sensor data from time t, estimate the target device at time t i Below the target device at time t i-1 The i-th fourth relative pose S that undergoes pose change 2i That is, calculate time t i The corresponding fourth relative pose S 2i In the aforementioned example, the fourth relative pose S 2i This may include: rotation information R 2i and / or displacement vector T 2i Rotation information R 2i Used to indicate the target device at time t i Below the target device at time t i-1 Below, the rotation angle occurring along the x, y, and z axes in three-dimensional space; displacement vector T 2i Used to indicate the target device at time t i Below the target device at time t i-1 The distance of displacement along the x, y, and z axes in three-dimensional space.

[0041] The time intervals t1 to t2 are obtained through the process described in the previous example. N After their respective fourth relative poses, and in accordance with the example described above based on time t1~t N The process of determining the first pose change information based on the third relative pose at N time points is similar, and can be based on time t1~t2. N The second pose change information of the target device is determined by the fourth relative pose corresponding to each of the N time points.

[0042] Various matching algorithms can be employed to match the predicted first pose change information with the second pose change information. For example, the first pose change information and the second pose change information can be converted into a first embedding vector and a second embedding vector, respectively. Then, a first similarity between the first embedding vector and the second embedding vector can be calculated. If the first similarity is less than a first preset threshold, it indicates that the first pose change information and the second pose change information do not match. As another example, when the first pose change information includes N first relative poses and the second pose change information includes N second relative poses, the i-th similarity between the i-th first relative pose and the i-th second relative pose can be calculated. Then, based on the calculated N similarities, the first similarity between the first pose change information and the second pose change information can be calculated. If the first similarity is less than a first preset threshold, it is determined that the first pose change information and the second pose change information do not match.

[0043] In one possible implementation, a first encoder can be used to process an environmental image sequence to obtain an environmental change representation vector; a second encoder can be used to process target sensor data to obtain a pose change representation vector. The first encoder and the second encoder are machine learning models trained through multi-modal contrastive learning; a second similarity between the environmental change representation vector and the pose change representation vector is calculated; and if the second similarity is less than a second preset threshold, the target video is determined to be a fake video.

[0044] Multimodal contrastive learning is a self-supervised learning method that aims to learn meaningful feature representations by comparing samples of different modalities in a data space. The core idea is to optimize by minimizing the loss function. Through one or more rounds of updating the parameters of the first and second encoders, the goal is to maximize the similarity between related samples (or reduce their distance) while minimizing the similarity between unrelated samples (or increasing their distance). The loss function used during training is typically defined as minimizing the negative log-likelihood or maximizing the cross-entropy loss function. Based on this, the first and second encoders can be trained on a first sample pair (positive samples) and a second sample pair (negative samples). The first sample pair includes a first environmental image sequence and first sensor data, with the first environmental image sequence obtained from a first video, and the first video and first sensor data collected by the same device within the same time period. The second sample pair includes a second environmental image sequence and second sensor data that are not related.

[0045] It should be noted that the lack of correlation between the second environmental image sequence and the second sensor data may include: the second sensor data and the second video corresponding to the second environmental image are not collected by the same device; or, the second sensor data and the second video corresponding to the second environmental image are collected by the same device at different times.

[0046] The first sensor data in the first sample pair can also be replaced with pose change information of the relevant device predicted based on the first sensor data; similarly, the second sensor data in the second sample pair can also be replaced with pose change information of the relevant device predicted based on the second sensor data. In this case, referring to FIG3, the second encoder can also be used to process the second pose change information in the aforementioned embodiment to obtain the pose change representation vector.

[0047] In conjunction with the various implementation methods described above, it can be determined that the detected target video is not a fake video, i.e., the target video is a real video collected by the relevant target device, when it is determined that the first pose change information and the second pose change information match, and the environmental change representation vector and the pose change representation vector are similar.

[0048] Based on the same concept as the aforementioned method embodiments, this specification also provides a video detection device 400. Referring to FIG4, the device 400 includes: a video processing unit 401 configured to obtain an environmental image sequence based on a video frame sequence corresponding to a target video; a data acquisition unit 403 configured to acquire target sensor data corresponding to the target video, wherein the target sensor data is provided by a target sensor configured in a target device that acquires the target video, and the target sensor is used to sense the motion of the target device; and a detection processing unit 405 configured to determine whether the target video is a fake video based on the environmental image sequence and the target sensor data.

[0049] This specification also provides a computer-readable storage medium storing a computer program / instruction thereon, which, when executed in a computer, enables the computer to perform the video detection methods provided in the foregoing embodiments.

[0050] This specification also provides a computing device in the embodiments, including a memory and a processor. The memory stores computer programs / instructions, and when the processor executes the computer programs / instructions, it implements the video detection methods provided in the foregoing embodiments.

[0051] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0052] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0053] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0054] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes the elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.

[0055] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0056] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0057] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0058] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0059] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0060] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0061] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0062] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0063] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0064] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0065] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.

Claims

1. A video detection method, the method comprising: Based on the video frame sequence corresponding to the target video, obtain the environmental image sequence; Acquire target sensor data corresponding to the target video, wherein the target sensor data is provided by a target sensor configured in the target device that acquires the target video, and the target sensor is used to sense the motion of the target device; Based on the environmental image sequence and the target sensor data, determine whether the target video is a fake video.

2. The method of claim 1, wherein obtaining the sequence of environment images based on the sequence of video frames corresponding to the target video comprises: For the video frame sequence corresponding to the target video, remove the human image regions in the video frames to obtain the environmental image sequence.

3. The method according to claim 1, wherein determining whether the target video is a fake video based on the environmental image sequence and the target sensor data includes: Based on the environmental image sequence, predict the first pose change information corresponding to the target device; Based on the target sensor data, predict the second pose change information corresponding to the target device; Determine whether the first pose change information matches the second pose change information; If not, then the target video is determined to be a fake video.

4. The method according to claim 3, wherein the first pose change information includes the i-th first relative pose of the target device when acquiring any i-th video frame, relative to the target device when acquiring the first video frame; and the second pose change information includes the i-th second relative pose of the target device when acquiring any i-th video frame, relative to the target device when acquiring the first video frame. wherein, Determining whether the first pose change information and the second pose change information match includes: calculating the i-th similarity between the i-th first relative pose and the i-th second relative pose, and calculating the first similarity between the first pose change information and the second pose change information based on the i-th similarity. If the first similarity is less than a first preset threshold, it is determined that the first pose change information and the second pose change information do not match.

5. The method according to claim 1, wherein determining whether the target video is a fake video based on the environmental image sequence and the target sensor data includes: The environmental image sequence is processed using a first encoder to obtain an environmental change representation vector; The target sensor data is processed using a second encoder to obtain a pose change representation vector, wherein the first encoder and the second encoder are machine learning models trained through multi-modal contrastive learning. Calculate the second similarity between the environmental change representation vector and the pose change representation vector; If the second similarity is less than the second preset threshold, the target video is determined to be a fake video.

6. The method according to claim 5, wherein the first encoder and the second encoder are trained based on a first sample pair and a second sample pair; the first sample pair includes a first environmental image sequence and first sensor data, the first environmental image sequence being obtained based on a first video, and the first video and the first sensor data being collected by the same device within the same time period; the second sample pair includes a second environmental image sequence and second sensor data that are not related.

7. The method of claim 6, the second environmental image is obtained based on a second video; wherein, The second video and the second sensor data are collected by different devices, or the second video and the second sensor data are collected by the same device at different time periods.

8. The method according to claim 1, wherein the target video is used to support biometric identification of the target user.

9. The method according to any one of claims 1-8, wherein the target sensor is at least one of the following types of sensors: accelerometer, gyroscope, and magnetometer.

10. A video detection device, the device comprising: The video processing unit is configured to obtain an environmental image sequence based on the video frame sequence corresponding to the target video. The data acquisition unit is configured to acquire target sensor data corresponding to the target video. The target sensor data is provided by a target sensor configured in the target device that acquires the target video. The target sensor is used to sense the motion of the target device. The detection processing unit is configured to determine whether the target video is a fake video based on the environmental image sequence and the target sensor data.

11. A computing device comprising a memory and a processor, wherein the memory stores a computer program / instructions, and the processor, when executing the computer program / instructions, implements the method of any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computing device, the computing device performs the method of any one of claims 1-9.