A method of counterfeiting detection of a synthetic video image

By analyzing the edge blurring and depth range of video images, and combining the human shadow-light intensity analysis model, the errors and loopholes in synthetic video identification were solved, and more accurate video identification was achieved.

CN116994120BActive Publication Date: 2025-11-25CHINA ACADEMY OF INFORMATION & COMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310971368.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-11-25
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

Existing technologies for identifying synthetic videos suffer from errors caused by frame-by-frame image manipulation and vulnerabilities due to a lack of content analysis, making it difficult for traditional methods to accurately identify synthetic videos.

Method used

By preprocessing video image data, extracting and analyzing the blurriness of target edge outlines and the average depth range, and constructing a human shadow-light intensity analysis model, combined with multiple data analyses, identification results are generated.

Benefits of technology

It improves the accuracy and reliability of video identification, can accurately identify spliced ​​and modified areas in videos, and enhances the identification logic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994120B_ABST
    Figure CN116994120B_ABST
Patent Text Reader

Abstract

The application discloses a kind of synthetic video image's method of identifying forgery, it is related to synthetic video identification technical field, specific steps are as follows: the sample is preprocessed, sample data is obtained, the sample data includes first image data and second image data;According to first image data, the first analysis label of sample is labeled by analysis processing;Second image data is analyzed and processed, and second analysis label is labeled in sample;The first analysis label in sample is integrated with second analysis label, and identification result is generated;The present application constructs human shadow-light intensity analysis model, uses the means of reflectivity estimation to analyze the illumination intensity of single-frame video image in sample, and the human shadow angle is matched by the illumination intensity analyzed, and the actual human shadow angle is compared with model human shadow angle, to realize the analysis of second data, strengthen the reliability and accuracy of identification data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of synthetic video identification, in particular to a synthetic video image identification method. BACKGROUND

[0002] With the rise of short videos, more and more video software has the function of frame-by-frame P picture, and some people use this means to forge and splice part of the video, thereby affecting the judgment of video viewers and video identifiers.

[0003] After searching, the contrast file with the publication number CN115662109A proposes a timing method, device and system for verifying the time credibility of multi-source fusion video. The contrast file controls the LED lamp beads in the minute timing display device, second timing display device and millisecond timing display device to light up according to the set program, so as to achieve the effects of precise timing and providing time anchor points for multi-source video. By analyzing the display screen of the multi-source video shooting timing device frame by frame, the frame interval time of the multi-source video can be determined, and then the frame rate of the multi-source video can be determined, so as to achieve the purpose of judging the time credibility of the multi-source video. The timing device has high display accuracy, wide visual range, convenient operation and low cost.

[0004] Referring to the contrast file, it is found that the prior art still has the following deficiencies:

[0005] 1. In the identification process of synthetic video, there are cases of frame-by-frame P picture. Based on such cases, the traditional video identification method may have errors in result analysis.

[0006] 2. In the process of video identification, video basic data is mostly used to determine the video, and the analysis is not based on the content of the video, so there is a loophole in the identification.

[0007] In order to solve the above-mentioned problems, a synthetic video image identification method is proposed. SUMMARY

[0008] The purpose of the present application is to provide a synthetic video image identification method to solve the problems in the background art.

[0009] In order to achieve the above-mentioned purpose, the present application provides the following technical scheme:

[0010] The synthetic video image identification method comprises the following steps:

[0011] The identification sample is pretreated to obtain sample data, and the sample data includes first image data and second image data;

[0012] According to the first image data, the identification sample is analyzed and processed, and the first analysis label is labeled.

[0013] analyzing the second image data to mark a second analysis label in the identification sample;

[0014] integrating the first analysis label and the second analysis label in the identification sample to generate an identification result.

[0015] In a preferred embodiment, the identification sample is divided into L identification intervals according to the frame number, where L = {0, 1, 2, 3, 4,..., n}, n is an integer greater than 1, and each identification interval contains 1 video frame.

[0016] In a preferred embodiment, the first image data includes an analysis target edge line blur degree and a depth of field range average value, the analysis target edge line blur degree is the clarity of the portrait edge line of the identification portrait, the smaller the analysis target edge line blur degree, the greater the suspicion of P image splicing, the depth of field range value difference is the average value between the maximum depth of field range value and the minimum depth of field range value in the image or video that is considered clear, the depth of field range represents the distance range from the foreground to the background in the image, that is, the visual range between the foreground and the background that can maintain clarity at the same time, the greater the depth of field range value average, the more objects in the image can be clearly presented, including the foreground and the background, the greater the synthesis probability, the smaller the depth of field range value average, indicating that only a few objects in the image are clearly displayed, and the synthesis probability is also large.

[0017] The second image data includes a shadow angle and an image illumination intensity, the shadow angle represents the angle ratio between the identification target and its shadow, and the image illumination intensity is the image illumination intensity calculated by reflectivity, the closer the closeness of the correlation between the shadow angle and the image illumination intensity, the lower the synthesis degree of the video, and the closer the closeness of the correlation, the higher the synthesis degree of the video.

[0018] In a preferred embodiment, the first image data analysis processing step is:

[0019] Obtain the video image of the nth identification interval in the L identification intervals, obtain the analysis target edge line blur degree x and the depth of field range average value s in the video image, obtain the first operation result by taking the logarithm of the square root of the analysis target edge line blur degree x, perform ratio operation on the first operation result and the depth of field range average value s, and multiply the ratio operation result by the blur degree correction constant P to obtain the first image coefficient ζ, the greater or smaller the first image coefficient ζ, the lower the authenticity of the video image, and the higher the image processing degree in the video image.

[0020] In a preferred embodiment, the first analysis label comprises a first true label and a first modified label, and the labeling step of the first analysis label is:

[0021] A first analysis label comparison threshold ki1 and ki2 are set, wherein the first analysis label comparison threshold ki1 is greater than the first analysis label comparison threshold ki2, and the first image coefficient ζ is substituted into the first analysis label comparison threshold ki1 and ki2 for comparison;

[0022] If the first image coefficient ζ is greater than 0 and less than the first analysis label comparison threshold ki2, the first modified label is labeled for the identification sample;

[0023] If the first image coefficient ζ is greater than the first analysis label comparison threshold ki2 and less than the first analysis label comparison threshold ki1, the first true label is labeled for the identification sample;

[0024] If the first image coefficient ζ is greater than the first analysis label comparison threshold ki1, the first modified label is labeled for the identification sample;

[0025] The first true label has a lower video image modification degree for the identification sample than the first false label.

[0026] In a preferred embodiment, the second image data analysis processing step is:

[0027] The angle a of the shadow in the video image and the image illumination intensity g are obtained, a shadow-illumination analysis model is constructed, the image illumination intensity g is substituted into the shadow-illumination analysis model for analysis processing, and a model shadow angle range a1-a2 is output;

[0028] The specific steps of constructing the shadow-illumination analysis model are as follows:

[0029] Data collection: collect shadow images under different lighting conditions and corresponding illumination intensity data for training and testing the model;

[0030] Feature extraction: extract useful features from the shadow images, such as leg key points, and illumination intensity information;

[0031] Reflectivity estimation: use existing illumination intensity data and feature information in the shadow image to construct a reflectivity estimation model for estimating the illumination intensity of each image.

[0032] Data processing: combine the human key points and illumination intensity data into sample data for training and testing the model.

[0033] Model training: use regression algorithms or classification algorithms to train the shadow-illumination analysis model to accurately predict the shadow angle corresponding to the illumination intensity;

[0034] Model output: converting the shadow-intensity relationship trained by the model into a predicted value of the shadow angle and a predicted value of the light intensity;

[0035] Specific range setting: setting a specific shadow angle range value for determining whether the detected shadow angle is within the range;

[0036] Prediction and comparison: inputting the detected shadow angle into the model to obtain the predicted light intensity, and comparing it with the specific shadow angle range value;

[0037] Result output: determining whether the shadow angle is within the specific range according to the model output result, and obtaining the final labeling result.

[0038] In a preferred embodiment, the second analysis label includes a second true label and a second modified label, and the labeling step of the second analysis label is:

[0039] Substitute the shadow angle a into the output model shadow angle range a1-a2 output by the shadow-intensity analysis model for analysis;

[0040] If the shadow angle a is in the model shadow angle range a1-a2, the second true label is labeled for the identification sample;

[0041] If the shadow angle a is not in the model shadow angle range a1-a2, a secondary determination is made;

[0042] The steps of the secondary determination are:

[0043] Add a correction number T to the shadow-intensity analysis model, adjust the model shadow angle range a1-a2 to a1-T-a2+T, and substitute the shadow angle a into the adjusted model shadow angle range a1-T-a2+T for analysis;

[0044] If the shadow angle a is in the model shadow angle range a1-T-a2+T, the second true label is labeled for the identification sample;

[0045] If the shadow angle a is not in the model shadow angle range a1-T-a2+T, the second modified label is labeled for the identification sample;

[0046] The authenticity of the identification sample of the second true label is greater than that of the second modified label.

[0047] In a preferred embodiment, the process of generating the identification result is:

[0048] Integrating and analyzing the first analysis label and the second analysis label in the identification sample;

[0049] statistically identify the number D of modified intervals in the sample that simultaneously have the first and second modification labels, wherein the number D is greater than or equal to 0 and less than or equal to the total number L of identification intervals;

[0050] setting a modified interval comparison threshold DL1 and DL2, wherein 0≤DL1<DL2≤L, and substituting the number D into the modified interval comparison threshold DL1 and DL2 for comparison analysis;

[0051] if the number D is greater than or equal to 0 and less than the modified interval comparison threshold DL1, generating a true result for the identification sample;

[0052] if the number D is greater than or equal to the modified interval comparison threshold DL1 and less than the modified interval comparison threshold DL2, generating a doubtful result for the identification sample;

[0053] if the number D is greater than or equal to the modified interval comparison threshold DL2 and less than or equal to the total number L of identification intervals, generating a modified result for the identification sample;

[0054] the greater the degree of modification of the modified result compared to the doubtful result, and so on.

[0055] In the above technical solution, the present application provides technical effects and advantages:

[0056] 1. The present application can help determine the foreground and background elements in the image or video by considering the target edge outline and the average depth of field range, which helps target detection, tracking and other image processing tasks, and uses two types of data that are difficult to think of, making it easier to identify the authenticity of the video.

[0057] 2. The present application constructs a human shadow-light intensity analysis model, uses reflectivity estimation to analyze the illumination intensity of single-frame video images in the identification sample, matches the human shadow angle through the analyzed illumination intensity, and compares the actual human shadow angle with the model human shadow angle, thereby realizing the analysis of the second data and strengthening the reliability and accuracy of the identification data.

[0058] 3. The present application accurately obtains the number of video image intervals in the identification sample that have splicing, modification and synthesis through superimposed analysis of multiple data, and performs secondary determination according to the number of video image intervals, further analyzes and determines the identification sample, increases the determination logic, and improves the accuracy of video identification. BRIEF DESCRIPTION OF DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0060] Figure 1 A flowchart of a synthetic video image authentication method according to an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will combine the drawings in the embodiments of the present application to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0062] Please refer to Figure 1 The synthetic video image authentication method according to an embodiment of the present application is shown in the figure. The synthetic video image authentication method includes the following steps:

[0063] Pretreat the identification sample to obtain sample data, wherein the sample data includes first image data and second image data;

[0064] Analyze and process the first image data, and label the identification sample with a first analysis label;

[0065] Analyze and process the second image data, and label the identification sample with a second analysis label;

[0066] Integrate the first analysis label and the second analysis label in the identification sample, and generate an identification result.

[0067] Divide the identification sample into L identification intervals according to the number of frames, wherein L={0, 1, 2, 3, 4,..., n}, n is an integer greater than 1, and each identification interval contains 1 video frame.

[0068] It should be noted that the pretreatment step and the technology involved are:

[0069] Video reading: read video frame data from a video file, which can use common image / video processing libraries such as OpenCV, etc.

[0070] Video frame extraction: split the video into multiple frames of images.

[0071] Image preprocessing: Preprocess each frame of image, including denoising, enhancement, cropping, scaling, etc. to improve the effect of subsequent analysis.

[0072] Feature extraction: Extract key features from each frame of image for subsequent analysis and processing. The data extracted in this invention is specifically sample data.

[0073] Analysis algorithm: Select appropriate analysis algorithm according to specific task, such as target detection, object tracking, action recognition, etc.

[0074] Data storage: Store the video frame data after frame processing and the extracted features, etc. for subsequent comprehensive analysis.

[0075] The invention is based on the processing in the video authentication scheme with human image;

[0076] The first image data includes the average value of the analysis target edge line blur and the depth of field range. The analysis target edge line blur is the clarity of the edge line of the human image in the human image authentication. The smaller the analysis target edge line blur, the greater the suspicion of P image splicing. The depth of field range value difference refers to the average value between the maximum depth of field range and the minimum depth of field range in the image or video that is considered clear. The depth of field range represents the distance range from the foreground to the background in the image, that is, the visual range between the foreground and the background that can maintain clarity at the same time. The larger the average value of the depth of field range, the more clearly multiple objects in the image are presented, including the foreground and the background, and the greater the synthesis probability. The smaller the average value of the depth of field range, the fewer objects in the image are clearly displayed, and the synthesis probability is also large.

[0077] It should be noted that: the larger the average value of the depth of field range, the more likely it is a synthesized image or a computer generated image, because in the real world, the depth of field is usually limited and cannot display distant and near objects clearly at the same time. Therefore, an excessively large depth of field range may make the image look too realistic, but it is actually synthesized or computer generated;

[0078] The smaller the average value of the depth of field range, the presentation of the object will be blurred or unclear, which may be a sign of editing using a blur effect or image processing software. An excessively small depth of field range may make the image look too processed, but it is actually edited or modified.

[0079] In addition, the technologies involved in the first image data extraction are as follows:

[0080] Edge line blur extraction: Use edge detection algorithms or edge blur calculation, such as Canny algorithm, Sobel algorithm, Laplacian algorithm, etc. to detect the edge line blur in the image.

[0081] Depth range mean extraction: Use depth estimation techniques in computer vision, such as structured light, stereo matching, deep learning, etc., to estimate the depth information of objects in the image.

[0082] The second image data includes a shadow angle and an image illumination intensity, the shadow angle represents the presented angle ratio between the identification target and its own shadow, and the image illumination intensity is an image illumination intensity calculated by reflectivity. The tightness of the correlation between the shadow angle and the image illumination intensity is used to determine the synthesis degree of the video. The tighter the correlation, the lower the synthesis degree of the video, and the looser the correlation, the higher the synthesis degree of the video.

[0083] It should be noted that:

[0084] The shadow angle uses a human body detection and key point positioning method with the leg as the key point. Specifically, the following steps can be followed:

[0085] Human body detection: First, use a deep learning model to detect the human body in the video frame image to find the position and posture information of the person. Common deep learning models include SSD (Single Shot Multibox Detector) and YOLO (You Only Look Once) etc.

[0086] Key point positioning: Use a key point positioning model such as OpenPose to position the key points of the detected person. Pay special attention to the leg key points, which generally include the connection between the thigh and the calf.

[0087] Shadow detection: In the video frame image, use image processing and segmentation techniques to detect and segment the position and shape of the shadow;

[0088] Angle calculation: With the position, key points and shadow position information of the person, the angle between the leg key points and the shadow can be calculated using geometric and trigonometric knowledge. The angle information can be obtained by calculating the included angle between the connecting line of the leg key points and the horizontal line or the included angle between the vertical line.

[0089] Angle correction: In some cases, the angle may need to be corrected to take into account the tilt of the camera or perspective transformation, etc.

[0090] The image illumination intensity is obtained by reflectivity estimation, specifically:

[0091] The reflectivity estimation is a method for estimating the illumination intensity in the image by the relationship between the reflectivity of the object and the illumination intensity. The reflectivity refers to the ability of the object surface to reflect light, and different materials of the object will have different reflectivity.

[0092] A specific reflectance estimation method can be implemented by the following steps:

[0093] Obtain a reference image: First, a reference image is needed, which is taken under known conditions, including known light intensity and camera parameters, to serve as a reference for reflectance;

[0094] Extract the object region: According to the object that needs to estimate the reflectance, use image segmentation or object detection algorithm to extract the object region.

[0095] Estimate reflectance: For the extracted object region, find the corresponding region in the reference image and obtain the pixel value of the region. Assume that the pixel value in the reference image is I_ref, and the pixel value in the image to be estimated is I_test;

[0096] Reflectance calculation: According to the definition of reflectance, reflectance R = I_test / I_ref. By calculating the reflectance of each pixel point, the reflectance of the object surface in the image to be estimated can be obtained.

[0097] Reflectance estimation is a relative estimation method, which needs to obtain a reference image in advance, and assumes that the light intensity and camera parameters are the same in the reference image and the image to be estimated. In order to obtain more accurate reflectance estimation, more complex techniques such as multi-view imaging or reflectance model fitting may be needed.

[0098] The steps of the first image data analysis processing are:

[0099] Obtain the video image of the nth identification interval in the L identification intervals, obtain the analysis target edge sketching blur x and the depth of field range average s in the video image, obtain the first operation result by taking the logarithm of the square root of the analysis target edge sketching blur x, and obtain the first image coefficient ζ by performing ratio operation on the first operation result and the depth of field range average s and multiplying by the blur correction constant P, the acquisition formula of the first image coefficient ζ is:

[0100] (P>0 and x is always greater than 0);

[0101] The greater or smaller the first image coefficient ζ is, the lower the authenticity of the video image will be, and the higher the image processing degree in the video image will be.

[0102] The first analysis label includes a first real label and a first modified label, and the labeling step of the first analysis label is:

[0103] A first analysis label comparison threshold ki1 and ki2 are set, wherein the first analysis label comparison threshold ki1 is greater than the first analysis label comparison threshold ki2, and the first image coefficient ζ is substituted into the first analysis label comparison threshold ki1 and ki2 for comparison;

[0104] If the first image coefficient ζ is greater than 0 and less than the first analysis label comparison threshold ki2, the first modified label is marked for the identification sample;

[0105] If the first image coefficient ζ is greater than the first analysis label comparison threshold ki2 and less than the first analysis label comparison threshold ki1, the first true label is marked for the identification sample;

[0106] If the first image coefficient ζ is greater than the first analysis label comparison threshold ki1, the first modified label is marked for the identification sample;

[0107] The first true label has a lower video image modification degree for the identification sample than the first false label.

[0108] The second image data analysis processing steps are:

[0109] The angle a of the shadow in the video image and the image illumination intensity g are obtained, a shadow-illumination intensity analysis model is constructed, the image illumination intensity g is substituted into the shadow-illumination intensity analysis model for analysis processing, and the model shadow angle range a1-a2 is output;

[0110] The specific steps of constructing the shadow-illumination intensity analysis model are as follows:

[0111] Data collection: collect shadow images and corresponding illumination intensity data under different illumination conditions for training and testing the model;

[0112] Feature extraction: extract useful features from the shadow images, such as leg key points, and illumination intensity information;

[0113] Reflectivity estimation: use existing illumination intensity data and feature information in the shadow image to construct a reflectivity estimation model for estimating the illumination intensity of each image.

[0114] Data processing: combine the human key points and illumination intensity data into sample data for training and testing the model.

[0115] Model training: use regression algorithms or classification algorithms to train the shadow-illumination intensity analysis model to accurately predict the shadow angle corresponding to the illumination intensity;

[0116] Model output: convert the shadow-illumination intensity relationship obtained by model training into the predicted value of the shadow angle and the predicted value of the illumination intensity;

[0117] Specific range setting: set a specific range of shadow angle value to determine whether the detected shadow angle is within the range;

[0118] Prediction and comparison: input the detected shadow angle into the model to obtain the predicted light intensity, and compare it with the specific shadow angle range value;

[0119] Result output: determine whether the shadow angle is within the specific range according to the output result of the model, and obtain the final labeling result.

[0120] It should be noted that the greater the image light intensity, the smaller the angle value between the person and the shadow in general. In the case of high light intensity, the shadow of the person will be clear and short, and the angle between the person and the shadow will be relatively small. In the case of low light intensity, the shadow will be relatively blurred and extended, and the angle between the person and the shadow will be relatively large.

[0121] The above-mentioned shadow-light intensity analysis model can be solved by a supervised learning regression model. In training the shadow-light intensity analysis model, a labeled data set containing shadow images, light intensity and corresponding shadow angle can be used. Through training, the shadow-light intensity analysis model can learn the relationship between light intensity and shadow angle, and then use this relationship to predict the shadow angle in the image to be detected.

[0122] The second analysis label includes a second true label and a second modified label, and the labeling step of the second analysis label is:

[0123] Substitute the shadow angle a into the output model shadow angle range a1-a2 output by the shadow-light intensity analysis model for analysis;

[0124] If the shadow angle a is in the model shadow angle range a1-a2, the second true label is labeled for the identification sample;

[0125] If the shadow angle a is not in the model shadow angle range a1-a2, secondary determination is performed;

[0126] The steps of the secondary determination are:

[0127] Add a modified constant T to the shadow-light intensity analysis model, adjust the model shadow angle range a1-a2 to a1-T-a2+T, and substitute the shadow angle a into the adjusted model shadow angle range a1-T-a2+T for analysis;

[0128] If the shadow angle a is in the model shadow angle range a1-T-a2+T, the second true label is labeled for the identification sample;

[0129] If the human shadow angle a is not in the model human shadow angle range a1-T~a2+T, the second modified label of the identification sample is marked;

[0130] The identification sample authenticity of the second true label is greater than the identification sample authenticity of the second modified label.

[0131] The flow of the identification result generation is:

[0132] The first analysis label and the second analysis label in the identification sample are integrated and analyzed;

[0133] The number D of the modified interval in which the first modified label and the second modified label coexist in the identification sample is counted, and the number D is greater than or equal to 0 and less than or equal to the total number L of identification intervals;

[0134] The modified interval comparison threshold DL1 and DL2 are set, wherein 0<=DL1<DL2<=L, and the number D is substituted into the modified interval comparison threshold DL1 and DL2 for comparison analysis;

[0135] If the number D is greater than or equal to 0 and less than the modified interval comparison threshold DL1, the true result of the identification sample is generated;

[0136] If the number D is greater than or equal to the modified interval comparison threshold DL1 and less than the modified interval comparison threshold DL2, the doubtful result of the identification sample is generated;

[0137] If the number D is greater than or equal to the modified interval comparison threshold DL2 and less than or equal to the total number L of identification intervals, the modified result of the identification sample is generated;

[0138] The greater the modified degree of the modified result compared with the doubtful result, and so on.

[0139] The present application can help to determine the foreground and background elements in the image or video by considering the target edge outline and the depth of field range mean and other conditions, which is helpful for target detection, tracking and other image processing tasks, and the two types of data are difficult to think of, which is more easy to identify the authenticity of the video; and by constructing a human shadow-light intensity analysis model, the reflectivity estimation method is used to analyze the illumination intensity of the single-frame video image in the identification sample, and the actual human shadow angle is matched with the model human shadow angle through the analyzed illumination intensity, so as to realize the analysis of the second data, and the reliability and accuracy of the identification data are strengthened.

[0140] In addition, by superimposed analysis of multiple data, the number of video image intervals with splicing, modification and synthesis in the identification sample is accurately obtained, and secondary determination is carried out according to the number of video image intervals, so as to further analyze and determine the identification sample, increase the determination logic, and improve the accuracy of video identification.

[0141] The above formulas are dimensionless values, and the formulas are obtained by collecting a large amount of data to simulate a formula of the nearest real situation, and the preset parameters in the formula are set by the person skilled in the art according to the actual situation.

[0142] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, and indicates that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist simultaneously, and B exists alone, wherein A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects, but can also represent an "and / or" relationship, which can be understood according to the context before and after.

[0143] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, wherein a, b, and c can be single or multiple.

[0144] It should be understood that in various embodiments of the present application, the size of the sequence number of the above processes does not mean the order of execution, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0146] The units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, which can be located in one place or distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the present embodiment.

[0147] In addition, each functioning described in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0148] The above descriptions are merely specific embodiments of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for detecting fake synthetic video images, characterized in that, The method includes the following steps: The identification sample is preprocessed to obtain sample data, which includes first image data and second image data. The first image data is analyzed and processed to label the identification sample with the first analysis label; The steps for analyzing and processing the first image data are as follows: The video image of the nth identification interval among L identification intervals is obtained. The blurriness x of the edge outline of the analysis target and the mean depth range s of the analysis target are obtained in the video image. The first calculation result is obtained by taking the square root of the blurriness x of the edge outline of the analysis target and performing a logarithmic operation. The first calculation result is compared with the mean depth range s and multiplied by the blurriness correction constant P to obtain the first image coefficient ζ. The larger or smaller the first image coefficient ζ is, the lower the authenticity of the video image will be, and the higher the degree of image processing in the video image. The second image data is analyzed and processed, and a second analysis label is added to the identification sample; The steps for the second image data analysis and processing are as follows: Obtain the human shadow angle 'a' and the image illumination intensity 'g' in the video image, construct a human shadow-light intensity analysis model, substitute the image illumination intensity 'g' into the human shadow-light intensity analysis model for analysis and processing, and output the human shadow angle range 'a1' to 'a2' of the model; The specific steps for constructing the human shadow-light intensity analysis model are as follows: Data acquisition: Collect images of human shadows under different lighting conditions and corresponding light intensity data for training and testing models; Feature extraction: Extracting useful features from human silhouette images, including key points on the legs and information on illumination intensity; Reflectance estimation: Using existing illumination intensity data and feature information in human shadow images, a reflectance estimation model is constructed to estimate the illumination intensity of each image; Data processing: Combine human body key points and light intensity data into sample data for training and testing models; Model training: The human shadow-light intensity analysis model is trained using regression or classification algorithms so that it can accurately predict the angle of human shadows corresponding to light intensity. Model output: Convert the shadow-light intensity relationship obtained from model training into predicted values ​​of shadow angle and light intensity; Specific range setting: Set a specific range of human shadow angle values ​​to determine whether the detected human shadow angle is within this range; Prediction and Comparison: Input the detected shadow angle into the model to obtain the predicted light intensity, and compare it with the value of a specific shadow angle range; Output results: Based on the model output, determine whether the angle of the human figure is within a specific range, and obtain the final annotation result; The first and second analytical labels within the identified sample are integrated to generate the identification results.

2. The method for detecting fake synthetic video images according to claim 1, characterized in that, The identification samples are divided into L identification intervals based on the number of frames, where L = {0, 1, 2, 3, 4.....n}, and n is an integer greater than 1. Each identification interval contains 1 video frame.

3. The method for detecting fake synthetic video images according to claim 2, characterized in that, The first image data includes the blurriness of the target edge outline and the average depth range. The blurriness of the target edge outline is used to identify the clarity of the edge outline of the portrait. The smaller the blurriness of the target edge outline, the greater the suspicion of image manipulation. The depth range difference refers to the average between the maximum and minimum depth range values ​​considered to be clear in the image or video. The depth range represents the distance range from the foreground to the background in the image, that is, the visual range between the foreground and the background that can maintain clarity at the same time. The larger the average depth range value, the more objects in the image can be clearly presented, including the foreground and the background, and the greater the probability of compositing. The smaller the average depth range value, the more only a few objects are clearly displayed in the image, and the probability of compositing is also relatively high. The second image data includes the shadow angle and the image illumination intensity. The shadow angle represents the angle ratio between the target and its own shadow. The image illumination intensity is the image illumination intensity calculated by reflectance. The degree of correlation between the shadow angle and the image illumination intensity is considered. The greater the degree of correlation, the lower the degree of video synthesis. The smaller the degree of correlation, the higher the degree of video synthesis.

4. The method for detecting fake composite video images according to claim 3, characterized in that, The first analysis label includes a first real label and a first modified label. The annotation steps for the first analysis label are as follows: Set first analysis label comparison thresholds ki1 and ki2, wherein the first analysis label comparison threshold ki1 is greater than the first analysis label comparison threshold ki2, and substitute the first image coefficient ζ into the first analysis label comparison thresholds ki1 and ki2 for comparison; If the first image coefficient ζ is greater than 0 and less than the first analysis label comparison threshold ki2, the identification sample is labeled with the first modified label; If the first image coefficient ζ is greater than the first analysis label comparison threshold ki2 and less than the first analysis label comparison threshold ki1, the identification sample is labeled with the first true label; If the first image coefficient ζ is greater than the first analysis label comparison threshold ki1, the identification sample is labeled with the first modified label; The more genuine the label, the less modified the video image of the sample it identifies compared to the first fake label.

5. The method for detecting fake synthetic video images according to claim 4, characterized in that, The second analysis label includes a second true label and a second modified label. The annotation steps for the second analysis label are as follows: Substitute the shadow angle 'a' into the output model of the shadow-light intensity analysis model within the shadow angle range a1 to a2 for analysis; If the angle 'a' of the human shadow is within the range 'a1' to 'a2' of the model's human shadow angle, the identification sample is labeled with a second true label. If the angle 'a' of the shadow is not within the range 'a1' to 'a2' of the shadow angle in the model, then a second determination is made; The steps for the secondary determination are as follows: Add a correction constant T to the human shadow-light intensity analysis model, and adjust the human shadow angle range a1~a2 in the model to a1-T~a2+T. Substitute the human shadow angle a into the adjusted human shadow angle range a1-T~a2+T in the analysis. If the angle 'a' of the human shadow is within the range of a1-T to a2+T of the model's human shadow angle, the identification sample is labeled with a second true label. If the angle 'a' of the human shadow is not within the range of a1-T to a2+T of the model's human shadow angle, the second modification label is applied to the identified sample. The authenticity of the sample identified by the second genuine label is greater than that of the sample identified by the second modified label.

6. The method for detecting fake composite video images according to claim 5, characterized in that, The process for generating the identification results is as follows: Integrate the analysis of the first and second analytical labels in the identified samples; The number of modification intervals D in the statistical identification sample that simultaneously contain the first modification label and the second modification label is greater than or equal to 0 and less than or equal to the total number of identification intervals L. Set the comparison thresholds DL1 and DL2 for the modified intervals, where 0 ≤ DL1 < DL2 ≤ L. Substitute the number of modified intervals D into the comparison thresholds DL1 and DL2 for comparison analysis. If the number of modified intervals D is greater than or equal to 0 and less than the comparison threshold DL1 of the modified intervals, then a true result is generated for the identified sample. If the number of modified intervals D is greater than or equal to the modified interval comparison threshold DL1 and less than the modified interval comparison threshold DL2, then a questionable result is generated for the identified sample. If the number of modified intervals D is greater than or equal to the modified interval comparison threshold DL2 and less than or equal to the total number of identification intervals L, then a modified result is generated for the identification sample. The degree of modification of the modified result is greater than that of the questionable result, and so on.

Citation Information

Patent Citations

  • Timing method, device and system for verifying time credibility of multi-source fusion video

    CN115662109A

  • Data authentic identification method and device, electronic equipment and storage medium

    CN114359811A