A Generalized Forged Media Detection Method for Realistic High-Traffic Scenarios
Through the multi-model joint decision-making and load balancing mechanism, the problems of traditional detection systems in high-traffic environments are solved, and efficient and accurate forged media detection are achieved.
Patent Information
- Application Number
- CN202510316802.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-18
AI Technical Summary
In the environment of high traffic Internet, traditional forged image and video detection systems are prone to encounter efficiency bottlenecks and resource overload problems when processing massive data, and the generalization ability and robustness of the single-model detection method are insufficient, making it difficult to adapt to diversified generation technologies.
The multi-model joint decision-making method is adopted to integrate multiple different detection models to improve the ability to respond to multiple forgery technologies and enhance the generalization performance and robustness of the algorithm. At the same time, a load balancing mechanism and standardized processing processes are introduced to improve the detection efficiency and reliability of the system.
It realizes efficient detection of forged media in large traffic scenarios, improves detection accuracy and reliability, and ensures that the system still maintains efficient performance under high load conditions.
Smart Images

Figure CN119854218B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically refers to a method for detecting generalized forged media for real large-traffic scenarios. Background Art
[0002] In the current Internet environment, the spread speed of forged images and videos increases rapidly with the large traffic of network platforms. Especially in high-concurrency scenarios such as social media and news websites, the diffusion range and speed of false information often exceed expectations, and the public is easily affected by these forged contents. Therefore, providing online detection of forged images / videos to the public is not only an inevitable choice to cope with the rapid spread of forged contents, but also an important means to maintain public trust and ensure network environment security. However, online forged detection needs to handle a large number of access requests. When traditional detection systems process large-scale data, they are prone to encounter problems such as efficiency bottlenecks and resource overload. In order to ensure that the detection system can still maintain efficient and stable performance when a large amount of data flows in, the present invention introduces a load balancing mechanism to reasonably allocate computing resources and avoid excessive single-point pressure, so as to ensure that the system can quickly process a large amount of data.
[0003] In this context of large traffic, the limitations of a single detection model become more prominent. Especially when the forgery means are increasingly diverse and complex, the detection difficulty increases significantly.
[0004] Existing methods for detecting generative models can be divided into three categories: biometric-based, statistical feature-based, and deep learning-based methods.
[0005] Biometric-based methods detect by using the slight differences in physiological characteristics between the generated image and the real image. These methods focus on capturing the natural biometric characteristics in real images, such as the regularity and consistency of pupil shapes, the consistency of eye specular reflections, or by detecting the natural manifestations of facial textures and detailed features to judge the authenticity of the image. These biometric characteristics are often interfered by forgery techniques during the generation process, resulting in unnatural manifestations, which become the basis for detection.
[0006] Statistical feature-based forged image detection methods identify forgery traces by analyzing the statistical attributes of images. For example, the gray-level co-occurrence matrix and cross-color-channel co-occurrence matrix of images are often used for texture analysis and serve as effective features for forgery detection. There is also frequency-domain feature extraction based on the discrete cosine transform (DCT), which identifies forgery patterns in images by converting the images to the frequency domain. By combining machine learning algorithms such as support vector machines (SVM), these statistical features can effectively identify forged images.
[0007] Deep learning-based methods automatically extract deep features in images by training neural networks, without relying on traditional hand-designed features. They are more adaptable and can effectively handle complex generation patterns. The core advantage of such methods lies in their ability to learn complex features of images from large-scale data and possess stronger discriminative capabilities. Network models such as VGGNet, ResNet, and Xception are often used for the detection of forged images or videos.
[0008] However, current forgery detection technologies face numerous challenges when dealing with large volumes of data. With the rapid development of generation technologies, detection systems need to process a large amount of data in a short period, which poses high requirements for the real-time performance and efficiency of algorithms. In a high-traffic environment, traditional detection methods may not be able to respond in a timely manner, resulting in potential forged content not being identified promptly. In addition, the system may experience a performance degradation under high load, thereby affecting the overall detection accuracy.
[0009] Furthermore, the limitations of single-model detection methods cannot be ignored. Although a single model performs well in certain specific tasks, its generalization ability and robustness are often insufficient, making it difficult to adapt to diverse generated images and videos. When faced with new generation technologies, a single model may not be able to comprehensively capture the features of forged images, leading to missed detections or false detections. In addition, the vulnerability of a single model when dealing with data noise and adversarial attacks also limits its effectiveness in practical applications. Therefore, single-model detection methods need to be combined with other technical means to improve the accuracy and reliability of their detections. Summary of the Invention
[0010] In view of the deficiencies of the prior art, the present invention proposes a generalization forgery media detection method for real-world large-traffic scenarios, introducing an algorithm for joint decision-making of multiple models. By integrating multiple different detection models, the ability to handle various forgery techniques is enhanced, and the generalization performance and robustness of the algorithm are strengthened. The detection efficiency of the system is improved through a modular process and streaming forgery detection, enabling efficient feedback even based on multiple models.
[0011] To solve the above technical problems, the technical solution of the present invention is as follows:
[0012] A generalization forgery media detection method for real-world large-traffic scenarios, comprising the following steps:
[0013] Step 1: Distribute the data of network requests to multiple servers through a weighted round-robin algorithm;
[0014] Step 2: Cache the access traffic of each server in an access queue, and each time, take out the data at the head of the queue for normalization processing to obtain a detection packet. The normalization processing includes frame division, cropping and segmentation, and encapsulation;
[0015] Step 3: Cache the detection packets in the detection queue, and each time, take out the detection packet at the head of the queue for forgery detection.
[0016] The method of the forgery detection is as follows:
[0017] First, construct a multi-modal forgery detection framework composed of three expert models: Xception (Model A), ResNet-50 (Model B), and ConvNeXt (Model C);
[0018] Then, construct a data set labeled with real images and corresponding generated forged images, and apply this data set to train and validate the multi-modal forgery detection network model;
[0019] Extract the detection packets to be detected, decompose and encapsulate the detection packets to obtain several detection clusters, decompose and encapsulate each detection cluster to obtain detection units, and the detection units output detection results through the multi-modal forgery detection network model.
[0020] Preferably, the data of the network request includes images and videos.
[0021] Preferably, in the step 1, the allocation method of the network request is specifically as follows:
[0022] Quantify the performance metrics of each server, select an image as a benchmark, and perform inference in all servers to obtain the inference time and memory occupancy ratio of all servers for this image, and obtain the performance quantization index of each server through the following formula:
[0023] ;
[0024] Among them, ;
[0025] Quantify the performance required for the network request data according to the size and number of key frames of the image or video frame. For an image with a size of , the quantization expression for the image or video frame is as follows:
[0026] ;
[0027] Among them is the number of key frames. For an image , for a video ;
[0028] Perform weighted round-robin according to the performance quantization of the server and the network request data.
[0029] Preferably, the method of the weighted round-robin is: First, initialize a value , which are respectively equal to , select servers in sequence , process requests in turn, and the performance requirement for processing each one is for an image or video, the corresponding subtract , if is insufficient, select the next server. When the of all current servers is insufficient, then is reset to the initial value.
[0030] Preferably, the method for video frame segmentation is: for a video, use the video codec tool FFmpeg to extract all key frames ( frames) in the video to obtain a frame set , where is the cardinality of the frame set, that is, the number of frames, and ; for an image, no frame segmentation is required, and the frame set only contains this image, that is ;
[0031] The cropping and segmentation method is: for each frame in the frame set, perform face cropping and image segmentation operations respectively;
[0032] The encapsulation method is: after cropping and segmentation, encapsulate the image blocks required by each model into corresponding detection units, and add the identifier of the specified detection model for the routing module in the detection stage to achieve the directional transmission of detection units.
[0033] Preferably, in step 3, three data sets containing real images and generated images of corresponding forgery types are respectively constructed, where the real images are labeled as 0, while the forged images are labeled as 1, 80% of the images are used as the training set, 10% of the images are used as the validation set, and 10% of the images are used as the test set.
[0034] Preferably, the training method of the multi-modal forgery detection framework is:
[0035] Use the three training sets as inputs to train Xception (model A), ResNet-50 (model B), and ConvNeXt (model C) respectively. Each model will predict the input image and finally output a confidence , representing the probability that this image is a forged image. The cross-entropy function used for training is:
[0036] ;
[0037] where represents the label of the image, that is, a real image , forged images , and the training objective of this loss is to make as small as possible until convergence;
[0038] Train the weights of the confidence of each model output. First, define the weights corresponding to the three models as , take the union of the images in the three validation sets, and each image can be predicted by the three trained models to obtain three confidences , and finally obtain a number of triples with real or forged labels as the dataset for weight training. First, aggregate the triple confidences, and then continuously train and optimize the three weights through this loss function , and the expression is as follows:
[0039] ;
[0040] ;
[0041] where represents the label of the image corresponding to the confidence triple, that is, a real image , forged images , and the training objective of this loss is to ensure that on the premise of making as small as possible until convergence, represents the aggregated confidence.
[0042] Preferably, in step 3, the three unpacked detection units apply a routing module to send them to the corresponding model for detection according to the identifiers on the detection units.
[0043] Preferably, in step 3, the confidences output by the three detection units in the detection cluster through the corresponding models are aggregated by the trained weights to obtain the confidence of each detection cluster , set a threshold , when , the image is predicted as a real image, , when , the image is predicted as a forged image, and the predicted label is recorded as
[0044] Preferably, the method for determining the threshold is: sort all from small to large, and select them as thresholds in turn. According to the authenticity judgment rule mentioned above, obtain the predicted label , and by comparing with obtain the prediction accuracy of this threshold. Finally select the threshold corresponding to the maximum prediction accuracy.
[0045] Preferably, in step 3, for video detection, a detection package usually contains multiple detection clusters, corresponding to multiple cluster confidence levels. Therefore, by combining sensitive parameters together for authenticity judgment, where , when the proportion of forged detection clusters among all detection clusters is greater than or equal to , the video will be determined to be forged. The smaller the sensitive parameter , the easier it is for the video to be judged as forged. Therefore, for a video containing detection clusters, when more than or equal to forged detection clusters have been obtained, the detection will not continue, and the detection result of forged will be directly returned.
[0046] In the present invention, a load balancing algorithm for dealing with both image and video formats is designed. For data of different formats, their performance requirements are uniformly quantified, and weighted round-robin is performed according to the performance of the server to achieve load balancing.
[0047] A set of standardized processing procedures is designed to encapsulate images and videos into a unified format, simplifying the data processing process and ensuring the compatibility and consistency between different input data types.
[0048] A multi-model joint decision-making mechanism is designed. For images and videos generated by multiple forgery methods, the generalization ability and detection accuracy of the system are improved by integrating the judgment results of multiple models.
[0049] Through load balancing, key frame extraction, image segmentation during the standardization process, and streaming authenticity judgment during the multi-model joint decision-making process, the present invention significantly improves the inference efficiency and achieves efficient detection feedback.
[0050] The present invention has the following characteristics and beneficial effects:
[0051] Adopting the above technical solution, this method is first designed based on the load balancing strategy for detecting large-traffic forged images and videos. By dynamically allocating computing resources and tasks, we can effectively process a large amount of data streams, ensuring that the system still maintains high detection performance under high load. On this basis, a multi-model joint decision-making framework is adopted, combining the outputs of multiple detection models to improve the recognition accuracy and robustness of forged content. Different models have their own advantages in dealing with diverse generation techniques, enabling the system to more comprehensively capture the characteristics of forged content. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0053] Figure 1 It is a schematic flowchart of the method according to the embodiment of the present invention.
[0054] Figure 2 It is a schematic flowchart of the data standardization process in the embodiment of the present invention.
[0055] Figure 3 It is a schematic flowchart of the multi-model joint detection process in the embodiment of the present invention. Detailed implementation manners
[0056] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0057] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "plurality" is two or more.
[0058] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "install", "connect", "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0059] Embodiment 1
[0060] The present invention provides a generalization forgery media detection method for real large - traffic scenarios, such as Figure 1 As shown, the overall framework consists of three parts: load balancing, data standardization, and multi - model joint decision - making.
[0061] Load balancing: To disperse traffic pressure, improve service performance, and ensure high availability of the system, the present invention distributes network requests to multiple servers through the weighted round - robin algorithm. Weights are assigned to the servers according to their processing capabilities, and the number of requests received by each server is proportional to its weight.
[0062] Specifically, the load - balancing implementation method is as follows:
[0063] Quantify the performance metrics of each server. Based on the memory and inference time of the server, quantify the performance metrics of each server as weights. Select an image as a benchmark, and perform inference in all servers to obtain the inference time and memory occupancy ratio of all servers for this image and memory occupancy ratio . The performance quantization index of each server is obtained through the following formula:
[0064] ;
[0065] where ;
[0066] Request performance quantization: Due to the different performance requirements for images and videos, quantify the performance required for network request data according to the size and number of key frames of the image or video frame. For an image or video frame with a size of , the quantization expression is as follows:
[0067] ;
[0068] where is the number of key frames. For an image , for a video ;
[0069] Weighted round - robin: Based on the quantization results of the first two steps, perform weighted round - robin to achieve load balancing. First, initialize values , which are respectively equal to . Select servers in order and process requests in turn. For each image or video with a performance requirement of , the corresponding is subtracted by . If is insufficient, select the next server. When the If it is insufficient, then is reset to the initial value.
[0070] Data standardization: The access traffic of each server is cached in the access queue. Each time, the image or video at the head of the queue is taken out for standardization processing, and they are encapsulated into detection packets with the same format. As Figure 2 shown, an image or video needs to go through three steps: frame splitting, cropping and segmentation, and encapsulation to obtain the corresponding detection packet.
[0071] It can be understood that the meaning of only taking out the head of the queue in this embodiment is that this queue is a buffer. Every time a detection request arrives, it is enqueued and placed at the end of the queue. According to the idea of first come, first served, each time a data (request) is taken out for detection, and after the detection is completed, the next one is taken out and processed in turn.
[0072] Specifically, the data standardization implementation method is as follows:
[0073] The method of video frame splitting is: for videos, use the video codec tool FFmpeg to extract all key frames ( frames) in the video to obtain the frame set , where is the cardinality of the frame set, that is, the number of frames, and ; for images, no frame splitting is required, and the frame set only contains this image, that is, ;
[0074] The described cropping and segmentation method is: for each frame in the frame set, perform face cropping and image segmentation operations respectively.
[0075] It should be noted that this step provides image blocks with the required content and size for all detection models. Xception (Model A) analyzes the face area of the image because DeepFake only tampered with the face area of the image, and the other areas are real; ResNet-50 (Model B) and ConvNeXt (Model C) do not need to crop out the face area because the Diffusion Model and GAN forgery methods detected by these two models are whole-image synthesis and there is no real area. However, if the image size is too large, it will cause problems such as excessive GPU load and slow inference speed. Therefore, the image needs to be segmented into several image blocks for parallel detection by the model.
[0076] The encapsulation method is as follows: After cropping and segmentation, the image patches required by each model are encapsulated into corresponding detection units, and the identifiers of the specified detection models are added to enable the routing module in the detection stage to send the detection units in a directed manner. The detection unit corresponding to Xception (Model A) contains only one cropped face image patch; the detection units corresponding to ResNet-50 (Model B) and ConvNeXt (Model C) contain segmented image patches. The detection units of the three detection models are further encapsulated to obtain detection clusters, and each detection cluster corresponds to one frame in the frame set. All the detection clusters of a video or image are encapsulated into a detection packet, and a detection packet is the standardized result of a video or image.
[0077] Joint decision-making of multiple models: After the image or video is standardized, detection packets with a unified format are obtained and cached in the detection queue. Each time, the detection packet at the head of the queue is taken out for forgery detection, and finally the detection results of the image or video corresponding to each detection packet are obtained.
[0078] As Figure 3 shown, the specific method is as follows:
[0079] Decapsulation and routing: The detection packet corresponding to each image or video is decapsulated to obtain several detection clusters, and each detection cluster is decapsulated to obtain three detection units. The routing module determines which model the detection unit should be sent to based on the identifier on the detection unit.
[0080] Model detection: A multi-modal forgery detection framework composed of Xception (Model A), ResNet-50 (Model B), and ConvNeXt (Model C) is constructed to detect images or videos of DeepFake, Diffusion Model, and GAN forgery types respectively.
[0081] Before being put into use, each model needs to be trained. The three models are trained using different data sets. Each data set contains real images and generated images of the corresponding forgery type. Among them, the real images are labeled as 0, while the forged images are labeled as 1. 80% of the images are used as the training set, 10% of the images are used as the validation set, and 10% of the images are used as the test set. When the model is trained on the training set, each model predicts the input image and finally outputs a confidence , representing the probability that this image is a forged image. The training uses the cross-entropy function:
[0082]
[0083] where represents the label of the image, that is, a real image , a forged image , the training objective of this loss is to make as small as possible until convergence, and the test set data is used for evaluation.
[0084] During the operation, the image patches in each detection unit are passed into the corresponding detection model as a batch, and the confidence of each detection unit can be obtained. Since there is only one face image in the detection unit corresponding to Xception (Model A), only one confidence will be obtained , while the detection units corresponding to ResNet-50 (Model B) and ConvNeXt (Model C) may contain multiple image patches. Therefore, the average value of the confidences of these image patches needs to be taken to obtain the confidence of this detection unit .
[0085] Joint decision-making: Before being put into use, in addition to training the model, it is also necessary to train the weights of the confidence of each detection unit. Models A, B, and C correspond to the weights . First, take the union of the images in the three validation sets. Each image can be predicted by the three trained models to obtain three confidences , and finally a number of triples with real or forged labels are obtained as the dataset for weight training. Then, the three weights are continuously trained and optimized through this loss function :
[0086] ;
[0087] ;
[0088] where represents the label of the image corresponding to the confidence triple, that is, a real image , a forged image , the training objective of this loss is to ensure that on the premise of making as small as possible until convergence, represents the aggregated confidence
[0089] During the operation, the confidences of the three detection units in the detection cluster will be obtained after model detection . Through the weights obtained by training, the confidence of each detection cluster is obtained by using the same calculation method as the confidence aggregation formula .
[0090] True or false streaming judgment: Before being put into use, it is necessary to pre-select the threshold for true or false judgment . When , the image is judged as a real image, When it is, the image is determined to be a forged image. In the joint decision-making step, we have obtained several confidence triples as the dataset for weight training. Here, the same dataset is also used in this embodiment to obtain the threshold . For each confidence triple , the aggregated confidence can be obtained through the trained weights and according to the confidence aggregation formula , and each corresponds to a true label . Sort all from small to large, and select them in turn as the threshold. According to the authenticity judgment rule mentioned above, the predicted label is obtained, and by comparing with , the prediction accuracy of this threshold is obtained. Finally , select the threshold corresponding to the maximum prediction accuracy.
[0091] It should be noted that during the actual use process, the joint decision-making will obtain the confidence of this detection cluster . By comparing with , the authenticity of this detection cluster is obtained. For an image, only one detection cluster is included in a detection package, so the authenticity of this detection cluster is the authenticity of this image.
[0092] Embodiment 2
[0093] The difference between this embodiment and Embodiment 1 is that during the detection process, when detecting a video, a detection package usually contains multiple detection clusters, corresponding to multiple cluster confidences. Therefore, by combining the sensitive parameter to make the authenticity judgment together, where is preset before actual use and is usually set to 0.5. When the proportion of forged detection clusters in all detection clusters is greater than or equal to , the video will be determined to be forged. The smaller the sensitive parameter , the easier it is for the video to be determined to be forged. Therefore, for a video containing detection clusters, when it has obtained more than or equal to forged detection clusters, the detection will not continue and the detection result will be directly returned as forged.
[0094] The above has described the embodiments of the present invention in detail in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations of these embodiments including components still fall within the protection scope of the present invention.
Claims
1. A generalized forged media detection method for realistic high-traffic scenarios, characterized in that: The steps include: Step 1: Distribute the data requested by the network to multiple servers through a weighted round-robin algorithm; Step 2: Cache the access traffic of each server in the access queue, and take out the data at the head of the queue each time to perform standardization processing to obtain a detection packet, wherein the standardization processing includes framing, clipping and segmentation, and encapsulation; Step 3: Cache the detection packets in the detection queue, and take out the detection packet at the head of the queue for forgery detection each time. The method of forgery detection is: Firstly, a multimodal forgery detection framework consisting of three expert models, Xception, ResNet-50 and ConvNeXt, is constructed; Then, a dataset with real images and corresponding generated forged images is constructed, and the dataset is used to train and verify the multimodal forgery detection network model; The training method of the multimodal forgery detection network model is: The three training sets are used as input to train Xception, ResNet-50 and ConvNeXt respectively. Each model predicts the input image and finally outputs a confidence level p, which represents the probability that the image is a forged image. The cross entropy function used in training is: L train =-y·log(p)+(1-y)·log(1-p); Where y represents the label of the image, that is, y = 0 for a real image and y = 1 for a fake image. The training goal of the cross entropy function is to make L train As small as possible until convergence; The weights of the confidence of each model output are trained. First, the weights corresponding to the three models are defined as ω A ,ω B ,ω C , the images in the three validation sets are combined, and each image is predicted by three trained models to obtain three confidence levels Finally, several triplets with real or fake labels are obtained as the data set for weight training. First, the triple confidence is aggregated, and then the three weights ω are continuously trained and optimized through the cross entropy function. A ,ω B ,ω C , the expression is as follows: Where y represents the label of the image corresponding to the confidence triplet, that is, y = 0 for the real image and y = 1 for the forged image. The training goal of the cross entropy function is to ensure ω A +ω B +ω C =1, so that L aggr As small as possible until convergence. Represents the confidence after aggregation; Extract the detection package to be detected, decapsulate the detection package to obtain n detection clusters, decapsulate each detection cluster to obtain a detection unit, and the detection unit outputs the detection result through the multimodal forgery detection network model. The confidence of the three detection units in the detection cluster output by the corresponding model is aggregated through the trained weights to obtain the confidence p of each detection cluster. T , set the threshold T, when p T When <T, the image is predicted to be a real image, p T ≥T, the image is predicted to be a forged image, and the predicted label is recorded as 2. According to claim 1, a generalized forged media detection method for realistic high-traffic scenarios is characterized in that: The data requested by the network includes images and videos.
3. According to claim 2, a generalized forged media detection method for realistic high-traffic scenarios is characterized in that: In step 1, the method for allocating network requests is as follows: Quantify the performance indicators of each server, select an image u×v as the benchmark, and perform inference in all R servers to obtain the inference time Inf1,…,Inf for all servers for the image R and memory usage ratio Prop1,…,Prop R , the performance quantitative index of each server is obtained by the following formula: Where r∈[1,R]; The performance required for network request data is quantified based on the size of the image or video frame and the number of key frames. The quantization expression of an image or video frame is as follows: Where m is the number of key frames, for images m = 1, for videos m ≥ 1; Weighted polling based on performance quantification of server and network request data.
4. According to claim 3, a generalized forged media detection method for realistic high-traffic scenarios is characterized in that: The weighted polling algorithm is as follows: first, initialize R values E1, ..., E R , which are equal to Perf1,…Perf R , select servers r∈[1,R] in order, process the requests in turn, and each time an image or video with a performance requirement of Req is processed, the corresponding E r Subtract Req, if E r If the number of E of all current servers is insufficient, the next server will be selected. r If E1,…,E R Reset to initial value.
5. According to claim 3, a generalized forged media detection method for realistic high-traffic scenarios is characterized in that: The frame division method is as follows: for a video, use the video codec tool FFmpeg to extract all key frames in the video to obtain a key frame set I = {I1, ...I n }, where n is the cardinality of the frame set, that is, the number of frames, and n≥1; for an image, no frame processing is required, and the frame set only contains the image, that is, n=1; The cropping and segmentation method is: for each frame in the frame set, perform face cropping and image segmentation operations respectively; The encapsulation method is as follows: after cropping and segmentation, the image blocks required by each model are encapsulated into corresponding detection units, and an identifier of a specified detection model is added to realize directional sending of the detection unit in a routing module in the detection phase.
6. According to claim 5, a generalized forged media detection method for realistic high-traffic scenarios is characterized in that: In step 3, three data sets containing real images and corresponding forged type generated images are constructed respectively, where the real images are labeled as 0 and the forged images are labeled as 1, 80% of the images are used as training sets, 10% of the images are used as validation sets, and 10% of the images are used as test sets.
7. According to claim 5, a generalized forged media detection method for realistic high-traffic scenarios is characterized in that: In step 3, the three decapsulated detection units are sent to the corresponding models for detection by a routing module according to the identifiers on the detection units.
8. The generalized forged media detection method for realistic high-traffic scenarios according to claim 6 is characterized in that: The method for determining the threshold T is: Sort from small to large and select them as thresholds in turn, according to the predicted label And by comparing y with The prediction accuracy of the threshold is obtained, and finally T selects the threshold corresponding to the maximum prediction accuracy.
9. The generalized forged media detection method for realistic high-traffic scenarios according to claim 6 is characterized in that: In step 3, for video detection, a detection package contains multiple detection clusters, corresponding to multiple cluster confidences, so the authenticity judgment is performed by combining the sensitive parameter S, where S∈(0,1], when the proportion of forged detection clusters in all detection clusters is greater than or equal to S, the video will be judged as forged. The smaller the sensitive parameter S is, the easier it is to judge the video as forged. Therefore, for a video containing N detection clusters, when greater than or equal to S×N forged detection clusters have been obtained, detection will not continue, and the detection result will be directly returned as forged.
Citation Information
Patent Citations
Video depth forgery detection method and device
CN116778545A
Image forgery detection method and device, server and storage medium
CN119048846A