Model training method and device, image detection method and device and related equipment
By training a set of image frames in multiple stages, a loss function optimization detection model is constructed, which solves the problem of inaccurate image tampering detection in existing technologies and achieves efficient image tampering detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-27
AI Technical Summary
In the existing technology, the detection models on electronic devices are not accurate enough in detecting image tampering, and the detection effect and performance need to be improved.
By acquiring a set of image frames, the first and second detection models are trained. By combining image features and predicted image frames, a loss function is constructed, and the model is optimized to improve detection accuracy.
It achieves accurate output of image tampering detection results, significantly enhances the feature discrimination power and robustness of the detection model, and improves detection effect and performance.
Smart Images

Figure CN121745341A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a model training method, an image detection method, an apparatus, and related equipment. Background Technology
[0002] With the development of internet technology, electronic devices need to detect tampered images. However, in related technologies, the detection results obtained by the detection models on electronic devices are not accurate enough, and the detection effect and performance of the detection models need to be improved. Summary of the Invention
[0003] This application provides a model training method, an image detection method, an apparatus, and related equipment, which can obtain a detection model that can output accurate image tampering detection results, and the detection model has good detection effect and high detection performance.
[0004] In a first aspect, embodiments of this application provide a model training method, the method comprising:
[0005] Get the image frame set;
[0006] The first image feature and the second image feature are input into the first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the first predicted image frame corresponding to each image frame. The first predicted image frame is the predicted image frame corresponding to each image frame output by the first prediction model.
[0007] Based on the image tampering detection results and the preset tampering labels corresponding to each image frame, a first loss function is obtained;
[0008] The first detection model is trained using the first loss function to obtain the second detection model.
[0009] Optionally, before inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model, the method further includes:
[0010] The image frame set is input into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model;
[0011] A second loss function is constructed based on each image frame and the corresponding second predicted image frame.
[0012] The second prediction model is trained using the second loss function to obtain the first prediction model.
[0013] Optionally, the second prediction model includes a first encoder and a first decoder; the step of inputting the image frame set into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model includes:
[0014] The target image frame is divided to obtain at least two image blocks corresponding to the target image frame; the target image frame is an image frame in the image frame set.
[0015] At least one image block in the target image frame is occluded to obtain the occluded target image frame;
[0016] The occluded target image frame is input into the first encoder to obtain the features corresponding to the occluded image block output by the first encoder.
[0017] The features corresponding to the occluded target image frame and the occluded image block are input into the first decoder to obtain the second predicted image frame corresponding to the target image frame output by the first decoder.
[0018] Optionally, the image frame set includes: at least two image frame subsets; the image frames in different image frame subsets are obtained by different shooting devices, and all image frames in the same image frame subset are obtained by the same shooting device;
[0019] The construction of the second loss function based on each image frame and the corresponding second predicted image frame includes:
[0020] A third loss function is determined based on at least one preset label and at least one second predicted image frame; the preset label is the label corresponding to each first image frame in at least one first image frame in the first image frame subset, and the at least two image frame subsets include the first image frame subset.
[0021] A fourth loss function is determined based on the first similarity and the second similarity; the first similarity is the similarity between image frames in the second image frame subset, and the second similarity is the similarity between image frames in the second image frame subset and image frames in the third image frame subset; the at least two image frame subsets include the second image frame subset and the third image frame subset; and the third image frame subset is different from the second image frame subset.
[0022] The second loss function is obtained by weighted summation of the third loss function and the fourth loss function.
[0023] Optionally, determining the fourth loss function based on the first and second similarities includes:
[0024] The fourth loss function is determined using the first formula;
[0025] The first formula is: ;
[0026] in, For the first similarity, For the second similarity, and These are the predicted image frames corresponding to the second and third image frames, where the second and third image frames are image frames from a subset of the second image frames. The predicted image frame is the image frame corresponding to the fourth image frame, which is an image frame in the subset of the third image frames.
[0027] Optionally, the step of inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model includes:
[0028] Align the first image features and the second image features;
[0029] The aligned first image features and the aligned second image features are concatenated to obtain the concatenated third image features;
[0030] The third image feature is input into the first detection model to obtain the image tampering detection result output by the first detection model.
[0031] Optionally, the first loss function is: ;
[0032] in, The image tampering detection result, The preset tampering label.
[0033] Secondly, embodiments of this application provide an image detection method, the method comprising:
[0034] Acquire the image frame to be detected;
[0035] The image frame to be detected is input into the first prediction model to obtain the third predicted image frame output by the first prediction model;
[0036] The fourth image feature and the fifth image feature are input into the second detection model to obtain the image tampering detection result output by the second detection model. The fourth image feature is the image feature corresponding to the third predicted image frame, and the fifth image feature is the image feature corresponding to the image frame to be detected.
[0037] The second detection model is a model trained according to the model training method described in the first aspect.
[0038] Thirdly, embodiments of this application provide a model training apparatus, including:
[0039] The acquisition module is used to acquire a set of image frames;
[0040] The first determining module is used to input the first image feature and the second image feature into the first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the first predicted image frame corresponding to each image frame. The first predicted image frame is the predicted image frame corresponding to each image frame output by the first prediction model.
[0041] The second determining module is used to obtain a first loss function based on the image tampering detection result and the preset tampering label corresponding to each image frame;
[0042] The training module is used to train the first detection model using the first loss function to obtain the second detection model.
[0043] Fourthly, embodiments of this application provide an image detection apparatus, including:
[0044] The acquisition module is used to acquire the image frames to be detected;
[0045] The first determining module is used to input the image frame to be detected into the first prediction model to obtain the third predicted image frame output by the first prediction model.
[0046] The second determining module is used to input the fourth image feature and the fifth image feature into the second detection model to obtain the image tampering detection result output by the second detection model. The fourth image feature is the image feature corresponding to the third predicted image frame, and the fifth image feature is the image feature corresponding to the image frame to be detected.
[0047] The second detection model is a model trained by the model training method according to any one of claims 1-7.
[0048] Fifthly, an electronic device includes: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the model training method as described in the first aspect, and the steps of the image detection method as described in the second aspect.
[0049] In a sixth aspect, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the model training method as described in the first aspect, and the steps of the image detection method as described in the second aspect.
[0050] In a seventh aspect, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the model training method as described in the first aspect, and the steps of the image detection method as described in the second aspect.
[0051] In this embodiment, the electronic device acquires a set of image frames. The electronic device then inputs a first image feature and a second image feature into a first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the predicted image frame corresponding to each image frame output by the first prediction model. The electronic device then obtains a first loss function based on the image tampering detection result and the preset tampering labels corresponding to each image frame. The electronic device then trains the first detection model to obtain a trained second detection model. It is understood that in this process, the first loss function is determined based on the first and second image features. This allows the detection model trained by the electronic device through the first loss function to have a cross-modal feature fusion tampering region detection mechanism, significantly enhancing the feature discrimination and robustness of the detection model. This results in a detection model that can output accurate image tampering detection results, and the detection model has good detection effect and high detection performance. Attached Figure Description
[0052] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 This is a flowchart of a model training method provided in an embodiment of this application;
[0054] Figure 2This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;
[0055] Figure 3 This is a schematic diagram of the structure of an image detection device provided in an embodiment of this application;
[0056] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0058] See Figure 1 , Figure 1 This is a flowchart of a model training method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0059] Step 101: Obtain the image frame set.
[0060] In some embodiments, the image frame set includes at least two image frame subsets; the image frames in different image frame subsets are obtained by different shooting devices, and all image frames in the same image frame subset are obtained by the same shooting device.
[0061] In this embodiment, the electronic device can acquire video frames or image frames captured by different shooting devices, wherein each video frame includes at least one image frame. Furthermore, the electronic device can obtain a set of image frames based on the video frames or image frames captured by the different shooting devices. For example, suppose the set of shooting devices is... , Let be the i-th capturing device. Each capturing device may include at least one of the following: a video sequence, multiple video sequences, a single image frame, and multiple image frames, i.e., Each shooting device has at least two video frames or at least two image frames that constitute a subset of image frames. That is, different subsets of image frames correspond to different shooting devices, and all image frames in the same subset of image frames are obtained by the same shooting device.
[0062] Step 102: Input the first image feature and the second image feature into the first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the first predicted image frame corresponding to each image frame. The first predicted image frame is the predicted image frame corresponding to each image frame output by the first prediction model.
[0063] In some embodiments, before inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model, the method further includes:
[0064] The image frame set is input into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model;
[0065] A second loss function is constructed based on each image frame and the corresponding second predicted image frame.
[0066] The second prediction model is trained using the second loss function to obtain the first prediction model.
[0067] In this embodiment, before inputting the first image features and the second image features into the first detection model, the electronic device needs to train the second prediction model to obtain the trained first prediction model. That is, in the embodiments of this application, the model training process is divided into two stages: the electronic device first trains the prediction model, and then trains the detection model after the prediction model training is completed.
[0068] Understandably, by training the model in two stages in this embodiment, the predicted image frame information output by the prediction model can be made richer after the prediction model is trained, reducing information loss and making it easier for subsequent electronic devices to obtain more accurate image features through the predicted image frames. The detection model trained with the image features corresponding to the predicted image frames can have stronger feature discrimination and robustness.
[0069] In some embodiments, the detection model and the prediction model can also be integrated into a single image model. The electronic device can also train the image model in stages; that is, the electronic device can first train the output prediction image frame portion of the image model, and then train the output image tampering detection result portion.
[0070] In some embodiments, the second prediction model includes a first encoder and a first decoder; the step of inputting the image frame set into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model includes:
[0071] The target image frame is divided to obtain at least two image blocks corresponding to the target image frame; the target image frame is an image frame in the image frame set.
[0072] At least one image block in the target image frame is occluded to obtain the occluded target image frame;
[0073] The occluded target image frame is input into the first encoder to obtain the features corresponding to the occluded image block output by the first encoder.
[0074] The features corresponding to the occluded target image frame and the occluded image block are input into the first decoder to obtain the second predicted image frame corresponding to the target image frame output by the first decoder.
[0075] In this embodiment, the electronic device first performs block processing on the target image frame, dividing the image frame into N non-overlapping image blocks of fixed size, and then maps the image blocks. The electronic device then occludes a preset proportion (e.g., 75%) of the image blocks, retaining only a small portion of the image blocks that are visible (i.e., the unoccluded image blocks). The electronic device then inputs this small portion of the visible image blocks into the first encoder. In the first encoder, the features corresponding to the occluded image patch are obtained. The electronic device then inputs the features corresponding to the unoccluded image patch and the occluded image patch into the first decoder. The predicted image frame corresponding to the target image frame is obtained. The predicted image frame is also called the predicted photo response non-uniformity (PRNU) image frame.
[0076] Understandably, in this embodiment, the electronic device employs the occlusion concept to directly learn the image-to-PRNU mapping within the encoder, integrating PRNU feature extraction and task representation learning. This reduces information loss and improves the adaptability of features to downstream detection tasks. In other words, it achieves end-to-end PRNU feature learning, avoiding noise accumulation and information loss from independent preprocessing. Furthermore, by introducing the PRNU feature learning task during the prediction model training process, the electronic device can directly achieve end-to-end image-to-PRNU reconstruction on the encoder, without relying on an external PRNU extractor. This changes the process in related technologies that relies on preprocessing to extract PRNU features, enabling PRNU feature extraction and representation learning to be co-optimized within the same network.
[0077] In some embodiments, when determining the predicted image frame output by the second prediction model, the electronic device also needs to determine the preset label corresponding to the image frame in the image frame set, so that a third loss function can be constructed based on the predicted image frame and the image frame in the image frame set.
[0078] In some embodiments, the electronic device can calculate a preset tag, i.e., a PRNU tag, using a PRNU extraction method (e.g., a content removal operation based on high-pass filtering). .
[0079] Its formula can be: ; .
[0080] in, These are the image frames in the image frame set. The PRNU tag corresponding to the image frame. It will serve as the reconstruction target during the pre-training phase of the prediction model. This represents the denoising function. , These are the mean and standard deviation, respectively. , The value is calculated by applying a normal distribution to all image frames within the subset of image frames to which the image frame belongs. N represents the number of sampled image frames, which can come from different subsets of image frames, i.e., images captured by different imaging devices. Electronic devices can employ 3D block matched filtering (BM3D) to remove large-scale content information, retaining only sensor noise components.
[0081] In some embodiments, the third loss function may be: .
[0082] Where M is the number of devices to which the sampled image frames belong. This is the predicted image frame corresponding to the original image frame.
[0083] In this embodiment, the electronic device can infer the complete PRNU noise distribution from partially visible image patches using a third loss function, thereby capturing and learning the features of the PRNU.
[0084] In some embodiments, constructing a second loss function based on each image frame and the second predicted image frame corresponding to each image frame includes:
[0085] A third loss function is determined based on at least one preset label and at least one second predicted image frame; the preset label is the label corresponding to each first image frame in at least one first image frame in the first image frame subset, and the at least two image frame subsets include the first image frame subset.
[0086] A fourth loss function is determined based on the first similarity and the second similarity; the first similarity is the similarity between image frames in the second image frame subset, and the second similarity is the similarity between image frames in the second image frame subset and image frames in the third image frame subset; the at least two image frame subsets include the second image frame subset and the third image frame subset; and the third image frame subset is different from the second image frame subset.
[0087] The second loss function is obtained by weighted summation of the third loss function and the fourth loss function.
[0088] In this embodiment, the electronic device obtains the second loss function through the third and fourth loss functions. The fourth loss function can serve as a constraint term for the PRNU characteristics. This is because the PRNU textures generated by the same shooting device under different shooting conditions exhibit high consistency, while the PRNU patterns of different shooting devices show statistically significant differences. The electronic device explicitly utilizes this characteristic. The fourth loss function is obtained through the first and second similarities. The first similarity is determined by image frames in the image frame subset corresponding to the same shooting device; therefore, the first similarity can characterize the consistency of the same shooting device. The second similarity is determined by image frames in the image frame subset corresponding to different shooting devices; therefore, the second similarity can characterize the differences between different shooting devices. Therefore, the fourth loss function is a comparative loss jointly optimized for both the same and different shooting devices, facilitating the subsequent acquisition of a more suitable second loss function by the electronic device and achieving accurate training of the prediction model.
[0089] In some embodiments, determining the fourth loss function based on the first similarity and the second similarity includes:
[0090] The fourth loss function is determined using the first formula;
[0091] The first formula is: ;
[0092] in, For the first similarity, For the second similarity, and These are the predicted image frames corresponding to the second and third image frames, where the second and third image frames are image frames from a subset of the second image frames. The predicted image frame is the image frame corresponding to the fourth image frame, which is an image frame in the subset of the third image frames.
[0093] In this embodiment, the first similarity characterizes the difference between PRNU predicted image frames from the same shooting device, with a value close to 1. The second similarity characterizes the difference between PRNU predicted image frames from different shooting devices. The electronic device samples image frames from multiple shooting devices and constructs corresponding constraints. These constraints simultaneously incorporate consistency constraints within the same shooting device and separability constraints between different shooting devices. Based on these constraints, a loss function is determined, and the prediction model is trained using this loss function. This significantly enhances the PRNU feature discriminative power and robustness of the trained prediction model, resulting in a stronger device differentiation capability.
[0094] In some embodiments, electronic devices can be implemented using the following formula: Thus, the second loss function is obtained.
[0095] Here, λ is a weight hyperparameter used to balance reconstruction accuracy (i.e., the accuracy of predicted images) and feature discriminability. The gradient optimizes the parameters of the first encoder and the first decoder simultaneously during backpropagation. That is, the electronic device can optimize the parameters of the first encoder and the first decoder through the second loss function, so that the first encoder not only learns the global PRNU space distribution, but also forms separable embedding representations between different shooting devices in the feature space.
[0096] In some embodiments, to ensure the effectiveness of the subsequent loss function, each sample (or image frame) corresponding to each shooting device in the image frame set has at least two samples from different video frames. That is, the image frame set is not just a copy of frames from the same video, but a PRNU under different shooting conditions, thereby avoiding the model from only capturing pseudo-similarity related to image content.
[0097] In some embodiments, the step of inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model includes:
[0098] Align the first image features and the second image features;
[0099] The aligned first image features and the aligned second image features are concatenated to obtain the concatenated third image features;
[0100] The third image feature is input into the first detection model to obtain the image tampering detection result output by the first detection model.
[0101] In this embodiment, after the prediction model training is completed, the electronic device first inputs each image frame in the image frame set to the first encoder to obtain the first image feature. Simultaneously, the electronic device inputs each image frame in the image frame set to the first encoder and the first decoder to obtain the predicted image frame corresponding to each image frame. The electronic device then obtains the second image feature from the predicted image frames corresponding to each image frame. The electronic device then performs feature alignment between the first and second image features to obtain the second image feature that matches the size of the first image feature. For example, the size of the first image feature could be... The format is as follows: b is the batch size corresponding to the image frame set, n is the number of extracted frames corresponding to the video frame set, x is the number of features, and h is the dimension of the hidden features. Electronic devices can align the second image features based on the first image feature size. After feature size alignment, the two image features are concatenated along the channel dimension to obtain the third image feature. .
[0102] The electronic device then inputs the third image feature into the second segmentation decoder included in the detection model. In this process, the image tampering detection result output by the second segmentation decoder is obtained. This image tampering detection result can be a tampering probability map with the same feature resolution as the third image. In the tampering probability map, the value of each pixel position represents the probability that the position belongs to the tampered area.
[0103] In this embodiment, the electronic device can obtain the image tampering detection result, which facilitates the subsequent acquisition of the first loss function by using the image tampering detection result and the preset tampering label corresponding to each image frame.
[0104] Step 103: Based on the image tampering detection results and the preset tampering labels corresponding to each image frame, obtain the first loss function.
[0105] In some embodiments, the first loss function is: ;
[0106] in, The image tampering detection result, The preset tampering label.
[0107] In this embodiment, the electronic device can obtain the first loss function through pixel-level binary cross-entropy (BCE), which facilitates the subsequent training of the detection model through the loss function.
[0108] Step 104: Train the first detection model using the first loss function to obtain the second detection model.
[0109] In this embodiment, the trained first prediction model has the ability to infer the complete PRNU noise distribution from partially visible image patches, and has formed a consistent embedding representation within the same shooting device and distinguishable between different shooting devices in the feature space. At this time, the first encoder in the first prediction model outputs high-dimensional image semantic features, and the first decoder outputs the corresponding PRNU reconstructed predicted image frame. The electronic device can then use the first loss function to further refine the second segmentation decoder included in the detection model. The system is trained to simultaneously utilize image features and PRNU features for tamper detection and localization.
[0110] In some embodiments, considering that the first encoder and first decoder have been trained, a relatively accurate image and PRNU representation can be output. During the training of the detection model, the electronic device can assign a small learning rate to the first encoder and first decoder to fine-tune them, preventing excessive drift of existing features and loss of PRNU discriminative ability during the training process. The electronic device can also randomly initialize the parameters of the second segmentation decoder and train it using a normal learning rate, enabling it to quickly adapt to and fully utilize bimodal features.
[0111] Understandably, in this embodiment, the electronic device adds a second segmentation decoder, enabling it to simultaneously receive image semantic features from the first encoder and PRNU features output by the first decoder, which are then used as input to the second segmentation decoder, forming a cross-modal feature fusion tampering region detection mechanism. Furthermore, by fusing image semantic features and PRNU features, the detection model can capture both the visual semantic differences of the tampered region and utilize the physical characteristics of sensor noise for auxiliary judgment, thereby improving the accuracy and robustness of the detection.
[0112] Through the training process of the above model, the embodiments of this application can learn robust PRNU feature representations under unsupervised conditions and transfer them to supervised tampering detection tasks, effectively improving detection accuracy and robustness in highly similar tampering scenarios. Compared with related technologies, the embodiments of this application do not rely on a dedicated PRNU extraction network, but directly integrate PRNU modeling into the model training, achieving deep coupling between feature learning and noise pattern modeling, thereby obtaining stronger generalization ability.
[0113] In practical applications, the model trained using the embodiments of this application achieves a test accuracy of 96%, which is 4 percentage points higher than the 92% accuracy of deep learning models in related technologies under the same settings. This fully verifies the accuracy and robustness of the embodiments of this application in tampering scenarios.
[0114] As described in steps 101-104 above, the electronic device acquires a set of image frames. The electronic device then inputs the first image feature and the second image feature into the first detection model to obtain the image tampering detection result output by the first detection model. Here, the first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the predicted image frame corresponding to each image frame output by the first prediction model. The electronic device then obtains a first loss function based on the image tampering detection result and the preset tampering labels corresponding to each image frame. The electronic device then trains the first detection model to obtain a trained second detection model. It can be understood that in this process, the first loss function is determined based on the first and second image features. This allows the detection model trained by the electronic device through the first loss function to have a cross-modal feature fusion tampering region detection mechanism, significantly enhancing the feature discrimination and robustness of the detection model. This results in a detection model that can output accurate image tampering detection results, and the detection model has good detection effect and high detection performance.
[0115] In some embodiments, this application further includes an image detection method, the method comprising:
[0116] Acquire the image frame to be detected;
[0117] The image frame to be detected is input into the first prediction model to obtain the third predicted image frame output by the first prediction model;
[0118] The fourth and fifth image features are input into the second detection model to obtain the image tampering detection result output by the second detection model. The fourth image feature is the image feature corresponding to the third predicted image frame, and the fifth image feature is the image feature corresponding to the image frame to be detected.
[0119] The second detection model and the first prediction model are models trained according to the model training method in the above embodiments.
[0120] In this embodiment, after both the second detection model and the first prediction model have been trained, the electronic device can acquire the image frame to be detected and input it into the first encoder and the first decoder of the first prediction model to obtain the predicted image frame corresponding to the image frame to be detected. The electronic device then inputs the predicted image frame into the first encoder to obtain the image features corresponding to the predicted image frame output by the first encoder, and inputs the image frame to be detected into the first encoder to obtain the image features corresponding to the image frame to be detected output by the first encoder. The electronic device then inputs the image features corresponding to the predicted image frame and the image features corresponding to the image frame to be detected into the second segmentation decoder of the second detection model to obtain an accurate image tampering detection result.
[0121] Understandably, the trained detection model can output accurate image tampering detection results, and the detection model has good detection effect and high detection performance.
[0122] Please refer to Figure 2 , Figure 2 This is a schematic diagram of a model training device according to an embodiment of this application. The model training device 200 includes:
[0123] Acquisition module 201 is used to acquire a set of image frames;
[0124] The first determining module 202 is used to input the first image feature and the second image feature into the first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the first predicted image frame corresponding to each image frame. The first predicted image frame is the predicted image frame corresponding to each image frame output by the first prediction model.
[0125] The second determining module 203 is used to obtain a first loss function based on the image tampering detection result and the preset tampering label corresponding to each image frame;
[0126] The training module 204 is used to train the first detection model using the first loss function to obtain the second detection model.
[0127] Optionally, the model training device 200 may also include:
[0128] The third determining module is used to input the image frame set into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model;
[0129] The construction module is used to construct a second loss function based on each image frame and the second predicted image frame corresponding to each image frame;
[0130] The fourth determining module is used to train the second prediction model using the second loss function to obtain the first prediction model.
[0131] Optionally, the second prediction model includes a first encoder and a first decoder; the third determination module may further include:
[0132] A segmentation unit is used to segment a target image frame to obtain at least two image blocks corresponding to the target image frame; the target image frame is an image frame in the image frame set;
[0133] An occlusion unit is used to occlude at least one image block in the target image frame to obtain an occluded target image frame.
[0134] The first determining unit is used to input the occluded target image frame into the first encoder to obtain the features corresponding to the occluded image block output by the first encoder.
[0135] The second determining unit is used to input the features corresponding to the occluded target image frame and the occluded image block into the first decoder to obtain the second predicted image frame corresponding to the target image frame output by the first decoder.
[0136] Optionally, the image frame set includes: at least two image frame subsets; the image frames in different image frame subsets are obtained by different shooting devices, and all image frames in the same image frame subset are obtained by the same shooting device;
[0137] Build modules may also include:
[0138] The third determining unit is used to determine a third loss function based on at least one preset label and at least one second predicted image frame; the preset label is the label corresponding to each first image frame in at least one first image frame in the first image frame subset, and the at least two image frame subsets include the first image frame subset.
[0139] The fourth determining unit is used to determine a fourth loss function based on a first similarity and a second similarity; the first similarity is the similarity between image frames in the second image frame subset, and the second similarity is the similarity between image frames in the second image frame subset and image frames in the third image frame subset; the at least two image frame subsets include the second image frame subset and the third image frame subset; and the third image frame subset is different from the second image frame subset.
[0140] The weighting unit is used to perform a weighted summation of the third loss function and the fourth loss function to obtain the second loss function.
[0141] Optionally, the fourth determining unit may also include:
[0142] The first determining subunit is used to determine the fourth loss function using a first formula;
[0143] The first formula is: ;
[0144] in, For the first similarity, For the second similarity, and These are the predicted image frames corresponding to the second and third image frames, where the second and third image frames are image frames from a subset of the second image frames. The predicted image frame is the image frame corresponding to the fourth image frame, which is an image frame in the subset of the third image frames.
[0145] Optionally, the first determining module 202 may further include:
[0146] An alignment unit is used to align the first image features and the second image features;
[0147] The stitching unit is used to stitch together the aligned first image features and the aligned second image features to obtain the stitched third image features;
[0148] The fifth determining unit is used to input the third image feature into the first detection model to obtain the image tampering detection result output by the first detection model.
[0149] Optionally, the first loss function is: ;
[0150] in, The image tampering detection result, The preset tampering label.
[0151] The model training device 200 provided in this application embodiment can perform the above-described... Figure 1 The method embodiments shown are similar in principle and technical effect, and will not be described again here.
[0152] Please refer to Figure 3 , Figure 3 This is a schematic diagram of an image detection device 300 according to an embodiment of this application. The image detection device 300 includes:
[0153] The acquisition module 301 is used to acquire the image frame to be detected;
[0154] The first determining module 302 is used to input the image frame to be detected into the first prediction model to obtain the third predicted image frame output by the first prediction model.
[0155] The second determining module 303 is used to input the fourth image feature and the fifth image feature into the second detection model to obtain the image tampering detection result output by the second detection model. The fourth image feature is the image feature corresponding to the third predicted image frame, and the fifth image feature is the image feature corresponding to the image frame to be detected.
[0156] The second detection model and the first prediction model are based on the above. Figure 1The model trained by the model training method in the illustrated embodiment is shown.
[0157] This application also provides an electronic device. Since the principle by which this electronic device solves the problem is similar to the model training method in the embodiments of this application, the implementation of this electronic device can be found elsewhere. Figure 1 The implementation of the method shown will not be repeated here. Figure 4 As shown, the electronic device according to an embodiment of this application includes: a processor 410, configured to read a program from a memory 420 and execute the following processes:
[0158] Get the image frame set;
[0159] The first image feature and the second image feature are input into the first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the first predicted image frame corresponding to each image frame. The first predicted image frame is the predicted image frame corresponding to each image frame output by the first prediction model.
[0160] Based on the image tampering detection results and the preset tampering labels corresponding to each image frame, a first loss function is obtained;
[0161] The first detection model is trained using the first loss function to obtain the second detection model.
[0162] Optionally, the processor 410 is also used to read the program from the memory 420 and perform the following steps:
[0163] Before inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model, the method further includes:
[0164] The image frame set is input into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model;
[0165] A second loss function is constructed based on each image frame and the corresponding second predicted image frame.
[0166] The second prediction model is trained using the second loss function to obtain the first prediction model.
[0167] Optionally, the second prediction model includes a first encoder and a first decoder; the processor 410 is further configured to read the program in the memory 420 and perform the following steps: inputting the image frame set into the second prediction model to obtain the second predicted image frames corresponding to each image frame in the image frame set output by the second prediction model, including:
[0168] The target image frame is divided to obtain at least two image blocks corresponding to the target image frame; the target image frame is an image frame in the image frame set.
[0169] At least one image block in the target image frame is occluded to obtain the occluded target image frame;
[0170] The occluded target image frame is input into the first encoder to obtain the features corresponding to the occluded image block output by the first encoder.
[0171] The features corresponding to the occluded target image frame and the occluded image block are input into the first decoder to obtain the second predicted image frame corresponding to the target image frame output by the first decoder.
[0172] Optionally, the image frame set includes: at least two image frame subsets; the image frames in different image frame subsets are obtained by different shooting devices, and all image frames in the same image frame subset are obtained by the same shooting device;
[0173] Processor 410 is also used to read the program in memory 420 and to perform the following steps: constructing a second loss function based on each image frame and the second predicted image frame corresponding to each image frame, including:
[0174] A third loss function is determined based on at least one preset label and at least one second predicted image frame; the preset label is the label corresponding to each first image frame in at least one first image frame in the first image frame subset, and the at least two image frame subsets include the first image frame subset.
[0175] A fourth loss function is determined based on the first similarity and the second similarity; the first similarity is the similarity between image frames in the second image frame subset, and the second similarity is the similarity between image frames in the second image frame subset and image frames in the third image frame subset; the at least two image frame subsets include the second image frame subset and the third image frame subset; and the third image frame subset is different from the second image frame subset.
[0176] The second loss function is obtained by weighted summation of the third loss function and the fourth loss function.
[0177] Optionally, the processor 410 is also configured to read the program from the memory 420 and perform the following steps: determining the fourth loss function based on the first and second similarities includes:
[0178] The fourth loss function is determined using the first formula;
[0179] The first formula is: ;
[0180] in, For the first similarity, For the second similarity, and These are the predicted image frames corresponding to the second and third image frames, where the second and third image frames are image frames from a subset of the second image frames. The predicted image frame is the image frame corresponding to the fourth image frame, which is an image frame in the subset of the third image frames.
[0181] Optionally, the processor 410 is also configured to read the program in the memory 420 and perform the following steps: inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model, including:
[0182] Align the first image features and the second image features;
[0183] The aligned first image features and the aligned second image features are concatenated to obtain the concatenated third image features;
[0184] The third image feature is input into the first detection model to obtain the image tampering detection result output by the first detection model.
[0185] Optionally, the first loss function is: ;
[0186] in, The image tampering detection result, The preset tampering label.
[0187] Among them, Figure 4 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 410 and memory represented by memory 420 together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides the interface.
[0188] The electronic device provided in this application embodiment can perform the above-described functions. Figure 1 The method embodiments shown are similar in principle and technical effect, and will not be described again here.
[0189] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described model training method and image detection method embodiments, achieving the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may include read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0190] This application also provides a computer program product, including computer instructions. When executed by a processor, these computer instructions implement the various processes of the above-described model training method and image detection method embodiments, and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0191] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0192] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0193] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A model training method, characterized in that, The method includes: Get the image frame set; The first image feature and the second image feature are input into the first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the first predicted image frame corresponding to each image frame. The first predicted image frame is the predicted image frame corresponding to each image frame output by the first prediction model. Based on the image tampering detection results and the preset tampering labels corresponding to each image frame, a first loss function is obtained; The first detection model is trained using the first loss function to obtain the second detection model.
2. The method according to claim 1, characterized in that, Before inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model, the method further includes: The image frame set is input into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model; A second loss function is constructed based on each image frame and the corresponding second predicted image frame. The second prediction model is trained using the second loss function to obtain the first prediction model.
3. The method according to claim 2, characterized in that, The second prediction model includes a first encoder and a first decoder; the step of inputting the image frame set into the second prediction model to obtain the second predicted image frame corresponding to each image frame in the image frame set output by the second prediction model includes: The target image frame is divided to obtain at least two image blocks corresponding to the target image frame; the target image frame is an image frame in the image frame set. At least one image block in the target image frame is occluded to obtain the occluded target image frame; The occluded target image frame is input into the first encoder to obtain the features corresponding to the occluded image block output by the first encoder. The features corresponding to the occluded target image frame and the occluded image block are input into the first decoder to obtain the second predicted image frame corresponding to the target image frame output by the first decoder.
4. The method according to claim 2, characterized in that, The image frame set includes at least two image frame subsets; the image frames in different image frame subsets are obtained by different shooting devices, and all image frames in the same image frame subset are obtained by the same shooting device; The construction of the second loss function based on each image frame and the corresponding second predicted image frame includes: A third loss function is determined based on at least one preset label and at least one second predicted image frame; the preset label is the label corresponding to each first image frame in at least one first image frame in the first image frame subset, and the at least two image frame subsets include the first image frame subset. A fourth loss function is determined based on the first similarity and the second similarity; the first similarity is the similarity between image frames in the second image frame subset, and the second similarity is the similarity between image frames in the second image frame subset and image frames in the third image frame subset; the at least two image frame subsets include the second image frame subset and the third image frame subset; and the third image frame subset is different from the second image frame subset. The second loss function is obtained by weighted summation of the third loss function and the fourth loss function.
5. The method according to claim 4, characterized in that, The determination of the fourth loss function based on the first and second similarities includes: The fourth loss function is determined using the first formula; The first formula is: ; in, For the first similarity, For the second similarity, and These are the predicted image frames corresponding to the second and third image frames, where the second and third image frames are image frames from a subset of the second image frames. The predicted image frame is the image frame corresponding to the fourth image frame, which is an image frame in the subset of the third image frames.
6. The method according to claim 1, characterized in that, The step of inputting the first image features and the second image features into the first detection model to obtain the image tampering detection result output by the first detection model includes: Align the first image features and the second image features; The aligned first image features and the aligned second image features are concatenated to obtain the concatenated third image features; The third image feature is input into the first detection model to obtain the image tampering detection result output by the first detection model.
7. The method according to claim 1, characterized in that, The first loss function is: ; in, The image tampering detection result, The preset tampering label.
8. An image detection method, characterized in that, include: Acquire the image frame to be detected; The image frame to be detected is input into the first prediction model to obtain the third predicted image frame output by the first prediction model; The fourth image feature and the fifth image feature are input into the second detection model to obtain the image tampering detection result output by the second detection model. The fourth image feature is the image feature corresponding to the third predicted image frame, and the fifth image feature is the image feature corresponding to the image frame to be detected. The second detection model is a model trained by the model training method according to any one of claims 1-7.
9. A model training device, characterized in that, include: The acquisition module is used to acquire a set of image frames; The first determining module is used to input the first image feature and the second image feature into the first detection model to obtain the image tampering detection result output by the first detection model. The first image feature is the image feature corresponding to each image frame in the image frame set, and the second image feature is the feature of the first predicted image frame corresponding to each image frame. The first predicted image frame is the predicted image frame corresponding to each image frame output by the first prediction model. The second determining module is used to obtain a first loss function based on the image tampering detection result and the preset tampering label corresponding to each image frame; The training module is used to train the first detection model using the first loss function to obtain the second detection model.
10. An image detection device, characterized in that, include: The acquisition module is used to acquire the image frames to be detected; The first determining module is used to input the image frame to be detected into the first prediction model to obtain the third predicted image frame output by the first prediction model. The second determining module is used to input the fourth image feature and the fifth image feature into the second detection model to obtain the image tampering detection result output by the second detection model. The fourth image feature is the image feature corresponding to the third predicted image frame, and the fifth image feature is the image feature corresponding to the image frame to be detected. The second detection model is a model trained by the model training method according to any one of claims 1-7.
11. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps in the model training method as claimed in any one of claims 1 to 7 and the steps in the image detection method as claimed in claim 8.
12. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps in the model training method as described in any one of claims 1 to 7 and the steps in the image detection method as described in claim 8.
13. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps in the model training method as described in any one of claims 1 to 7 and the steps in the image detection method as described in claim 8.